Guide
Collecting data is the easy half. This is the other half: what to set up, what to teach your testers, what the report actually answers, and how many plays it takes before any of it means anything.
Set up
Playtesting is not a separate app or a website you log into to run a test. It lives inside the NexTurn app you already use to run a game night.
A project is one game under test. Not one session and not one version. The project is the thing that outlives both, and every play, survey and version comparison hangs off it.
THE TWO MINUTES THAT MATTER
WORTH DOING BEFORE THE FIRST SESSION
ONLY IF YOU NEED IT
Every project has a join code that looks like TEST-4KP2. Uppercase, four characters, from an alphabet with no letters you can mistake for digits. That code is the whole invitation.
THREE WAYS IN, IN ORDER OF FRICTION
Testers see only their own notes and plays, never each other’s. Reading someone else’s review first tends to colour your own, and at the sample sizes a playtest runs on, one contaminated opinion is a meaningful fraction of the evidence.
By default a tester can pass the code on. That is usually what you want: the realistic shape of a playtest is you sending five people to five tables, and each of them needs to pull their own group in. If you are running something confidential, switch it off in the project settings. The code still works, you just hand it out yourself.
Run the test
One phone in the middle runs the game, or everyone joins from their own. Either works, and neither needs our devices.
STARTING A PLAYTEST SESSION
THE ONE HABIT WORTH TEACHING
Tap the flag. One tap marks that worked, that didn’t, or rules issue, and it is stamped with the exact turn, round, phase and whose turn it was. Nobody has to stop playing or write anything. We ask for the “why” at the next round break, when the table has already paused.
AT GAME END
Scores go in as they always do. Then the app asks for the few things only the table knows: how it felt, which faction each player had if you listed any, and your after-each-play survey. A tester answering on their own phone answers privately; a single phone in the middle answers for the table, and the report keeps those apart because “12 responses” means something very different if it is three tables rather than twelve people.
Read the results
| Chart | The question it answers |
|---|---|
| Phase share of runtime | Which part of my game is long, not whether it's long |
| Turn pace by round | Does it bog down, and when |
| Convergent moments | Where several testers independently noticed the same thing |
| Experience split | Is it long, or long until you know it |
| Seat win rates | Whether turn order matters (usually not, at your sample size) |
| Version compare | Did the change do what you wanted |
If you listed variants at setup, the report adds a balance card. It answers two different questions and keeps them apart, because they need different responses from you.
Balance is the one you expected: how each faction did, with the number of plays behind it always printed beside the figure. Coverage is the one designers tell us they didn't know they needed: which of your options nobody has played yet. Four untested factions in a list of eleven is not a balance problem, it's next week's schedule, and no percentage above it can speak for them.
Co-operative games get a different reading, clearly labelled. In a co-op everyone at the table shares one result, so a per-faction win rate computed the normal way isn't a weak signal, it's meaningless. Every faction at a winning table would read 100%. Instead we report “tables that included this one won 3 of 5”, which is a real question, and we never mix it in with a competitive rate.
| Claim | Plays needed | Why |
|---|---|---|
| Median game length | 3 | Length is stable; you'll know it early |
| Phase share of runtime | 3 | Built on hundreds of turns per play |
| Turn-pace claims | 8 | Needs enough rounds across enough tables |
| Version comparison | 12 per version | You're comparing two noisy things |
| Seat / turn-order advantage | ~40 | A 4-player game is a coin flip with four sides |
If you take one number away from this page, take ~40. Seat advantage is the claim designers most want to make and least often have the plays to support. A 36% / 14% split across 14 games is exactly what a fair game looks like.
Label a ruleset before the session that uses it: v0.7, post-Essen, 3.2 cards. Free text, your scheme. Every play is stamped with the label that was current when it ran, so you never have to reconstruct which rules a game used.
Twelve plays per version is the bar. Below it we'll still show you both columns; we just won't call the difference real.
Good to know
We won't tell you something is true when we don't have the plays to back it up. Every number comes with how many games it's built on, and the thin ones say so instead of pretending.
We don't score your game, predict its rating, or tell you whether it's good. We tell you what happened at the tables, with the sample size attached. What it means is your job; you're the designer.
Reply to any NexTurn email, or use the contact box on our site. It comes straight to us and we read every message. If something in this guide is wrong or unclear, that's the most useful thing you can tell us.
Read the guide, then go and get three plays. Three is enough to know how long your game really is.
Start a project