What Did Your Unscripted Format Test Actually Prove?
What Did Your Unscripted Format Test Actually Prove?
Two sentences stand between a format test and a pitch that can survive scrutiny.
The first belongs to the pitch: players worked it out on their own, or the round lands every time, or people lean in. The second belongs to the record: what was on the table, who was sitting at it, what the host said after the rules were read, and how long the whole thing took.
Most pitches contain the first sentence and assume the second. The useful work is to write both, then shrink the first one until it fits inside the second.
Four different claims dressed as one word
When a creator says a test "worked," at least four separate claims may be hiding in that word, and they don't cost the same to support.
The weakest is comprehension: people understood what the round was asking them to do. Almost any recording can support this, including one where the players got everything wrong.
Next is unaided operation: the group produced the outcome without a hand on the wheel. This requires the record to show what the hand did or didn't do, which is exactly the part edited reels tend to omit.
Then repeatability: this happens across groups, not just this one. A single round cannot support this claim no matter how well it went. Two rounds can begin to, if the conditions were comparable.
And last, appeal: people want to watch it. A room of participants is not an audience. Six people concentrating at a table have told you about six people concentrating at a table.
Evidence accumulates from left to right. A recorded test round can support the first claim outright, sometimes the second, and points toward the third without settling it. The fourth claim needs entirely different work, and no amount of enthusiasm in the room substitutes for it.
The artifact is not the record
Sort your material by what each piece can actually bear.
A two-minute cut is evidence about what the format can look like, and about how someone chose to shape it. That's real and useful. It is evidence about frequency, difficulty, duration or unassisted success only if the underlying conditions are still recoverable — and the usual reason you have a two-minute cut instead of a forty-minute one is that the other thirty-eight minutes weren't the point.
Condition notes — the rule card as read, the deal, the timing, the restart log — are what let a second person reproduce the round. Without them, you have an anecdote with a nice ending.
A full recording holds the waiting, the false starts, the moment someone asked whether they could pass, and the answer. It is the least watchable artifact and the most informative one.
A participant account tells you what the round was like from inside. It is recollection, offered after the fact, by someone who now knows how it turned out. Worth having. Not interchangeable with the tape.
The polish on an artifact measures how much someone wanted it to be persuasive, not how completely it documents what happened. Keep the two questions separate: what does this show, and what was this made to do?
The condition ledger
Here is the teaching tool worth building before you spend anything: a blank ledger, one row per test round, every value entered from the record. Where there is no value, you write not recorded, which is a finding, not a gap.
- The rules version actually in play, and whether it matches the rule card
- The deal: fixed set, random draw, or stacked, and who decided
- What participants were told before the clock started
- What participants already knew: each other, the game type, the answer's subject matter, the designer's intent
- Every host utterance after the start, with timing
- Restarts, and what caused each one
- Time from start to the group's commitment
- The group's answer against the written answer
- What was cut in editing, by whom, and why
- What was never captured at all
That last line does more work than it looks like. Ledgers usually fail at the bottom, not the top.
The ledger's job is not to make the test look rigorous. It's to make the test repeatable — to let a stranger five months from now re-run the round and know which differences matter. A test whose conditions can't be reconstructed isn't a repeatability study. It's a memorable afternoon.
Six Clues
Six Clues is a proposed format, described here as a design on paper. It has not been run. Nothing below is a result.
The draft: six players around a table. Each is dealt one clue card, face-down, visible only to its holder. A single answer card sits face-down in the center. Each clue is individually consistent with several answers; the six together narrow to one. Players may describe their own clue in any words they like but may not show it. After the rules are read, the host says nothing. The group must speak one answer aloud within a five-minute window. The answer card is turned over only after the group commits.
The mechanism under test is narrow and worth naming precisely: can six strangers holding partial information converge on the intended answer without help? Not "is this a good show." Not "will audiences enjoy it." One mechanism, one condition.
Now the three branches, each of which you might find yourself holding.
If the record documents an unaided round — six clues described, no host speech between "go" and the group's answer, no restart, the spoken answer matching the card — then you can say exactly that, and no more. Six players holding these six clues reached the intended answer inside that window with no intervention. What you have established is that the clue logic closes and that the conversation can get there. What you have not established is episode shape. A five-minute convergence is a mechanic. A forty-five-minute episode needs escalation, stakes, a reason to care which answer wins, and a shape that makes minute thirty different from minute four. The mechanic is necessary and very far from sufficient.
If the group converged after the host said something — has everyone heard the third clue? — the result is not ruined. It is reclassified. What you have is mechanism evidence with a documented dependency: the round needed a clarification to close. Trace the chain honestly. The host said X, X supplied Y, and Y was something the format's design assumed the clues would supply. Then you have two real options, and both are better than editing it out. Change the design so the clues supply Y on their own — clearer wording, a printed protocol, a different distribution of information. Or accept the dependency, and design the host's role as a deliberate, specified part of the format rather than an improvisation you hope nobody notices. Assistance is not a verdict. It is a readout of where the format ends and the room begins.
If what you have is an edited reel and the conditions can't be recovered — no notes, no full tape, nobody who remembers whether the deal was restacked — then the honest claim is about the artifact. This is what the round can be made to look like. That's a legitimate thing to put in front of someone. It is not evidence about frequency, difficulty, unaided operation or duration, because the missing conditions are precisely the ones that would decide those questions. The temptation here is strong, because the reel is the only thing you have and it looks like proof. It isn't. It's a demonstration.
Mechanism evidence is a small, sharp object
A documented successful choice can establish that one mechanism was intelligible under those conditions. Hold that sentence and notice how little it contains. One mechanism. These conditions.
It does not establish that the mechanism recurs across groups who haven't played a deduction game before. It does not establish that the mechanic fills an episode, or that a season of it varies enough to stay interesting, or that the casting bench exists at the scale a series order needs. A mechanism that works beautifully on a table can still fail to become an episode. An episode that works can still fail to become a season. Those are three different jobs and your test addressed the first one.
This is why the phrase the format plays itself is worth striking from a pitch on sight. It compresses a mechanism observation into a whole-show claim, and the compression is invisible until someone asks to see the tape.
The next question has to separate explanations
Suppose a round stalled. The group never committed, or committed wrong, and you have the tape.
Two explanations predict that same tape, and they call for opposite fixes.
The first is an underdetermined clue set: the six clues written for the intended answer also admit a second answer, and a careful group found it. That's a design problem in the card text, and no amount of rehearsal fixes it.
The second is a coordination failure: the group held the information and had no protocol for exchanging it, so the five minutes went to turn-taking and one person's theory. That's a facilitation problem, and the clue set is fine.
These are distinguishable, and that's the whole point of designing a next test rather than running another round. Run two bounded comparisons, each changing one thing.
In the first, hold the clue set fixed and add a speaking order — each player gets the floor once, in sequence, before open discussion. If the stall disappears, coordination was the binding constraint. In the second, hold the protocol fixed and hand the same clue set to cold readers whose only instruction is to write down every answer consistent with all six clues. If several answers come back, the set was the problem, and the deal was never solvable in the way you thought.
A third explanation sits in the ledger rather than in the design: the participants may have known each other, or played deduction games regularly, or heard the pitch beforehand. That's not a flaw, it's a condition — and it matters most when you're tempted to generalize from it. Testing with a colder group changes the group, not the deal, so hold the deal and protocol constant and change only who's sitting there.
What doesn't help is another favorable round under identical conditions. A second success on the same terms narrows nothing. It restates the branch you already had.
And a real trial with real participants carries its own production and participant requirements, which this article doesn't certify and isn't qualified to. Treat that as a separate workstream with its own competent people, not a footnote to your pitch.
Rewording the pitch
Here is the shape of the transformation, with brackets where a real value belongs:
Before: The format is repeatable. Players work it out on their own.
After: In [one] recorded round under the [date] rules, [six] players reached the written answer in [duration] with no host speech after the start. Whether this holds for groups who haven't played the format before is untested. The next two steps hold the [same] deal fixed: a turn protocol, to test whether a coordination stall rather than the clue set was binding, and a cold-reader check of the clue set, to test whether it admits a second answer.
The second version is longer, narrower, and much harder to attack — because every clause is anchored to something a stranger could check. It also does something the first version can't: it names the uncertainty you intend to reduce, which tells a commissioner you know the difference between a demonstration and a study.
Keep the four categories separate in your own head, because they'll run together the moment you're excited:
- A paper walkthrough with colleagues tests whether the rules are legible and whether the clue logic closes. Your colleagues know what you meant. That is a real limitation, not a small one.
- An actual trial with participants who don't know you tests the mechanism under conditions you can document. It speaks to that group, that deal, that day.
- An edited demonstration shows what the round can look like at its best.
- Audience evidence is a separate body of work with separate methods.
Calling all four "the test" is how a narrow result becomes a broad claim without anyone deciding to make it one.
A claim no broader than its record
The result of a format test is not a verdict on the format. It is a narrowed question. You started with does this work? and you should end with something closer to does this work when the turn protocol changes, for a group that has never seen it? — which is answerable, and which your next round could actually settle.
Write the claim you can defend. Mark the part you can't, in the ledger, in the pitch, in the sentence you'll say out loud. Then choose the next test for what it could change your mind about, not for how it would look.
The strongest sentence you'll ever write about a format test is the one you could still defend after someone asks to see the unedited recording.
Frequently asked questions
What four different claims can 'the test worked' hide?
Comprehension, unaided operation, repeatability, and appeal. Evidence accumulates from left to right: a recorded test round can support comprehension outright, sometimes unaided operation, and point toward repeatability without settling it. Appeal needs entirely different work.
Why keep a condition ledger if the test already seemed to work?
The ledger makes the test repeatable, not just rigorous-looking. One row per round records the rules version, deal, what participants were told and knew, host utterances, restarts, time to commitment, answer, cuts, and what was never captured. Where no value exists, mark it 'not recorded'; that is a finding.
If the host said something during the round, is the result ruined?
No; it is reclassified as mechanism evidence with a documented dependency. Trace what the host supplied. Then either change the design so the clues supply it, or make the host's role a deliberate, specified part of the format. Assistance is a readout of where the format ends and the room begins.
What does a documented successful round actually establish?
That one mechanism was intelligible under those conditions. It does not establish that the mechanism recurs across groups, that the mechanic fills an episode, that a season varies enough, or that a casting bench exists at series scale. Mechanism, episode, and season are three different jobs.
How can the next test separate an underdetermined clue set from a coordination failure?
Run two bounded comparisons, each changing one thing. Hold the clue set fixed and add a speaking order; if the stall disappears, coordination was the binding constraint. Hold the protocol fixed and give the same clues to cold readers writing every consistent answer; if several answers come back, the clue set was the problem.