Skip to content

Design a Competition Pitch for Contestants With Unequal Skill Levels

Television

Design a Competition Pitch for Contestants With Unequal Skill Levels

A pitch that says "our cast spans the whole range of ability" has said something about the cast, not about the competition. The format question is still open: when this ends, what does winning show? Every later decision—round order, scoring, teams, roles—will answer that question whether you intend it to or not.

Here is how quickly it happens. Five rounds. Round one is a speed puzzle; rounds two through five are something else; the lowest scorer leaves after each round. Whatever you said about a varied cast, round one has done part of the selection before the variety arrives, and anyone whose strength lived in round four may already be gone. Nothing about that rule is unequal or unfair. It has simply answered the ability question in a way nobody wrote down.

That is the territory here: not whether the rules are equal, but what the equality is for.

Name the difference the contest wants to compare

"Unequal skill" is shorthand for at least four different comparisons. Raw proficiency at a defined task. Adaptability, meaning performance on something unfamiliar. Rate of learning. Coordination with other people. A format can be built around any of them. A format that intends one and scores another is broken in a way no scoring tweak repairs, which is why the comparison comes first and the challenge comes second.

Deciding who to cast and deciding what the contest compares are two jobs, and finishing the first does not finish the second. Variety in the cast is a programming need. It may also be a genuine dramatic engine. Neither of those is a statement about what a win will mean.

A quick test helps. Take a difference in your proposed cast and ask: if we removed it, what would the competition lose? If the honest answer is nothing, you were casting for texture, which is a perfectly good reason to cast and a bad reason to design. And one distinction to keep straight: "no demonstrated specialty" is a fact about an audition record. It is not a finding about a person. If you write the second when you mean the first, you will design a challenge around an assumption you never tested.

A proposed fictional format, used throughout. R is a stipulated spatial specialist with years of practice on rotation, assembly and constraint puzzles. V is a stipulated verbal specialist—crosswords, anagrams, cryptic clues. N is a stipulated newcomer with no demonstrated puzzle specialty and no competition record. These are capabilities as described, not assessments, and not predictions about how anybody behaves under studio lights. R, V and N do not change for the rest of this article. Swapping the cast so that each design looks successful is the most common way this analysis goes wrong.

Before going further, separate three claims that pitches routinely merge. Equal rules. Fair treatment of participants. Watchable television. They are different, they come from different kinds of work, and the fact that a design has one says nothing about the others. There is also a fourth area this discussion does not cover at all: putting people of very different ability into a public contest raises questions about selection, pressure and duty of care that a format exercise is not the place to resolve.

And do not open the pitch with a promise that everyone is equally likely to win. That is a different promise, it usually means the result is either a lottery or a showcase you cannot read, and it lets you skip the actual question. If the honest answer is that a deep specialist rarely wins because the crown is for breadth, price that in now rather than discovering it in the edit.

Compare common tasks with different strengths

One spatial challenge, everyone solves it. This is the most legible thing you can build. One task, one attempt each, one ranking. A win demonstrates proficiency at that task, on that day, under those conditions—and the viewer can infer all of that without a lecture on scoring.

What stays invisible is everything else: V's verbal capability, N's capability in any untested domain, adaptability, coordination. If R wins, you have learned that the challenge sat within reach of trained expertise. If someone else wins, you have learned only that the task was solvable without the declared specialty, or that V or N was better on the day. One result cannot separate those two, and whether the challenge measured what you told the commissioner it measured is a question for review rather than a finding. Both outcomes are useful. Only one of them is really about R.

Keep the virtue in view. Legibility is a product benefit, and a lot of formats trade it away for variety and receive ambiguity back.

A sequence of different puzzle types. Give R a spatial round and V a verbal round, and N has no scheduled domain at all. Now the pitch has to answer an aggregation question, because the rounds do not combine themselves.

If results accumulate into a total, a win demonstrates a weighted average, and you have thereby declared how much spatial ability and verbal ability are worth relative to each other. That declaration is the real content of the pitch. Write it down, defend it, and expect to be asked. If instead each round eliminates, then the order decides whether the variety is ever exercised. Round one can quietly become the format.

Notice what "varied rounds" is and is not. It is not automatically more inclusive; it is a breadth contest whose breadth you selected. Add axes and the winner becomes more of an average performer. Score by peak and you have a specialist's contest with extra ceremony. Neither is wrong. They are different shows, and the difference is a sentence in the rules, not a nuance in the edit.

N's position in a varied design is the exposed one. There is no domain to schedule N into, so the variety gives N nothing of their own. The remaining routes for N are the team round, the improvement route, or a new axis—which is a casting decision wearing a rule's clothing.

A team information-dependency round. Each member holds part of what the team needs; the team must produce one solution. Two consequences arrive immediately. The compared quantity stops being a person and becomes a group, so no individual win is demonstrated. And a group result can be carried.

The design assumption underneath the round is that pooling is necessary. Check it. If a member's contribution is a word that anyone could read aloud, that member is a channel rather than a solver, and the round is not testing what the pitch says it tests. This is where V's apparent strength can turn out to be less relevant than the team assumed—not because V lacks skill, but because the design asked for the wrong thing. A paper model can flag that risk. It cannot tell you whether real players route around V.

There is also fragility. If the outcome hinges on one member's piece, the group's result is that member's result with administrative overhead. What you gain is genuine: teams test communication, which nothing in the single-task or varied-round designs touched. What you pay is that the format no longer tests the abilities those designs tested.

Examine improvement and asymmetric roles separately

A declared improvement comparison. This answers a different question from every route above: not who is best, but who changed most. Four things have to be fixed before the format means anything—the task, the practice allowed between attempts, the measure, and the window. Use the same task at baseline and at re-measurement, or you are comparing two different achievements and calling the difference growth.

Then choose your gain, and notice that the choice decides who can win. Absolute gain favors whoever started with more room. Proportional gain favors whoever started low, because a small denominator is generous. If the measure has a ceiling, someone near it has little headroom by construction. Whichever you pick, you have chosen a rule that shapes the outcome, and the pitch should say so rather than let an audience find out by feel.

The harder unresolved question is what counts as change. A newcomer's early sessions tend to consist of learning the structure of a task. An experienced solver's late sessions tend to consist of refinement inside a structure they already have. Whether those are the same kind of event is a judgment the format has to make explicitly, because viewers will make it implicitly and will not all agree.

Two hazards worth naming to whoever reviews the design. A baseline taken once, on one day, in one room will move for reasons that have nothing to do with learning, and a single weak baseline produces a large measured gain. And repeating an identical task invites memorization, which is an incentive problem rather than an ability problem—a separate discipline, with its own failure modes.

Asymmetric roles. Two people, contrasting strengths, different access: one can see and not touch, the other can touch and not see. Roles can be assigned by the format or chosen by the players, and the two versions are not equivalent. Assignment removes strategy and can look arbitrary. Choice converts your ability comparison into a decision comparison, which is a different show.

The trouble is influence. Whichever role makes the decisions that determine the solution is the role that matters; the other may be executing. And the pair's result may be mostly about how the two of them communicate, which is fine if coordination is your intended comparison and a problem if it is not.

State the unresolved question instead of solving it in prose: role-level attribution. If a two-person team wins, what did the person who never touched the pieces win? A progression rule that advances both equally has answered that question without arguing for it.

Introduce a participant who breaks the assumption

Add G: stipulated strong in both the spatial and the verbal domains. Not predicted to win. Simply capable across the axes your format tests.

Run G through the routes.

Under the single spatial challenge, G does not break anything—G sharpens it. The comparison becomes specialization against general strength, which is exactly the comparison a one-task format performs well.

Under varied rounds, G is a problem. If the pitch's argument was "our variety exists because no single person can carry all of it," G falsifies the premise in the first episode. If G wins a varied-round format, the win demonstrates breadth, which is a fine thing for a win to demonstrate and not the thing you claimed. You cannot fix this by adding axes without checking which way each addition cuts: a new axis may just widen the breadth contest G is already winning, or it may hand someone else a domain and shrink G's share. Both are legitimate moves. Both are casting adjustments presented as rule adjustments.

Under the team round, a team containing G makes the dependency optional. The round then measures whether the team organizes around G—a leadership and coordination question—rather than whether pooled information was required. That may be more interesting than what you planned, but it is not what you pitched.

Under the improvement route, G has the least headroom under an absolute measure and the smallest proportional gain from a high starting point. An improvement ladder systematically disadvantages the strongest starter. That is not a design flaw; it is the price of the comparison. It does mean a pitch that runs an improvement ladder and still promises anyone can win has chosen an outcome shape while denying it.

Now ask, for each route, whose actions still matter. In the single challenge, N's actions matter as contrast. In varied rounds, N has no domain. In the team round, N's actions matter only if the team genuinely cannot pool without N's piece—which depends on how the pieces were written. In the improvement route, N's actions matter most, and the measure is most likely to have been designed with N's kind of starting point in mind.

Then find where advancement separates from the intended comparison. Elimination order in the varied design. Whole-team advancement in the team round. The baseline rule in the improvement design, where an early advantage can be banked if the measure is taken once. In each case the fix is to redesign the comparison, not to patch the scoreboard.

State the chosen trade-off and the required test

Pick one architecture and name its cost.

For a format that needs a varied cast, the defensible choice is a varied-round structure with declared weights, no elimination until every scheduled domain has been played and each contestant who has none has had the team round or the improvement measure, and one team round whose dependency you have checked on paper. The gate has to say what happens to N, whose domain the variety never scheduled; a condition N cannot satisfy is not a protection. What winning then demonstrates is the broadest useful command of the domains the show has selected. The cost is visible: the crown is for breadth, the specialist's excellence becomes the midgame rather than the ending, and the audience has to be told the weights instead of inferring them. That is a real cost. It is also a clear sentence you can say out loud to a commissioner.

The alternative is equally defensible and easier to explain: one task, one crown, casting for texture rather than for comparison. Narrow, legible, honest. What is not defensible is running the second and describing the first.

A paper model can do real work here. It can show you that a role is decorative, that two routes reward different things, that your stated comparison and your scored one have drifted apart, that a dual-strong entrant collapses a premise. It cannot certify fairness, suitability, safety or appeal. It cannot tell you what an audience will find gripping, and it cannot tell you what any stipulated participant does on the day. Nothing described above has been run—not the challenge, not the profiles, not the scoring.

So the test is specific. A playtest with people who did not write the rules. A review of the challenge and progression design by someone qualified to catch a dependency you believe is necessary and that is not. Particular scrutiny of the improvement measure and the asymmetric roles, because those are the two places where a format's stated claim and its mechanics most often part company. And before the pitch locks, a format creator's own account of how they designed challenges across a mixed-ability cast—not as a citation to wave, but because an argument and a practice are different objects.

Winning this format should demonstrate the broadest useful command of the domains you have chosen; the architecture that makes that comparison visible is a varied-round structure with declared weights, a non-elimination gate with a place in it for the contestants the variety never scheduled, and a team round whose dependency you have checked on paper; and the part still untested is whether real players under studio pressure produce that comparison or a legible version of something else.

Frequently asked questions

What different comparisons can “unequal skill” stand for in a competition format?

At least four: raw proficiency at a defined task, adaptability on something unfamiliar, rate of learning, and coordination with other people. A format can be built around any of them, but intending one and scoring another is a break that no scoring tweak repairs.

What does a single-task format make visible, and what does it leave invisible?

It is legible: one task, one attempt each, one ranking, and a win shows proficiency at that task on that day under those conditions. It leaves other capabilities invisible, including verbal or untested domains, adaptability, and coordination. One result also cannot separate trained expertise from a task solvable without the declared specialty.

In a varied-round format, what must the pitch decide about combining rounds?

Rounds do not combine themselves. If results accumulate into a total, the win demonstrates a weighted average, and the pitch has thereby declared the relative worth of the domains. If each round eliminates, round order decides whether the variety is ever exercised; round one can quietly become the format.

What assumption sits underneath a team information-dependency round, and how can it fail?

The round assumes pooling is necessary. If one member’s contribution is merely a word anyone could read aloud, that member is a channel rather than a solver, and the round is not testing what the pitch says. A paper model can flag that risk but cannot tell how real players would route around that member.

What testing does the article recommend before the pitch locks?

A playtest with people who did not write the rules; a review of challenge and progression design by someone qualified to catch a dependency believed necessary but not actually necessary; particular scrutiny of the improvement measure and asymmetric roles; and a format creator’s own account of designing challenges across a mixed-ability cast. A paper model cannot certify fairness, suitability, safety, or appeal.

More in Television Browse all articles