Skip to content

Explain Expert Judgment in a TV Competition Without Pretending It Is Measurement

Television

Explain Expert Judgment in a TV Competition Without Pretending It Is Measurement

A pitch deck says the work will be "assessed by a panel of experts." That isn't a format yet. It's the promise of one. What the panel looks at, which qualities it's been asked to weigh, and what happens when two strong entries call for different things — that's the format, and if a reader can't find those answers in your pitch, the phrase does no work.

There are two easy ways to avoid the problem, and both leave the reader knowing less than they did before. The first dresses judgment up as measurement: three judges, a score out of ten, a decimal place, a table. The number looks like a finding, so nobody asks what it found. The second gives up: "It's subjective, it's a matter of taste, the panel decides." True enough, and useless to a viewer trying to follow along or a producer trying to picture the episode.

The workable middle is to name what can be observed, name the criteria, and then show the reasoning that connects them. A score can record that reasoning, tally it, put it on a scoreboard. It can't stand in for it.

What a judge can actually see

Test the difference with one question: could two people with working eyes, shown the same thing, disagree about it?

"Entry A placed nine objects." No disagreement. Nine is nine. That's an observation — a count, a duration, a placement, a recorded action. "Entry A's arrangement is sparse" is close to the same thing and still mostly a count; sparse only starts to mean a shortfall once you've added a view about how much a room ought to hold. "Entry A's room feels unresolved" isn't an observation at all. It's a conclusion, arrived at through a standard the viewer can't see.

The slide happens fast and quietly. A pitch asserts something countable — the chair is turned out from the table; the panels are finished flush; the model took eleven hours — and then banks the credit for a quality verdict that the count doesn't support on its own. "The chair is turned out" is observable. "The chair reads as casual" is a first interpretation. "The arrangement is careless" is a verdict wearing an observation's clothes. Each step needs the step before it, and if you skip one, a viewer who disagrees with you has nothing to grab hold of.

Listing observable evidence is also the easy part of writing the pitch. You can do it almost mechanically, and it pays twice: it tells the panel what to look at, and it tells the audience what they're about to see on screen.

Criteria, priorities, and the discretion that remains

A round should say what it's asking for. Two or three criteria are usually enough: the qualities under assessment, why they matter to this brief, and what each one means here rather than in general. "Craft" is a heading. "Construction, finish and consistency of scale" is something a judge can be held to.

Criteria alone rarely decide anything, because strong entries are strong in different ways. So the round also needs a shared view of which criteria carry the most weight when they pull apart. There are two honest ways to handle that. Either the format states the priority in advance — this round rewards X first — or the format states that the panel resolves competing strengths in discussion, and treats that discussion as part of the method rather than a gap in the design. Both are legitimate. The mistake is leaving the priority unstated and letting the audience assume the criteria sorted it out by themselves.

Then there's the discretion that stays. Even with criteria and a stated priority, someone still has to decide what a criterion asks for when it's contested. That someone should be named in the pitch. Deciding whether readability means legibility at a glance or depth on a second look is an interpretive act. So is deciding whether a deviation is a lapse in craft or a deliberate choice the criterion was written to reward. More decimal places do not remove that choice. They only make it harder to see.

Which is worth being precise about, because scoring isn't the villain here. A number can encode a judgment perfectly well; it can be tallied, ranked, put on a leaderboard, and read out on air. What it cannot do is disclose the weighting that produced it. If a format wants to weight three criteria 40/35/25, that's a rule, and stating it as a rule is fine. Trouble starts when the weighting is hidden and the total is presented as a fact about the work.

A worked comparison

The example below is invented to show the mechanics. It isn't drawn from any broadcast, and no real format's published judging rationale was examined while writing it.

A hypothetical making competition runs a round with the brief: build a miniature room that feels like somewhere a person actually lives. Makers are shown one fixed camera position — the hero angle — before they start, and told the finished room will be judged from it. The round declares three criteria: lived-in quality (does this read as a place in use, rather than a set arranged to be looked at), craft (construction, finish, scale consistency), and readability of the space (from the hero angle, how quickly a viewer can identify the room's purpose, its zones, and the kind of person who uses it). The declared priority for this round is readability. That's the format's rule for this round, decided in advance — not a general claim about what makes a miniature good.

Two entries.

Entry A. Nine placed objects. The floor plan separates into three zones — a work corner, a sleeping alcove, a passage between them — and all three resolve from the hero angle without occluding one another. Every object sits inside its zone; nothing overlaps. Surfaces are finished uniformly, paint continuous across joins. No visible wear, spills or displaced objects. Joins clean, scale consistent.

Entry B. Thirty-eight placed objects. Zones overlap: a stack of boxes intrudes into the passage, and the front row of a shelf partly hides the row behind it. Trace evidence throughout — a chair turned out from the table rather than squared to it, a ring left on a surface, a garment over a chair back, objects layered in a way that suggests things were set down and not put back. From the hero angle the work corner and the sleeping alcove partially block each other. Joins clean, scale consistent.

Both entries are complete. Nobody ran out of time. The count of objects is observable in both cases, and it settles neither criterion on its own.

A reading that favours A. Under readability, A resolves in one look. A viewer can name the room's purpose and how it's used before the shot ends. Under lived-in quality, A is the weaker entry: nine objects can't carry the density of use that makes a place feel inhabited, and the uniform finish with no trace of wear reads as a set dressed to be photographed. Craft is strong on both. Since this round puts readability first, A takes it.

A reading that favours B. The same evidence, and mostly the same judgments. Under lived-in quality, B is the strongest entry in the round — the ring, the turned chair, the objects set down and not tidied are exactly the traces of use, and a room arranged this cleanly would look like nobody had ever sat in it. Under readability, B is weaker on a single glance but not, in this reading, weaker at all: an arrangement with overlaps and buried objects gives a viewer something to look into, and the criterion never said the room had to be understood in one second.

Both readings accept the same observations. They differ on what readability asks for and on how much weight it should carry. Notice what neither one did: neither misread the evidence, and neither is describing a procedural failure. A disagreement about what a criterion means is not the same thing as a rule violation, and a format that can't tell the two apart will end up defending itself instead of explaining itself.

The panel. Say the format's rule is that a three-judge panel talks the round through and the lead judge — one of those three — decides, with any dissent recorded in the round summary. Two of the judges, the lead among them, read the evidence as A better meeting readability as written — including its emphasis on speed — and B better meeting lived-in quality, with craft level. The third judge argues the priority was misapplied and would choose B on that basis; the dissent is recorded. The lead judge applies the declared priority and selects A.

Nothing in that required anyone to be wrong, and the outcome wasn't inevitable. Had the round declared lived-in quality first, the same three judges, looking at the same two rooms and reading the same evidence, would plausibly have landed on B. The criteria and the priority did real work. That's the point of stating them.

The 8.7 problem

Here's the sentence to watch for: "Entry A scored 8.7, so it was objectively better."

Almost everything is missing. Which criteria produced the 8.7, and at what weighting? Did the judge who scored B's 8.5 apply the same reading of the evidence? If the round's priority had been lived-in quality, would the same judges have reached the same number? If a viewer can't answer those, the score is a summary of a decision they can't audit — and the reported precision exceeds the precision of the judgment behind it. A gap of 0.2 doesn't mean A was better than B by a fifth of a point of quality. It means two compressed judgments landed near each other.

This isn't an argument against numbers. A score can be the most efficient way to tally a panel and rank a field, and audiences read leaderboards without difficulty. It becomes a problem when the number starts doing the explaining, because the number can't say why the readability criterion outranked the density of B's detail. Only the reasons can.

Writing the decision into the pitch

Four things need to be somewhere in your format description, and they're easy to confuse with one another.

What the judges see. The evidence the round puts in front of them — the fixed camera angle, the model itself, the timed element, whatever it is — and how much of that the audience will see too.

What they're asked to weigh. The criteria, their meaning for this brief, and the round's stated priority if it has one.

Who decides, and how. The named method: discussion then a lead judge, majority vote, consensus, a single judge with an advisor. Whichever it is, say it plainly, and say what happens to a disagreement — recorded, aired, or simply not part of the result.

Why this one. The reasoning offered, tied back to evidence and criterion.

That last item is the one that usually goes missing, and it's the one the audience is actually waiting for. Give the Performance evidence, the criteria and the reasons enough room around the reveal that a viewer can follow the decision and still prefer the other entry. That's not a flaw in the format. It's the format working: you've made the decision arguable, which means you've made it legible.

Keep it separate from the procedure, though. The procedure tells you who decided and by what method; the reasons tell you why this entry, on this criterion, this round. A pitch that offers the procedure and calls it the reasoning has explained the machinery and not the choice.

One boundary worth stating once and not labouring: making reasoning visible isn't a fairness audit. It doesn't establish that the process was lawful, consistent across rounds, or resistant to pressure. It establishes that a viewer can see what the decision turned on, which is a different and more modest achievement — and usually the one a format actually needs. If your pitch cites a real show's judging explanation as precedent, the check is whether the published criteria, the remarks that actually aired and the official account of the decision line up with one another. If the deliberation wasn't recorded, a result can't be reverse-engineered back into reasons. That verification is a separate piece of work from writing the format, and it's worth doing before you lean on the show as evidence.

A judging passage that would work

On this round's stated priority, Entry A takes it. From the hero angle, all three zones resolve without occlusion, so a viewer can read the room's purpose and how it's used inside a single look — which is what this round put first. Entry B carries the strongest evidence of habitation in the round: the turned chair, the ring, the objects set down and left. Under lived-in quality, B is the better entry, and craft is level between them. A judge who reads readability as a room's depth rather than its speed would reasonably choose B, and that reading is recorded. The decision rests on the priority this round declared, not on A being the better room in every respect.

That's a result a viewer can disagree with and still understand — which is the most a judged format can honestly promise.

Frequently asked questions

What should a pitch actually specify if it says work will be assessed by a panel of experts?

It should name what the panel looks at, which qualities it has been asked to weigh, and what happens when two strong entries call for different things. The phrase by itself is only a promise of a format, not a format.

How can a pitch separate an observation from a quality verdict?

An observation is something two people with working eyes, shown the same thing, would not disagree about—a count, duration, placement or recorded action. A conclusion such as 'the room feels unresolved' depends on a standard the viewer cannot see. A count does not support a quality verdict on its own, so each step needs the step before it.

How can a format handle criteria that pull against each other?

It can either state the priority in advance—this round rewards X first—or state that the panel resolves competing strengths in discussion and treat that discussion as part of the method. Leaving the priority unstated and letting the audience assume the criteria sorted it out by themselves is the mistake.

Does scoring necessarily undermine expert judgment?

No. A score can encode a judgment, be tallied, ranked and read out. What it cannot do is disclose the weighting that produced it. Trouble starts when the weighting is hidden and the total is presented as a fact about the work; reported precision can exceed the precision of the judgment behind it.

What four things need to be in a format description?

What the judges see, what they are asked to weigh, who decides and how, and why this one—the reasoning tied back to evidence and criterion. The last item is the one that usually goes missing and the one the audience is waiting for.

More in Television Browse all articles