Skip to content

Judges Use Different Score Ranges: What Would Changing the Scoring System Change?

Television

Judges Use Different Score Ranges: What Would Changing the Scoring System Change?

Two judges, one score sheet out of ten and one out of a hundred. The pitch sounds simple: add the marks, highest total wins. Everyone in the room agrees that judging is the hard part and the arithmetic is the boring part.

The arithmetic is not boring, and it is the part you actually control. On the same three performances, scored identically by the same two judges, the result changes depending on which rule you use to combine the columns. Nobody changes their mind about who was better; the rule decides what their opinions add up to. That is not an argument for one method over another. It is an argument for choosing the method before you can see which one produces the winner you wanted.

Two problems wearing the same sentence

"Our judges use different ranges" describes two unrelated situations, and they need different remedies.

The first is a unit mismatch. One judge is handed a 0–10 sheet, another a 0–100 sheet. A point does not mean the same thing in the two columns, and no amount of good faith on the panel reconciles that. A judge has done nothing wrong to produce this problem; you created it when you printed the sheets.

The second is a spread mismatch. Both judges receive the same 0–10 sheet, and one uses 6 through 9 while the other uses 2 through 8. The units are identical. What differs is how much of the scale each judge's opinions travel across.

These look alike in a results table and are not alike at all. And one consequence is easy to miss: if both judges declare 0–10, then a rule that maps each score onto its share of the declared range does nothing. It multiplies every score by the same constant and changes no ordering. The remedy that sounds most official does not touch the most common version of the complaint.

A related caution: an unused stretch of scale is not evidence of incompetence. A judge holding back the top of a 0–100 sheet may be reserving it for a performance nobody has delivered yet. Declared bounds say what a judge is permitted to express. The observed range says what this panel happened to do in this round, which is a much thinner fact.

A fixture, and the baseline that looks neutral

The performances and scores below are invented for the exercise. No judges, contestants or panels were observed, and nothing here predicts how a real one would score.

Judge One works from a declared 0–10 range. Judge Two works from a declared 0–100 range.

Performance Judge One (0–10) Judge Two (0–100) Raw total
A 9 60 69
B 5 90 95
C 3 50 53

Plain addition puts B first. The scores are used exactly as given, which is why this rule feels like the honest one — no transformation, no adjustment, no editorialising. It is also making a decision.

Judge Two's scale is ten times longer than Judge One's. In an unweighted sum, that means every point Judge Two awards carries ten times the leverage over the differences between performances. A performance that Judge One rates a tenth of a range above its rival moves the total by one point. A performance Judge Two rates a tenth of a range above its rival moves the total by ten. You have two judges and one of them has ten times the say, without anyone declaring that this is what they meant.

One more detail: the top possible total under this rule is 110, a number with no meaning attached to it. That is the first sign that the columns are not the same kind of thing.

Rescaling to the declared bounds

The standard move is to bring both judges onto a common scale using the bounds they were given. For a 0–10 target:

10 × (score − lower bound) / (upper bound − lower bound)

Judge One is already there. Judge Two's 60 becomes 6, its 90 becomes 9, its 50 becomes 5.

Performance Judge One Judge Two (rescaled) Total
A 9 6 15
B 5 9 14
C 3 5 8

The winner changes to A. It does not matter which direction you rescale — stretching Judge One up to 0–100 gives 150, 140 and 80 — because the totals are just the 0–20 figures multiplied by ten. What matters is the ratio between the judges' influence, and this rule makes it one to one.

It is easier to see the reasoning in fractions of a range than in rescaled points. Judge One rated A at 0.90 of its range and Judge Two at 0.60 of its. Averaged, A sits at 0.75. B, at 0.50 and 0.90, sits at 0.70. C, at 0.30 and 0.50, sits at 0.40. Under this rule B's 30-point win on the hundred scale is a smaller separation than A's 4-point win on the ten scale, and A edges ahead.

Three results, three different stories about the margin. Under raw addition B leads by 26 of a possible 110. Under declared-bound rescaling A leads by 1 of a possible 20. The two rules do not merely disagree about the winner; they disagree about whether there is anything to disagree about.

What the declared-bound rule assumes

Rescaling to declared bounds treats a fixed fraction of each judge's declared range as the common unit. It keeps the relative distances inside each judge's column intact: if Judge One puts four points between A and B and two points between B and C, those stay in a 2:1 relationship. What it assumes is that equal fractions of a declared range mean equal strength of preference across judges — that Judge One moving a contestant from 0.5 to 0.6 is expressing the same separation as Judge Two making that move.

It also quietly hands each judge an influence proportional to how much of their declared range they actually use. A judge who declares 0–10 and stays between 8 and 10 is contributing a fifth of a scale while a colleague contributes the whole thing. The declared ranges have been equalised; the working ranges have not.

That is a real limitation and not automatically a flaw. If you want each judge's maximum possible contribution to be equal, this is the rule. If you want each judge's actual influence to be equal, you need to look at how they behave, which is a different project.

The same fixture, stretched to whoever showed up

There is a more aggressive version of rescaling that stretches each judge's observed minimum and maximum, rather than the declared bounds, to fill the scale. In this fixture, Judge One's lowest score was 3 and highest 9; Judge Two's were 50 and 90. Both columns get stretched.

Performance Judge One stretched Judge Two stretched Total
A 10.00 2.50 12.50
B 3.33 10.00 13.33
C 0.00 0.00 0.00

B wins. But this method has a property the others do not: its numbers move when the cast changes. Add a fourth contestant, D, who scores 1 from Judge One and 10 from Judge Two — a weak round, last place on both sheets. Under the raw, declared-bound and rank rules below, the totals for A, B and C do not move at all. Under this one, A rises to 16.25 and B falls to 15.00, and A wins.

Nothing about A, B or C changed. Their scores are the same. Their judges awarded the same marks. What changed is that a fourth person competed badly, which widened both judges' observed ranges and reshuffled the distances of everyone above them. The bottom of the field is now defined as zero by construction, so the same performance can be worth 2.50 in one episode and 6.25 in another.

The European Commission's composite-indicator guidance, in the normalisation step of its ten-step guide (page dated 1 December 2020), separates bringing values onto a common scale from ranking, which discards absolute distances, and it notes that the outcome is sensitive to the minima and maxima you choose. That is methodology for building composite indicators, not evidence about television judging — but the sensitivity it flags is exactly what the fourth contestant exposes here. If your transformation depends on the contestant set, your scorecard is reporting something about the episode, not the performance.

Ranks: the robust route that throws the most away

Rank aggregation ignores scores entirely. Each judge orders the performances; the placements are summed; lower is better.

Performance Judge One rank Judge Two rank Rank sum
A 1 2 3
B 2 1 3
C 3 3 6

A and B tie. C is last on every route.

Ranks have one enormous practical advantage: they do not care what scale a judge uses, or whether a judge switches scales mid-season, or whether one judge's sheet is out of ten and another's out of a hundred. Any transformation that preserves order leaves the ranks untouched. All the trouble in the first sections of this article simply evaporates.

What evaporates with it is the size of the gap. Under rank sums, A's 4-point win with Judge One and B's 30-point win with Judge Two are the same object: one place. If Judge One was emphatic and Judge Two was lukewarm, the rank table cannot tell you, and neither can anyone reading the leaderboard. For a format where the drama lives in the margin — the runaway favourite, the near miss — that is the whole story being deleted to prevent a unit error.

Ties have to be declared, and the convention decides results

The combined tie between A and B is not a footnote. It is the result, and it has to be handled by a rule that existed before anyone saw it. If you invent a tie-break now — highest single score, Judge One wins, audience vote — you are selecting a winner after the fact, and that is a different kind of format than the one you advertised.

Within-judge ties matter just as much, and they are easier to overlook. Take a variant of the fixture: suppose Judge One had given B and C both a 5. Judge Two's ordering is unchanged.

  • If tied entries share the better rank, Judge One reads A 1, B 2, C 2. Sums: A 3, B 3, C 5. A and B still tie.
  • If tied entries receive the average of the ranks they span, Judge One reads A 1, B 2.5, C 2.5. Sums: A 3, B 3.5, C 5.5. A wins outright.
  • If tied entries share the worse rank, Judge One reads A 1, B 3, C 3. Sums: A 3, B 4, C 6. A wins outright by a wider margin.

Same two judges, same opinions, a tie or an outright win depending on a clerical convention that no viewer will ever see explained. Pick the convention, publish it, and check what it does to a tied input before you need it.

What no rule fixes

None of these routes reconciles different criteria. If Judge One is scoring technical execution and Judge Two is scoring charisma, adding their columns asserts that the two are commensurable, and rescaling or ranking the columns does not make that assertion true. It only makes it tidier. You can decide to treat craft and presence as equally important, or as weighted three to one, but that is a decision about what the show is, and the arithmetic will faithfully carry whichever answer you supply.

This is the honest version of the promise a normalised scorecard appears to make. Bringing judges onto a common scale makes the calculation legible. It does not make the taste objective, it does not make the panel accurate, and it does not make the competition fair. Those would be separate claims requiring evidence that a formula cannot produce.

Choosing before play

The rule you pick is a statement about weighting and about what you are willing to lose.

Raw addition gives the box the weight the scale designer handed it, which is a weight you probably did not intend. Declared-bound rescaling gives each judge equal maximum influence and preserves the size of their gaps, at the cost of assuming declared ranges are equally meaningful and at the cost of equalising ranges nobody actually used. Rank aggregation gives each judge equal influence and total immunity to scale problems, at the cost of every gap's size. Stretching to the observed range gives you numbers that move with the cast, which is not a weighting at all — it is an unowned editorial decision that makes itself each episode.

Notice that the fourth route's result in this fixture, B, is the same as raw addition's. It should still be rejected, and rejected on grounds that have nothing to do with which name it produced. That is the discipline the whole comparison is asking for: judge the method by what it preserves and what it discards, not by whether you like its output on one set of numbers.

The same scorecard under four rules

Rule A B C Result
R1 Raw total 69 95 53 B
R2 Normalised to declared bounds (each judge to a common 0–10, totals out of 20) 15 14 8 A
R3 Rank sum, lower better 3 3 6 A and B tie
R4 Stretched to observed range (figures move if the cast changes) 12.50 13.33 0.00 B

Three of these give a stable answer on these performances when the cast changes. One does not.

Three things remain untested in this fixture and worth running before the format locks: what happens when three or more performances tie, whether rank sums produce different ties when you add a third or fourth judge, and how each rule behaves in a round where every judge's scores cluster inside a narrow band. Each is a small calculation and each has the power to surprise you in front of an audience.

Rule 2. Every judge scores within a declared range published in the judging brief. Each score is converted to the fraction of that judge's declared range it represents, and the fractions are summed, so the maximum total equals the number of judges and every judge's maximum possible contribution is identical. Ties: performances finishing within 0.01 of each other share the placement, and the format carries no further tie-break. This rule was chosen because it expresses equal maximum weight for each judge and preserves the size of the gaps each judge draws — not because it identifies the best performance, and not as a claim that any judge is accurate.

Frequently asked questions

What two different problems can judges using different ranges describe?

A unit mismatch—different sheets, such as 0–10 versus 0–100—and a spread mismatch, where both judges use the same sheet but one uses 6–9 and another uses 2–8. Unit mismatch is created by sheet design; spread mismatch concerns how much of a scale each judge actually uses.

Why does raw addition give one judge more say?

A longer scale gives each point more leverage. On 0–10 versus 0–100, a tenth of a range can move the total by 1 versus 10. The top possible total in the fixture is 110, a number with no meaning attached, and the winner changes under different rules on the same scores.

What does rescaling to declared bounds do, and what does it assume?

It converts each score to the fraction of that judge's declared range it represents and sums the fractions, giving each judge equal maximum possible contribution while preserving the relative distances inside each column. It assumes equal fractions of a declared range mean equal strength of preference; it equalizes declared maximums, not the working ranges judges actually use.

Why reject stretching to the observed minimum and maximum?

The numbers move when the cast changes. Adding a weak fourth contestant can raise A and lower B even though A, B, and C have the same scores from the same judges. That transformation reports something about the episode set, not just the performance.

What do rank sums preserve and discard, and how do tie conventions matter?

Ranks are immune to scale and unit differences and preserve order, but discard the size of gaps, so an emphatic win and a lukewarm win become one place. Within-judge ties handled by sharing the better rank, averaging ranks, or sharing the worse rank can produce a tie or an outright win, so the convention must be declared before it is needed.

More in Television Browse all articles