Skip to content

Show the Judgment Around an AI Output, Not Just the Best Result

Business

Show the Judgment Around an AI Output, Not Just the Best Result

A demo that lands and a demo that informs are two different achievements. The first needs one good output at the right moment. The second needs the audience to understand what went in, what could have gone differently, and who carries the decision the output feeds.

The gap tends to open about a minute after the applause, when someone asks how often the system gets it wrong. "We're still measuring" is an honest answer, but it lands badly, because ten minutes of polish implied that an answer existed.

There is a version of this that holds up. Show the workflow rather than the artifact: the input and what it leaves out, the output the system actually produced, and the review that turns a proposal into something usable or throws it away. Say what preparation shaped what the audience just watched. Then be explicit about what the demonstration establishes. One output shows behaviour on one input under stated conditions. It does not show typical performance, and it should not be described as if it did.

That is the whole argument. What follows is how to build it.

Start with the decision the output supports

Producing a suggestion and deciding that it is correct are separate acts, and demos blur them constantly. So before building anything, name the decision the output feeds and what happens when it is wrong. A draft summary that misstates a date wastes a reader's minute. A routing decision that misstates a date sends work to a person who never agreed to do it. The same underlying capability deserves a different demonstration depending on which of those is true.

You can usually tell a demo was built backwards from what is on the screen. An impressive paragraph, a clean image, a formatted table — and no way for the audience to know what anyone is supposed to do with it. When the purpose is unclear, people fill the gap with whichever inference suits them. Some assume the system handles everything the artifact implies. Others assume the whole thing is theatre. Both are guessing, and the presenter has no way to correct either guess.

The counterweight is worth stating too. If the output genuinely is low-stakes — a first draft that a named person edits, with no external consequence and no deadline attached — say that, and keep the demo short. Over-dramatising the review is its own distortion. A demo of a spell-checker that spends four minutes on fallback procedures has misrepresented the stakes just as surely as one that hides a correction.

An invented example: three lines, three verdicts

Everything in this section is made up. No model was run to produce it, and it says nothing about how any real system behaves. It exists to make the shape of a review concrete, and the numbers and names in it are stipulations, not observations.

Suppose the fictional feature is a meeting-summary assistant that turns a short note into an action list. The input is one note:

Thu sync — Priya, Marcus, Ana. Rate table: Lena's team owns the source data. Priya to reconcile by Friday. Release note: Ana drafted it, waiting on legal. Marcus to send redlines to Ana. Next sync Tuesday.

And the invented output is this:

  1. Lena — reconcile the rate table by Friday.
  2. Marcus — send redlines to Ana.
  3. Ana — follow up with legal before Tuesday.

Three lines, three different verdicts, and only one of them is a straightforward pass.

Line 1 misassigns the owner. The note attaches the reconciliation to Priya; Lena appears in the previous clause, attached to a different fact, which is that her team owns the source data. Whether Lena owes anyone anything is a question for Priya, not something the note settles. The correction is small — replace the name — but the consequence is not. Sent as written, the request reaches Lena, who may not own the task, while Priya may reasonably assume it is covered. The human action here is not editing a word. It is deciding who receives the request.

Line 2 matches the note, so it passes. Passing is narrower than it sounds. It means the content agrees with the source. It does not mean the line is complete: the note gives no date for the redlines, so the item is undated, and someone still has to decide whether that matters. A gap in the source is not the same as an error in the output, and the review should treat them differently. If nobody owns the gap, the demo should show the item going out undated, or being held. Quietly supplying a plausible date is the failure mode we are trying to avoid, and it is exactly what the next section is about.

Line 3 has no basis in the note. The note says the release note is waiting on legal. It does not say Ana is chasing legal, and its grammar leaves open who is waiting in the first place. The assistant resolved that ambiguity by naming a person and turning a dependency into a task. You cannot verify this line against the source, because the source does not contain it. The honest disposition is to strike it or to hold it pending a message to Ana. Keeping it because it sounds like something that ought to be true is how invented commitments enter a real workstream.

Two things about that review are easy to skip in a demo and shouldn't be. First, it is not always reading. Resolving line 3 involves contacting a person and waiting, which takes time and interrupts someone else's afternoon; that cost is part of the workflow. Second, the person doing the review owns the sending decision. The model proposed a list. A human is accountable for the list that goes out, and the demo should make that transfer of responsibility visible rather than implying the artifact arrived finished.

The preparation is part of the claim

Now suppose the fictional list above was the third attempt at the same note. Attempt one produced actions with no owners at all; the instruction was revised to keep the name attached to each action. Attempt two kept the owners and invented a deadline for the redlines; the instruction was revised again, to use only dates present in the note. Attempt three is what we have, and the revisions did not fix everything. The invented redline deadline is gone, but an owner is still wrong and a line is still invented. Even the "only use dates in the note" rule holds in letter and slips in inference: line 3 borrows the word Tuesday from the sync line and adds a constraint, before, that the note never states.

That is the honest shape of instruction-following: partial, per-line, and checkable only against the source. It is not a property of the system you can assert once and be done with.

So label it. "Attempt three of three; instructions revised after attempt two" costs one sentence and prevents the audience from believing they watched a first try. The same applies to selection: if you generated four summaries and showed the best one, the number four belongs in the demo. And it applies to edits. If the demo ends with a polished email rather than the model's raw list, the audience has watched a person work and credited it to the system. The document you send is a human document that a model helped draft, and the difference is worth a slide.

In a real demonstration these are logged events, not stipulations. Keep the prompts, the number of runs, what was edited, and which output was chosen. That log is also what lets you answer the reliability question without guessing — not with a rate, which you may not have, but with the composition of what the audience saw.

Prepared example or live run

A prepared example can walk an audience through a complicated workflow: several handoffs, an ambiguous source, a correction that changes who gets asked to do what. You can narrate it precisely, pause where the judgment happens, and keep the room oriented. Its limits are real, though. It shows nothing about variation, and it is fully compatible with a system that only works when you have prepared for it.

A live run does something the prepared example cannot: it operates on input the audience can inspect, and its failures are more credible because nobody arranged them. It also buys you less than it appears to. One run on one note is one case, and audiences generalise from a single live result faster than they do from a rehearsed one, precisely because it felt unscripted. If you have changed anything since the prepared example — the prompt, the model version, a setting — the live run is a different system from the one you just described, and the demo should say so out loud.

The combination works better than either alone. Run the prepared example to make the workflow legible. Then run live on a note you are authorised to show, and treat that result as a second case rather than a reliability test. Together they establish that the workflow is intelligible and that these outputs followed from these inputs. They do not establish an error rate, and if someone asks for one, the useful answer is what evidence would be needed to produce it.

One practical constraint on the live half: a real meeting note carries other people's names and sometimes content they did not agree to broadcast. Either use a note whose participants have cleared it, or write one for the demo — and if you write it, say you wrote it, because a tidy synthetic note is easier than the real thing and that difference changes what the audience should conclude.

Review labels are not confidence scores

Google's PAIR guidance on explainability and trust is the closest thing to a source here, and it is worth being precise about what it is: design guidance about helping users calibrate their trust, communicating limits, accounting for situational stakes, and making human control and recovery visible. It also raises the question of whether displaying confidence helps anyone make a decision. It is not an evaluation of any product, and it does not tell you that your workflow is reliable.

The practical consequence is simple enough. Don't put a percentage on a line. A confidence display is not automatically a trustworthy probability of correctness, and unless someone measured how often a claim at that level turns out right, the number is a guess wearing the clothes of a measurement. Worse, it is a guess that changes behaviour: reviewers defer to high numbers, which means an uncalibrated score makes the review slightly worse than no score at all. If the number would not change what anyone does, it is decoration. If it would, it needs calibration you probably do not have.

What you can show is a record of the review, and it is a different kind of object. "Checked against the note." "No basis in the source — held." "Date missing from input." These describe what a person did, and they are honest because they are about the process rather than about the model's inner state. They are also more useful to an audience than a percentage, because they show where the work happened and who did it. The distinction is worth keeping clean: the review status is knowable, and the model's certainty is not.

Plan the unusable output in advance

If the live run produces something you cannot use — a mangled list, a hallucinated attendee, an empty field — decide beforehand what happens next, because the room will be watching you rather than the screen.

Three options, roughly in order of how much they ask of you. Continue and show the correction, which works best when the review is the point of the demo and you can narrate it calmly. Stop the demo and describe the fallback, which is honest as long as you say whether that fallback has ever been exercised; describing untested fallback behaviour as though it works is its own small exaggeration. Or switch to the prepared example and label the switch: "this is the prepared one, because the live run just gave us something we can't use." The label is the whole trick. Unlabelled, a switch looks like a save.

The failure path is often the most convincing part of a demonstration, because it is the part nobody believes you rehearsed. Which is why staged failure needs the same care as staged success. If you scripted a wrong output to illustrate the review, say it was scripted. A manufactured error presented as spontaneous is the same lie as a third attempt presented as a first.

What the audience should leave with

The thing worth handing over is not the output. It is a description they can repeat: here is what was automated, here is what varies, here is what a person decided and on what basis, and here is what we would need to measure before anyone quotes a number.

Build the demo that way and the question after the applause stops being a problem, because the answer is already on the screen. Where it isn't, you will at least know what would put it there — and saying so costs you far less than the alternative, which is standing in front of a good result and implying it means more than it does. Not having a reliability figure is a legitimate state. Pretending to have one is not, and the difference is entirely in what you choose to show.

Frequently asked questions

What should an AI output demonstration show besides the best result?

Show the workflow: the input and what it leaves out, the output the system produced, and the review that turns a proposal into something usable or discards it. Also say what preparation shaped what the audience saw and what the demonstration establishes.

Why can one good output not be treated as reliability evidence?

One output shows behavior on one input under stated conditions. It does not show typical performance or an error rate. A live run is still one case, and audiences generalize from it quickly. If asked for a rate, say what evidence would be needed to produce it.

Should review labels include confidence percentages?

No, not unless the scores have been measured and calibrated. A confidence display is not automatically a trustworthy probability of correctness, and uncalibrated scores can make reviewers defer to high numbers. Show a record of human review instead: checked against the source, no basis in the source and held, or date missing from input.

How do prepared examples and live runs differ?

Prepared examples can make a complicated workflow legible and allow precise narration, but they show nothing about variation. Live runs operate on inspectable input and failures are more credible, but one run is one case. Combining them works better, with the live result treated as a second case, not a reliability test.

What should you do if a live run produces unusable output?

Decide beforehand among continuing and showing the correction, stopping to describe the fallback while saying whether it has been exercised, or switching to the prepared example and labeling the switch. If a wrong output was scripted, say so; manufactured error presented as spontaneous is the same lie as a third attempt presented as a first.

More in Business Browse all articles