Agree on a Customer Benchmark Before Showing the Winning Result
Agree on a Customer Benchmark Before Showing the Winning Result
The comparison usually arrives pre-decorated. Someone on the sales side has already run a product against something, liked what came back, and now wants two columns on a slide. The number itself may be perfectly real. The problem is that no benchmark exists yet — nobody has agreed what work the products are doing, what counts as finished output, which conditions each side ran under, or what any of it would be allowed to mean. When the metric is chosen before the definition, the rules of fairness get written after the outcome is known. That is the one sequence a customer has good reason to distrust.
So decide the comparison before you can see it. For most customer-facing evaluations, one of three shapes is the honest one: a common protocol that runs both products through the same agreed procedure, a configured demonstration that shows what one route can do under conditions you disclose in full, or a narrower comparison that isolates one answerable piece of a task that can't fairly be run end to end. Pick the shape that matches the question, and say out loud which questions it leaves open.
The rest of this is about getting that agreement in writing before anyone runs anything: the customer's task and acceptance criteria, which conditions must match and which must be allowed to differ, which protocol the available evidence can carry, how every attempt gets recorded, and what sentence each possible outcome would justify. The example used throughout is invented — a paper protocol, not a benchmark, and no result anywhere in it.
Start with the customer's task and acceptable output
A comparison is useful only in relation to a decision someone has to make: replace the current tool, fund the migration, add the module, keep the incumbent. That decision tells you the task the products must perform, and the task tells you what "done" looks like. Write the acceptance definition before the measurement, because a metric chosen first will quietly redefine done.
Concretely: not "processes requests fast," but "produces a routing report our triage lead can send to a queue owner without rewriting it." The second version names the consumer, the format, and the standard of usability. It also implies what fails: a report that arrives quickly and gets half-reworked by hand was never finished. Speed, unit cost, operator minutes, and accuracy are all legitimate headline measures, but they are legitimate downstream of an agreed acceptance definition. A fast output that misses the customer's quality bar cannot be promoted to "the winning result" simply because it was fast.
Two habits keep this honest. First, ask whether the customer would have named this criterion a week ago, before seeing either column. If not, you have selected a measure rather than found one. Second, keep the claim bounded to the task. A comparison can support "better on this task under these documented conditions." It does not support "better product," which is a claim about everything the products do, most of which you did not test.
If the customer hasn't articulated their criteria, help them write it. That work is part of what they're buying, and it costs far less than discovering the criteria after a run.
Separate conditions that must match from differences worth observing
List the conditions twice. On the first list, everything that must be equivalent for the comparison to mean anything: the input set (one frozen batch, the same file to both sides), product versions pinned and recorded, the output format specification, the scoring key, the definition of the observation period, the rule that neither side sees the key during a run, and the tuning each product is permitted. On the second list, everything that may differ — with a written reason for each.
The second list is where most proposals fail. Forcing identity onto a step that one product genuinely performs differently is not fairness; it is a claim that the difference doesn't matter. Sometimes that claim is true. Sometimes it deletes the exact capability under consideration. The most common version of this mistake is insisting both sides receive identically pre-processed input, which is defensible only if input preparation is outside what the customer is buying. If a product's selling point is that it swallows the customer's messy file as-is, then pre-cleaning the file for everyone has quietly removed the thing being evaluated.
Human work needs its own treatment because it hides in gaps between steps. For every manual action in the workflow, record who performs it, whether the customer's own staff could do it, whether it is one-time setup or repeats per request, and whether it sits inside or outside the headline measure. A route that needs a person at the beginning and a route that needs a person at the end can produce identical totals and completely different operational experiences. If the assistance isn't in the number, say so next to the number.
One more condition deserves suspicion: the answer key. If the key is built from the customer's historical routing, then "correct" partly means "agrees with the current process," including whatever that process gets wrong. That can be exactly what a customer wants to measure — just state which question you're answering, because agreement with the incumbent and accuracy on the task are not the same thing.
No single superficial rule guarantees fairness. "Both sides got the same time limit" can mean one side received a ceiling it can't reach and the other a ceiling it never needed to approach. Name the trade-offs explicitly, and let the customer argue with them before the run rather than after.
Choose the comparison the evidence can support
Here is an invented illustration, not a description of any product on the market. Suppose a customer wants to compare two products on one task: turning incoming service requests into an owner-ready routing report. The customer supplies a frozen export of past requests. The agreed acceptable output, per request, is a category from the customer's category list, a priority, an owning queue drawn from the customer's own queue names, and a short quote from the request that justifies the placement. Complete and usable means all four fields present, category and owner queue matching the customer's agreed key, priority within one level of that key, and the quote appearing verbatim in the source request.
Call the two candidates Route A and Route B. Route A does not accept the customer's export directly: its taxonomy and field layout differ, so someone converts the file — mapping the customer's category strings onto Route A's taxonomy and reformatting fields — before a run. Because that mapping is defined in advance, Route A's output can be rendered back into the customer's own category names mechanically. Route B takes the raw export without conversion, but its queue names come from its own list, so after a run a person reconciles each placement into the customer's queue names.
Three protocols are available, and they answer three different questions.
A common end-to-end protocol. Both routes start from the same frozen raw export. Every human step on either side appears in the record: the conversion work on Route A's side, the reconciliation on Route B's side, who did each, and whether the customer's own staff could do it. This answers the customer's most likely real question — from our file to a report we can act on, what does each route take, and what does each report look like? Its limitation is accounting: it charges per-batch conversion to a route whose mapping might become standing configuration, and per-attempt reconciliation to a route whose queue mismatch might be a one-time load. Unless one-time setup and recurring per-request work are kept in separate columns, the summary overstates the ongoing cost of whichever side has the one-time work.
A narrower post-normalization demonstration. Both routes receive input already mapped into Route A's taxonomy and reformatted, and both outputs are converted into the customer's queue names before scoring. This answers a clean question: on identically prepared input, which route's placements match the key? It is also not neutral. Route B's ability to accept the raw file has been removed from the test — that capability was part of what the customer might be buying — while Route B's reconciliation cost has been moved outside the record. The protocol strips one route's advantage and hides one route's work. If the customer's actual question is "which of these classifies better once the plumbing is equal," this is a good instrument. It is not the question on the slide.
A configured demonstration of one route. Route B runs from the raw export, with the customer's queue names loaded during setup. The setup is disclosed: who loaded them, how, and what it involved. This answers "can this route produce an owner-ready report from our raw file at all, and what did that require?" It cannot support any comparative claim, because only one route ran. That is a scope limit, not a defect. When a customer wants to see what a route can do before investing in a full comparison, or when the other side can't be run under matched conditions without distorting the task, a disclosed configured demonstration is often the right proposal.
For this fictional routing task, the common end-to-end protocol answers the decision, provided one-time setup and per-request work are reported separately. The proposal should say that, and then name what remains unsettled: whether Route B's queue-name mismatch is a one-time configuration or a recurring reconciliation. That question changes what the run measures, so it has to be answered before running, by the vendor's own statement or by an agreed pre-step — not by a guess after the numbers come in.
Agree how every attempt will be recorded
Define the outcome classes before the run, per request, so that every attempt lands in exactly one of them: complete; incomplete or incorrect — any field missing, format unmet, category or owner queue not matching the agreed key, priority more than one level from the key, or the quote not appearing verbatim in the source request; corrected — delivered, then repaired by a person, logged as corrected, with the correction counted; and failed — nothing produced, or produced from the wrong input. Then fix the two rules that decide how much a result is worth: the rerun rule and the stopping rule. If a route may retry a request, state the limit in advance and keep every attempt visible, including the ones that preceded the delivered version. If the run stops when the batch is exhausted, say that — not when some accuracy target is reached.
Three practices corrupt a record faster than anything else. Discarding runs after seeing the result. Reporting a "success rate" without stating the number of attempts and the population they came from. Treating a handful of attempts as evidence of a stable difference between products. The first is a protocol violation. The second is an incomplete sentence. The third is a statistical claim the design cannot carry, and no attractive outcome repairs it. If you did not fix the attempt count and the population in advance, you have a description of what happened, not a comparison you can generalize from.
Keep setup and correction work visible even when the headline measure excludes it, and explain the boundary in the same breath — "the reported time covers processing only; conversion and post-run reconciliation are recorded separately below." That sentence costs one line and prevents the most common follow-up question in the room.
Disclosure discipline in computing benchmarks offers a useful, bounded parallel. SPEC's CPU2017 run and reporting rules require disclosure of relevant hardware, software, configuration, and performance-affecting changes so that results can be understood and reproduced (https://www.spec.org/cpu2017/Docs/runrules.html, recorded as checked on 2026-09-08). That is the rulebook for one specific benchmark, not a universal protocol for vendor evaluation, and it proves nothing about your customer's task. What transfers is the habit: configuration and anything that could affect performance belong in the record, because a result nobody can explain or reproduce is not evidence.
Write the interpretation before the results slide
Draft one paragraph per plausible outcome while the outcome is still unknown. For each, state what that result would support and what it would not settle. If the customer's incumbent performs well, or your product doesn't, the paragraph should already exist. And if you find you can't write a fair interpretation for one side's likely win, that asymmetry is a finding about the protocol, not a rhetorical problem.
Writing interpretation in advance blocks the most damaging later move: discovering, after a disappointing run, that a second configuration or a subset of the batch is "actually the fairer comparison." That revision may even be technically reasonable, and it is still unavailable, because the reason it appeared at that moment is the result.
Keep three documents distinct. The proposal says what will be measured and how. The run record says what happened, including the attempts that didn't work. The sales claim is what you say afterward to a different audience, and it can't be wider than the record. A proposal is not a benchmark result, and an appealing outcome does not retroactively justify the design that produced it.
What belongs on the slide, then, is not a number. It's the task, the acceptance definition, the conditions that were held equal and the ones deliberately allowed to differ with reasons attached, the attempt rules, and the one design question still open — in this fictional case, whether the queue-name reconciliation is standing configuration or recurring work. That last item is a better thing to put in front of a customer than a pair of columns, because it gives them a decision to make rather than a result to argue with. When the comparison is finally run, the run will have been defined by the people who have to live with it, and the interpretation will already be written.
Frequently asked questions
Why not simply show the winning result?
The comparison often arrives pre-decorated: someone has already run a product against something and liked the outcome. But if no benchmark has been agreed, the metric may be chosen before the definition, which lets the rules of fairness get written after the outcome is known. Decide the comparison before you can see it.
What are the three honest comparison shapes?
A common protocol runs both products through the same agreed procedure. A configured demonstration shows what one route can do under conditions disclosed in full. A narrower comparison isolates one answerable piece of a task that cannot fairly be run end to end. Pick the shape that matches the question and say what it leaves open.
What conditions must match, and what may differ?
Conditions that must match include the input set, product versions, output format specification, scoring key, observation period, rule that neither side sees the key, and permitted tuning. Conditions that differ need a written reason. Forcing identity onto a step one product genuinely performs differently can hide the capability under consideration.
What is wrong with pre-normalising both sides before a comparison?
Pre-normalising can strip one route's real advantage and hide its work. In the invented routing example, converting everything into Route A's taxonomy and queue names removes Route B's ability to accept the raw file, while moving Route B's reconciliation cost outside the record. That may answer a classification question with plumbing equal, but not the customer's real task.
How should attempts and outcomes be recorded?
Define outcome classes before the run: complete, incomplete or incorrect, corrected, and failed. Fix the rerun rule and stopping rule in advance, and keep every attempt visible. Do not discard runs after seeing the result, report a success rate without attempts and population, or treat a handful of attempts as evidence of a stable difference.