Skip to content

Your Pilot Succeeded With Heavy Support. What Does That Prove About a Wider Rollout?

Business

Your Pilot Succeeded With Heavy Support. What Does That Prove About a Wider Rollout?

A pilot summary often reads like this: schedules delivered daily, the customer's team used them, the result was favorable. Somewhere else in the account—usually an operations channel rather than a slide—there is a second list: the files someone cleaned every evening, the questions someone chased across three emails, the exceptions someone fixed by hand before the output went out the door. Both lists are true. The trouble starts when the proposal keeps the first and quietly drops the second, and then asks the result to prove something about a rollout that will not have the second.

The pilot is evidence about the arrangement you tested. It becomes evidence about the software alone only if the software was the whole arrangement, which it almost never is. Necessary human work is not automatically a defect to hide, and it is not automatically a virtue either—a service can be genuinely staffed and genuinely unaffordable, and you will not find out which by assuming. What follows is how to put the tested arrangement next to the result, compare continuing it against a version with less support, and write the proposal that the evidence can actually carry.

Rebuild the arrangement before you quote the result

Take an invented case, used here as a teaching example and not reported from any real engagement. A supplier sells delivery-scheduling software. Across a twelve-week pilot with one regional distributor operating three depots, the supplier delivered a schedule file every working day, and the customer's dispatchers ran routes from it. That is the result: usable daily schedules, used. Everything else in this example is stipulated, not observed.

Producing it took three kinds of work that the schedule file does not show.

First, preparation. The order export arrived each evening in whatever format the customer's system produced—addresses written in several styles, some site codes blank. Supplier staff cleaned each file, normalizing addresses and filling blanks from a mapping they maintained for this customer.

Second, resolution. Some delivery sites were genuinely ambiguous; two depots could plausibly serve them. Supplier staff raised these with the customer's dispatcher, who answered that same evening, in time for the next morning's routes.

Third, repair. Some orders could not be routed as given—a load too large for one vehicle, a time window no sequence could meet, a site with a locked gate. Supplier staff fixed these by hand before delivery: splitting loads, moving an order to the next day, overriding a sequence the dispatcher knew was impossible.

Now notice the conditions wrapped around those three jobs: a tight evening window, a customer contact who answered inside it, a supplier team able to do the work in that window, one customer's data, twelve weeks. And notice what the record does not contain. It names the work; it does not measure the hours it took. It shows schedules being used; it does not show whether deliveries themselves improved. It does not establish whether the schedules would have been usable without the repairs, because the repairs never stopped.

Those absent measurements are not sloppiness in the pilot. They are the difference between two questions that look similar and are not. The UK government's Digital Trade blog published a piece on 14 April 2025, "Understanding the evaluations role in measuring the impact of AI interventions across Government," whose section on evaluation techniques separates monitoring, process evaluation, impact evaluation, and value-for-money evaluation—that is, it distinguishes examining how an intervention was delivered from examining what difference it made. The context is UK public-sector AI evaluation, and the page is about how to approach evaluation rather than a report of any results, so it settles nothing about commercial pilots of any size. But the distinction travels. In the invented case, the record is mostly a delivery record: something was produced, daily, and used. Whether anything got better for the distributor is a separate question the twelve weeks did not answer. Keep those apart in the write-up, or the favorable delivery record will quietly start doing impact work it cannot do.

The same discipline applies to attribution. "The dispatchers used the schedules" is an observation. "Planning is easier now" is an inference, reasonable but unmeasured. "The software works" assigns to the software an outcome produced by the software plus evening preparation plus joint ambiguity resolution plus manual repair. Each of the first two is defensible within its own evidence. The third is not what the pilot establishes; it is the claim a rollout would need to establish separately.

Resist the urge to file the cleaning and the repairs under implementation details. That work is part of the service the customer received; if it is uncomfortable on the page, the decision in front of you is whether you want to sell it, not whether to leave it out and keep the result.

Continue it, change it, or stop claiming it

Three routes follow from the pilot, and they differ in how much of the tested arrangement they preserve.

Continue the staffed service. This is the route closest to the evidence. The same kind of preparation, resolution, and repair stay inside the offer; the customer's contact and turnaround stay as stated expectations rather than assumptions. What it inherits is not a guarantee. Continuing for the same customer at the same scale is close to what was tested. Continuing for more customers is already a change: more evening files competing for the same attention, more dispatchers who may or may not answer that night, new export formats, and no inherited mapping for the new accounts—that mapping was this customer's, and it took weeks to accumulate. The staffing that made the window work may not divide cleanly. So the honest version of this route is a supported service with a described staffed commitment, not a rescaled copy of the pilot.

Change the model. In the tempting version, customers upload untreated exports and receive no routine repair. Name what leaves the room: the file cleaning, the joint resolution of ambiguous sites, the manual exception work. What remains is the routing logic applied to inputs as they actually arrive. If the proposal keeps the pilot's result as its expected outcome, it is claiming exactly what it deleted. You cannot take the support out of the diagram and leave the outcome as proof of what stays. The pilot still tells you something useful here—that the logic can produce usable schedules when the input is fit for it, and that when it isn't, somebody handles it. Under the changed model, that somebody is the customer or nobody, and a schedule that reaches a dispatcher with unresolved exceptions is not the same object as the one the pilot delivered.

Hold the wider claim. Holding is not the same as stopping. You can keep serving the customer in front of you under the arrangement you already ran while declining to promise the rollout. This is the right route when neither offer's necessary conditions are established: you do not know whether the staffed service can be staffed at larger volume, and you have not tested the self-serve version at all. In that state, the only claim the record supports is about what was tested—staffed scheduling produced usable daily schedules for this customer under named conditions for twelve weeks. That sentence is small and true, and it is worth more than a large one you cannot defend.

Neither route inherits unmeasured reliability or unmeasured margin. "Scalable" is not a property a pilot confers; it is a claim that needs a stated condition every time you use it—scalable to how many customers, at what support level, over what mix of workloads.

What the follow-up has to be able to show

If you go the changed route, the evaluation should test the dependency you removed rather than the one you already demonstrated. Nobody needs another twelve weeks of proof that the routing logic can sequence prepared orders. The open question is whether usable schedules survive without preparation and repair, and who does the work when they don't.

Use the actual input. Take the customer's real export, unmodified, in the form it is genuinely produced—including the blank site codes and the inconsistent addresses. A curated sample tests your cleaning, which is the thing you are trying to remove from the arrangement.

Record exceptions and their handling, case by case: what the tool could not resolve, who resolved it, how the schedule was affected before and after. Do not establish rates in advance; establish that you will count.

Record customer effort as a finding, not a footnote. If a dispatcher now spends part of each evening normalizing addresses before upload, that effort did not disappear. It moved, and it moved to someone who may have less capacity to absorb it than your team did. A customer who cannot sustain the new routine may not complain; they may drift back to manual scheduling, and the failure may look like churn rather than like a design problem.

Decide before the trial begins what would end it. If your staff are still repairing exceptions most evenings, the dependency did not change—it relocated. If schedules regularly arrive unusable without repair, the self-serve offer has not been demonstrated, whatever the routing accuracy looks like in isolation. Write those conditions down while nobody has a stake in the answer, because a trial that can only be reported as a success is not a trial; it is a delay with better formatting.

And scope the run honestly. A short period under light load may not surface the kinds of exceptions the staffed pilot encountered. Say what did not come up, and do not treat its absence as reliability.

Write the proposal without the rollout verdict

The document you need has three parts, and none of them is a scale verdict.

What was observed. The delivery facts, in plain terms, with the conditions attached. Something like: Over twelve weeks, we produced a usable daily schedule for [customer]'s three depots, delivered before the morning shift and used by their dispatchers. Producing those schedules involved evening file preparation on our side, same-evening clarification from [customer]'s dispatcher on ambiguous sites, and manual handling of exceptions before delivery. The pilot did not measure delivery performance or the cost of that support work. Three sentences, no adjectives doing load-bearing work.

What the intended offer preserves. If you continue the staffed service, say so and describe the evening preparation and exception handling as part of the service, with the customer's response window named as an expectation. If you propose the changed model, say clearly that it does not preserve those conditions and that its expected result is a hypothesis rather than the pilot's outcome. If you hold, say what continues unchanged and what claim you are not making yet.

What still needs evaluating. One question, appropriately bounded: whether the changed arrangement produces usable schedules from untreated inputs, where the residual work lands, and what the customer must absorb for it to hold.

Then give each route its practical consequence, because readers make decisions from consequences, not from method. Continuing means an ongoing staffed commitment described as such. The changed model means a bounded trial with a stated question and no assurance of the pilot's output in the meantime; it also means telling the customer what work moves to them. Holding means the current service continues, the wider promise waits, and you name the trigger that would change it—either the staffed arrangement holding for more customers than one, or the changed arrangement producing usable schedules with the work accounted for wherever it lands.

The offer you can make today is the one you actually ran: usable daily schedules, produced by a named combination of software, staff, and customer participation, for this customer, under conditions you can describe out loud. The wider claim is not forbidden; it is just unwritten until one of those triggers fires. Until then, the sentence that survives contact with the customer who just renewed is the small, precise one—and it is available immediately.

Frequently asked questions

What does a pilot prove when it succeeds with heavy support?

It proves the tested arrangement produced a result under named conditions. It does not prove the software alone can. If the proposal keeps the pilot result but drops the evening preparation, ambiguity resolution, and manual repairs, it is claiming exactly what it removed.

What work supported the invented scheduling pilot?

Supplier staff cleaned each evening order export, normalized addresses, filled blank site codes, raised ambiguous delivery sites with the customer's dispatcher, and repaired unroutable orders by hand. Those jobs depended on a tight evening window, a responsive customer contact, a supplier team able to work in that window, one customer's data, and twelve weeks.

What are the three routes after such a pilot?

Continue the staffed service, change the model, or hold the wider claim. Continuing is closest to the evidence but scaling already changes conditions. Changing the model removes support work, so it cannot inherit the pilot outcome as proof. Holding keeps current service and declines the rollout promise until a named trigger fires.

How should a follow-up test the changed model?

Test the dependency being removed, not what was already demonstrated. Use the customer's real, unmodified input with its blank codes and inconsistent addresses; record exceptions and who handles them; count customer effort as a finding; decide before the trial what would end it; and scope the period and load honestly.

What should the proposal say if it cannot yet claim a wider rollout?

It should have three parts: what was observed with conditions and support work; what the intended offer preserves or does not preserve; and what still needs evaluating. Then give each route its practical consequence. The small, precise sentence about what was tested is available immediately and is stronger than an unsupported rollout verdict.

More in Business Browse all articles