Why do AI pilots impress and production disappoint? Because most pilots are built to show it can work, not whether it will. The Three-Gap Test: material, people and ownership.
The pilot went well. The team was enthusiastic, the outputs were impressive, and the steering group agreed to scale. Six months later the tool is used by a handful of people, the promised savings have not appeared, and nobody can quite say why the success did not travel.
We see this pattern often enough to give it a name: pilot theatre. It is not a failure of technology and rarely a failure of effort. It is a design flaw in the pilot itself. Most pilots are built to demonstrate that something can work. Very few are built to discover whether it will work under the conditions of ordinary Tuesday-afternoon operations.
A pilot that cannot fail is not evidence. It is a rehearsal.
Pilots share four conditions that production does not. The people are volunteers who want it to succeed. The documents are chosen, cleaned or well-behaved. The support is close, because the project team is watching. And the stakes are low, because nobody depends on the output.
Research on professional AI use points at the same danger from a different direction. In a 2023 field experiment at Harvard Business School with 758 management consultants, participants using AI completed tasks faster and did more of them. On one task deliberately placed outside what the tool did well, those using AI were 19 percentage points less likely to produce a correct answer. The lesson for pilots is uncomfortable: performance inside the frontier can hide the fact that you have not yet met the tasks outside it. A pilot that only uses well-behaved material never meets them.
Before scaling any pilot, we ask whether it has closed three gaps. Each one corresponds to a way that a promising demonstration quietly stops resembling real work.
The question is never whether the tool can do the task. The question is whether the organisation can do the task with the tool, on a bad day, without the project team in the room.
The remedy is not more pilots. It is a different kind of pilot, designed before it starts around evidence rather than momentum.
Write the decision rule in advance. State what result would lead you to scale, adjust or stop, in numbers, before anyone sees a single output. A rule written after the results is a justification.
Include a deliberate stress sample. Reserve part of the pilot for the ten or twenty worst examples of the task, chosen by the people who normally suffer them. Report results for that sample separately. If quality falls apart there, you have learned where the boundary lies.
Measure the baseline first. You cannot show improvement against a method nobody timed. Record how long the task takes now, how often it needs correction, and what it costs, before the pilot begins.
Keep the checking visible. Track review time and rework as part of the pilot, not as an afterthought. A pilot that ignores them will overstate its result.
Set an exit criterion and a named owner. Decide what a good production state looks like, who will hold it, and how they will know when quality drifts.
By the end of a well-designed pilot, the steering group should be able to answer five questions on one page. What did we test, and on what material? Who used it, and how were they chosen? What did it cost per accepted output, all-in? What happened on the hardest examples? And who owns it from Monday morning?
If any answer is missing, the honest recommendation is not to scale. It is to extend the pilot in the direction of the gap. That is not a delay. It is the cheapest point at which to discover the problem.
Pilot theatre persists because everyone involved has a reason to like a good result. Sponsors want a success story, vendors want a conversion, and project teams want recognition. None of that is cynical, but it means the design of the pilot needs to be protected from the enthusiasm around it. Independent review of the pilot design, before launch, is inexpensive and often decisive.
Consider an illustrative case, not drawn from any real engagement. A programme team pilots AI to summarise field reports for a monthly briefing. The pilot uses a dozen recent reports written by the most careful staff, in one language, in a consistent template. Results are excellent, and the sponsor approves rollout to all regional offices.
In production, reports arrive in three languages, some as photographed pages, some with contradictory figures across sections. The summaries look fluent, so reviewers begin to skim. A wrong figure reaches a briefing. Nothing was wrong with the tool. The material gap, the people gap and the ownership gap were all open, and the pilot had been designed in a way that could not reveal them.
Had the pilot included the twenty worst reports, a sample of sceptical regional staff and a named owner for the summary workflow, the same tool might still have been adopted, but with a narrower scope, a review step matched to the risk, and a person responsible for noticing when quality slipped.
Scaling does not need to be a single step. We recommend widening in stages, each with its own check. First extend to a second team with different material, and repeat the stress sample. Then extend to a second location or language. Then hand over ownership formally, with the operating instructions written down and a review date set.
At each stage, ask the same question: what did we learn that the previous group could not have told us? If the answer is nothing, the pilot was probably not testing anything. If the answer is a surprise, the design is working, and the surprise is cheaper to meet now than after full rollout.
A pilot earns the right to scale when it has been tested on the awkward material, run by ordinary users and handed to a named owner, not when it has produced a good demonstration. Design for the street, not the stage.
If this is the question on your desk, a thirty-minute call tells you whether the service fits, or that you do not need us yet.