The pattern is so consistent it's almost boring. A company runs an AI pilot. The demo is impressive. Everyone nods. Six months later the pilot is quietly shelved, the budget line disappears, and the organisation concludes that "AI isn't ready for our industry." The model was rarely the problem. The pilot design was.
The three ways pilots die
Death by missing metric. The pilot was chartered to "explore AI capabilities" — which means nobody defined the number it had to move. Without a metric, there is no finish line; without a finish line, budget review kills it by default. A pilot should be an experiment with a hypothesis: "this model can cut first-pass review time by half at equal accuracy." Pass or fail, you learn something budgetable.
Death by demo data. The pilot ran on clean, curated examples. Production runs on scanned faxes, forwarded email chains and the one supplier who still sends TIFFs. The gap between demo accuracy and production accuracy is where trust dies — and once operators stop trusting the tool, no accuracy improvement wins them back. The fix is unglamorous: build the evaluation set from your ugliest real cases first.
Death by missing owner. The pilot belonged to an innovation team; the workflow belonged to operations. When the pilot ended, no one whose bonus depends on the process was accountable for adopting it. Pilots survive when the process owner — not the sponsor — asks for them.
A pilot is not a small demo. It's a small production system — with a metric, real data and an owner.
Three that made it
Contract review at a law firm. The charter named one number: first-pass review time. The evaluation set was built from the firm's own past contracts, including the messy ones, and scored against the partners' own past markups. Lawyers saw the model's confidence on every clause and could reject suggestions in one click — rejections fed back into evaluation. Adoption wasn't mandated; it happened because the tool made the annoying part of the job shorter.
Exception triage at a retailer. Instead of aiming the model at the whole order flow, it was aimed only at exceptions — the 4% of orders that ate 60% of the team's time. Confidence-based routing meant the model handled the routine exceptions and escalated the strange ones. Nobody's job changed except the worst hour of it.
Demand forecasting embedded in purchasing. The previous "AI dashboard" had been ignored for a year. The rebuilt version put the forecast inside the purchasing tool, at the moment of the order decision, with the model's error range shown honestly. Buyers overrode it freely — and the override rate itself became the adoption metric. It fell month over month.
The checklist
One metric, named in the charter. An evaluation set from your ugliest real cases. A process owner who wants it, not a sponsor who funds it. Confidence routing to humans from day one. And a feasibility sprint before the build — so the go/no-go decision costs weeks, not quarters. That's the whole difference between a pilot and a press release.