Why AI pilots never reach production and how to fix it
Why do AI pilots fail to reach production, and how do you fix that?
Pilots die without owners, evals, integration and running budgets; scope for production from day one and pilot on real data. Decide before the pilot what production would require, build the pilot inside those constraints, and set a dated decision point with the evidence needed to proceed written down in advance.
Getting an AI pilot to production fails for four reasons, and none of them is the model. The pilot has no owner who will run the system after launch. It has no evaluation suite, so nobody can say whether it is good enough. It was built on a laptop with a public API key and synthetic data, so the integration, security and data work has not started. And its running cost was never budgeted, so the finance conversation begins after the demo instead of before it. The fix is the reverse of each: scope for production on day one, pilot on real data inside the real constraints, name the owner, build the evals first and put the running cost on the plan. This article works through why pilots stall and what we do differently in AI product strategy and build engagements to avoid it.
What AI pilot purgatory looks like
The pilot demoed well in month two. Leadership was pleased. Then the questions started: can it use real customer data, who signs off the outputs, what happens when it is wrong, how much will it cost at full volume, who maintains it. Each question was reasonable and each was answered with "we will look into it". Six months later the pilot is still running for a handful of users, the sponsor has moved on to a new idea, and the team that built it is on another project. Nobody killed it and nobody shipped it. That state is expensive because it consumes the budget and the credibility that the next attempt needs.
The four causes and the fix for each
| Cause | How it shows up | The fix |
|---|---|---|
| No owner | The sponsor was excited; nobody's job depends on the system working | Name the process owner before the pilot; they define success and present at the decision point |
| No evals | "It seems to work"; quality is judged from a few demo prompts | Build a golden set of real cases first and measure every version against it |
| No integration path | Public API key, synthetic data, no auth, no audit trail; production needs a rebuild | Pilot on real data inside the real security and data constraints, even at small scale |
| No running budget | Cost per request was never measured; finance meets the bill after the demo | Measure tokens per action in the pilot and put the monthly cost on the plan before the decision point |
Scope for production from day one
The single most effective change is to write down, before the pilot starts, what production would require: which systems it must integrate with, which data it must use and where that data may go, who approves its outputs, what audit record must exist, what latency and availability users need, and what it may cost per month. Then build the pilot as a small version of that, not as a separate thing. A pilot built inside the production constraints is slower to demo and faster to ship, and the trade is worth it every time. Our ProofRun programme exists to do exactly this: three weeks against real data, inside the constraints, with an evaluation result at the end. The difference between that and a demo is set out in AI proof of concept vs demo.
Pilot on real data, not on a sample that flatters
Pilots built on cleaned samples or synthetic data succeed in ways that production cannot repeat. Real data has missing fields, inconsistent formats, edge cases and permission boundaries, and finding those in week two of a pilot is the point. Real data also forces the security and privacy conversation early, which is where most pilots would otherwise stall later. If real data cannot be used at all in the pilot, that is itself a finding: the data governance work is the actual first project, and the roadmap should say so.
Build the evaluation suite before the feature
An evaluation suite is a set of real inputs with expected outputs, scored automatically or by reviewers, that every version of the system is run against. Without it, the production decision rests on anecdote and the argument about whether the system is good enough never ends. With it, the decision is a number compared with a threshold agreed in advance. The suite also carries into production as the regression test for every prompt or model change. Building it first, from a few hundred real cases, takes a week or two and is the best investment a pilot can make. The approach is described in evals: the practice that separates AI demos from AI products.
Shadow mode is the bridge
Between pilot and production sits shadow mode: the system runs on live traffic, produces outputs, but a person still acts. Its outputs are compared with what the person did, per intent, for two to four weeks. That produces the evidence for autonomy per intent, which is far more convincing to a cautious stakeholder than any demo. The mechanics are in shadow mode: the right way to launch AI agents.
Scaling AI pilots: the decision point
Every pilot needs a dated decision point with three possible outcomes: proceed to production, stop, or return to foundations because a data or integration gap was found. The evidence required for each outcome is written before the pilot starts: the evaluation score, the shadow-mode agreement rate, the measured cost per action, and the owner's confirmation that they will run it. At the decision point the owner presents the evidence and the sponsor decides. Stopping is a legitimate result and should be treated as one; a pilot that stops cleanly at week six has cost a fraction of one that lingers for a year.
Production AI failure after launch
Some pilots do reach production and then fail there, usually because the operating model was never designed. Nobody reads the dashboard, escalations pile up unreviewed, a provider updates the model and quality drops unnoticed, and users quietly stop using it. The fix is the operating model: a weekly review owned by the process owner, a dashboard of quality and cost, an evaluation run before every change, and engineering time to act on what the review finds. That is what a Care Plan provides when the in-house team cannot, and it is a line on the plan from the start, not an afterthought.
A worked example
A field-service software company had run a pilot of an in-app assistant for six months with a small group of friendly customers. It answered questions from product documentation well in demos, but it had been built outside the product with a public API key, had no access to customer job data, and had no measurement beyond a thumbs-up button. The restart scoped production first: the assistant would live inside the product, respect each customer's data permissions, answer from documentation and the customer's own job history, and be measured on a golden set built from real support tickets. The evaluation suite was built in the first two weeks, the assistant was rebuilt inside the product's permission model, and shadow mode ran against support agents' real answers for three weeks. It shipped to all customers within a quarter of the restart. The build is described in our in-app copilot case study.
Team and timeline
For a new use case, a ProofRun runs for three weeks at $6,250 to $10,500, or from ₹4,00,000, and ends with an evaluation result, a measured cost per action and a production plan; a Launch 6 build over six weeks from $26,500 then takes it to production. For a stalled pilot, a Sprint Zero of ten working days assesses what exists, builds the evaluation suite and produces the production scope and cost. Your side names the process owner and provides real data under the agreed constraints; we provide the evaluation harness, the build and the operating model. All programme prices are on the pricing page, and the fixed price and fixed date are part of the discipline that stops pilots drifting.
Before you start: a checklist
- A named process owner whose job depends on the outcome
- The production requirements written down: integrations, data, approvals, audit, latency, cost
- Real data available under the real constraints, even at small scale
- A golden set of a few hundred real cases with expected outputs
- Success thresholds agreed before the first version runs
- Tokens per action measured and a monthly cost forecast on the plan
- A dated decision point with proceed, stop and return-to-foundations as outcomes
- An operating model for after launch: who reviews, who changes, who pays
Questions clients ask
- Our pilot works; why can't we just switch it on? Because it was built outside the production constraints. Integration, permissions, audit and cost have to be designed, and the evaluation has to exist to show it is safe.
- How long should a pilot take? Three to six weeks, on real data, with a decision point at the end. Longer pilots without a decision point are the purgatory this article describes.
- Should we pilot several use cases at once? No. One use case with a full path to production teaches more than three demos, and its evaluation harness and platform carry over.
- What if the pilot fails the threshold? Stop or return to foundations. A clean stop protects the budget and credibility for the next attempt.
Related reading
Building an AI roadmap that survives the first quarter, what a six-week AI MVP actually contains, and copilot adoption: why most AI features die in a month. The NIST AI Risk Management Framework is a useful reference for the governance items a production system needs.
Decide what production requires before the pilot starts, build inside those constraints on real data, and set a date to decide; pilots built that way either ship or stop, and neither outcome is purgatory.
Frequently asked questions
Why do most AI pilots fail to reach production?
▾
They lack a process owner, an evaluation suite, an integration path built inside real security and data constraints, and a running-cost budget. The demo succeeds and the production questions arrive afterwards with nobody to answer them.
How do you move an AI pilot to production?
▾
Write down the production requirements first, rebuild the pilot inside them on real data, build a golden evaluation set, run shadow mode against live traffic, and hold a dated decision point with thresholds agreed in advance.
How long should an AI pilot last?
▾
Three to six weeks on real data with a decision point at the end. A pilot running for months without a decision is in purgatory and should be stopped or restarted with production scope.