azyware
Business

What changes when AI moves from pilot to production

EZ
Eazyware
· 7 min read
Quick answer

What changes when an AI pilot moves into production?

Moving AI from pilot to production changes six things: who owns it, how it is measured, what it is allowed to do, how fast it must respond, what it costs every month, and who answers at 2am. The model is usually the only part that stays the same.

When AI moves from pilot to production, six things change: ownership moves from an innovation team to a system owner, measurement moves from opinion to an evaluation suite, permissions become scoped and gated, latency becomes a target, cost becomes a monthly run-rate line, and someone has to answer at two in the morning. The model usually stays the same.

Most pilots that die do not die of bad accuracy. They die because nobody budgeted for the six changes above, so the pilot has nowhere to go once it has proved its point. This article is the delta list, the cost of each part, and where we advise clients not to make the jump.

A pilot and a production system are different objects

A pilot exists to answer a question: can this work at all, on our data, for our users. It is allowed to be slow, narrow, manually operated and supervised by the person who built it. Its success criterion is a decision.

A production system exists to do work every day without its authors present. It has an owner with a budget line, users who did not volunteer, uptime expectations, an audit trail, a cost that appears on someone's forecast, and consequences when it is wrong. Its success criterion is a metric that stays inside a threshold, month after month.

The confusion between the two is why so many organisations have a folder of successful pilots and nothing running. The pilot proved feasibility and was then asked to become a service, which is a different build with different work in it. We charge for that work openly rather than describing it as deployment, and why AI pilots never reach production covers the organisational side of the same failure.

The delta, line by line

Read this as a scope list for the production build, not as criticism of the pilot. A pilot that did all of this would have been a slow and expensive way to answer a yes or no question.

DimensionIn the pilotIn production
OwnerThe person who built itA named system owner with a budget line
UsersVolunteers who forgive failuresEveryone, including people who did not ask for it
Quality checkSomeone tried it and liked itFrozen eval set, scored on a schedule, gated on merge
Data accessA broad service accountScoped credentials, per-user permissions, row-level rules
ActionsSuggestions a human retypesGated tool calls with limits and an audit log
LatencyWhatever it wasA target with percentiles, and a fallback when it is missed
CostA trial credit nobody watchedA monthly run-rate with budgets, alerts and cost per task
FailureTry again laterRetries, degraded mode, escalation, incident process
ChangeEdit the prompt and refreshVersion control, regression tests, staged release, rollback
SupportAsk the builder on SlackA named response time and an on-call route

The two rows that surprise people are permissions and cost. In a pilot, one service account reads everything, which is fine for ten friendly users and unacceptable the moment the system answers questions for people with different access rights. Permission-aware retrieval is production work that the pilot was right to skip and wrong to ignore.

What production costs that a pilot did not

Three costs appear at the transition, and all three are predictable.

Build cost for the production version. This is engineering: permissions, observability, evaluation harness, error handling, admin tooling and integration hardening. It is usually a similar size to the original pilot or larger, which is the number that shocks a sponsor who thought the pilot was most of the work. An AI-accelerated MVP starts at $26,500 or ₹17,60,000, and LLM application development at $21,000 or ₹13,60,000; every starting price is published on the pricing page.

Run cost. Model usage scales with adoption, and adoption is the thing you were hoping for. You pay providers through your own accounts, and we set budgets, routing and dashboards so the number is forecastable rather than discovered. Model the figure at your expected volume with the AI agent ROI calculator before you commit, so the business case survives success rather than being broken by it.

Support cost. Someone maintains this now. Eazyware Care Plans start at $1,000 or ₹68,000 a month for Essential cover with eight-hour response, $2,500 or ₹1,60,000 for Standard at 24 x 5 and four-hour response, and $5,250 or ₹3,40,000 for Enterprise at 24 x 7 with one-hour response and a named engineer. The AI system add-on at $750 or ₹40,000 a month covers evals, cost monitoring, prompt regression and re-indexing. What that work involves in practice is the subject of care plans for AI systems.

How long does the transition take?

For a pilot that genuinely worked, six to twelve weeks is the usual range for the production build, including a shadow-mode period. Most scoped builds at Eazyware take eight to sixteen weeks, and the transition sits at the shorter end because the hard question has already been answered.

The sequence we run is consistent. First, freeze the pilot and write down what it proved, in numbers. Second, build the evaluation set from real usage, which is the artefact that makes everything after it measurable. Third, rebuild the data access properly with scoped permissions. Fourth, add observability so every request can be traced end to end. Fifth, run in shadow mode against live traffic, with humans reviewing proposals. Sixth, release to a limited group behind a flag, then widen.

Skipping step five is the most common shortcut and the most expensive one. Shadow mode is where you find out that the pilot's users were unrepresentative, which they almost always were.

The organisational changes nobody puts in the plan

Someone owns it now

A production AI system needs a named owner in the business, not only in engineering: the person who decides what an acceptable error rate is, who signs off approval thresholds, and whose team absorbs the escalations. Without that person the system has no one to defend it at budget time, and it quietly stops being maintained.

Service levels become a conversation

Production means agreeing what good looks like and what happens when it is not met. Borrowing the discipline of service level objectives from site reliability practice helps here; Google's SRE book chapter on service level objectives sets out how to choose a small number of indicators and hold yourself to explicit targets rather than aspirations.

The budget line moves

Pilots are funded from innovation or project budgets. Production systems are funded from operating budgets, which are annual, defended and scrutinised differently. Making that move early, while the sponsor is enthusiastic, is easier than making it in month nine when the trial credits run out.

When a pilot should not go to production

We advise against the jump more often than clients expect, for three reasons.

The pilot proved the technology, not the value. If people used it because it was new and cannot tell you what it saved them, productionising it buys a running cost against an unproven benefit. Go back and measure the workflow first.

The volume is too low. A workflow that happens forty times a month is rarely worth a production system, its evaluation harness and its support contract. Automate the surrounding process, or wait until volume justifies it.

The underlying data is not ready. If the pilot worked on a curated extract and the real source is inconsistent, out of date or scattered across systems, the production system will be worse than the pilot in a way nobody will forgive. That is a data project, and it should be scoped and funded as one.

A worked example

For a D2C brand we built personalisation and a WhatsApp support agent, described in the personalisation and WhatsApp agent case study. The production work that a pilot version would not have included was the part that mattered: handling peak-season load, making sure the agent respected customer consent and messaging rules, gating anything that touched an order, and reporting cost per resolved conversation so the finance team could see the trade. The model choice was close to the least interesting decision in the project.

Checklist for the transition

  • Write down what the pilot proved, as a number, and what it did not test
  • Name the business owner and get their budget line confirmed
  • Build the evaluation set from real pilot usage before anything is rebuilt
  • Re-do data access with scoped credentials and per-user permissions
  • Define the latency target and the fallback when it is missed
  • Instrument cost per completed task and set a budget alert
  • Agree approval gates for every action the system can take
  • Plan shadow mode and book the reviewers' time
  • Choose a support tier before launch, not after the first incident

From POC to production: the checklist is the engineering companion to this article, What makes an LLM application production-ready covers the technical bar in detail, and our fixed-price programmes set out how we structure discovery, proof and build so the transition is planned rather than improvised.

A pilot answers a question and a production system carries a promise, and the work between them is the actual project.

Frequently asked questions

Why do so many AI pilots never reach production?

▾

Because the production work was never budgeted. Permissions, evaluation harnesses, observability, error handling, approval gates and support are a build of similar size to the pilot itself. Sponsors who assume the pilot was most of the work find there is no funding left at exactly the moment the results look good.

How long does it take to move an AI pilot into production?

▾

Six to twelve weeks for a pilot that genuinely worked, including a shadow-mode period where the system runs alongside humans on live traffic. Most scoped Eazyware builds take eight to sixteen weeks; a pilot-to-production transition sits at the shorter end because the feasibility question is already answered.

Does the model change between pilot and production?

▾

Rarely at first, and often later. The pilot's model choice is usually fine to launch with. What changes is that production systems route between two or three models for cost and reliability, get benchmarked against your own evaluation set on every provider release, and need a migration plan for when a version is deprecated.