Why AI systems degrade after launch and what to do about it
Why do AI systems degrade after launch and what should you do about model drift in production?
AI systems degrade as providers update models, data shifts and prompts age; evaluation regression and retraining loops keep them working. AI model drift in production is a condition you manage, not a bug you fix once: a golden eval set run on every change, production sampling, pinned model versions and a named owner.
AI model drift in production has three causes, and none of them is your code changing. Providers update or retire the models behind an API, so the same prompt returns different behaviour. The world shifts, so the documents, customers and questions the system sees drift away from the ones it was built on. And prompts age: the examples, rules and edge cases written in month one stop matching the product in month nine. An AI system that was accurate at launch and untouched since is almost certainly less accurate now. The fix is not heroics; it is an evaluation regression loop, production sampling, version pinning and a retraining or re-tuning cadence, run by someone whose job it is. That is what our maintenance and support service exists to do.
Traditional software degrades slowly and visibly: a dependency goes end-of-life, a server fills up. AI systems degrade quietly, because the output still looks plausible. A support agent that started giving slightly wrong refund policies does not throw an error. Someone has to be looking.
What AI model drift in production actually is
Drift is any gap between how the system behaves now and how it behaved when it was evaluated and accepted. It shows up as lower accuracy on the same questions, more hallucinations, changed tone, longer or shorter outputs, more escalations, or a jump in cost and latency. The term covers several different mechanisms, and the response depends on which one is at work, so the first job of maintenance is diagnosis rather than a blanket "retrain".
The four causes and their fixes
| Cause | What happens | How you notice | What to do |
|---|---|---|---|
| Provider model updates | The model behind an alias changes; behaviour, format and refusals shift | Eval scores move with no change on your side; format parsing errors rise | Pin explicit model versions; run evals on every provider release; migrate behind routing |
| Model deprecation | The pinned version is retired on a published date | Provider notice; eventual API errors | A migration plan with a regression run before switching traffic |
| Data drift | Documents, products, policies or user language change | Retrieval misses; questions about things the corpus does not cover | Refresh the index; add new golden questions; monitor unanswered queries |
| Prompt and logic ageing | Rules and examples no longer match the product or the policy | Specific failure clusters in production samples | Prompt versioning with evals; scheduled review with the business owner |
| Trained-model decay (ML) | The relationship the model learned no longer holds | Precision or recall falls on recent labelled data | Retraining loop on a cadence with a holdout comparison |
Provider updates and LLM drift
Most API providers offer both an alias that points at the latest model and dated or versioned snapshots. If your system calls the alias, the provider can change its behaviour under you at any time. Pinning a snapshot removes the surprise but adds the obligation to migrate before the snapshot is retired. The right pattern is pin, evaluate each new release against your golden set, and migrate deliberately behind a routing layer so that the switch is a decision with evidence rather than an event you discover from users. Handling model deprecations walks through the migration itself, and a model-agnostic routing layer across OpenAI, Anthropic, Google and open-weight models is how we keep that switch cheap.
Data drift in retrieval systems
A RAG system is only as current as its index. Policies change, products launch, the help centre is rewritten, and the retrieval layer keeps serving last quarter's chunks. Data drift is detected by watching two things: questions the system could not ground in any retrieved document, and golden questions whose expected answers have changed. The fix is operational: an index refresh pipeline that runs on a schedule and on source changes, with the eval suite run after each refresh to catch chunks that now retrieve badly. How to measure RAG quality covers the metrics that make drift visible.
Prompt ageing
Prompts are code that nobody reviews. The few-shot examples written for the launch product describe features that have since changed; the rule that said "never discuss pricing" is wrong now that pricing is public; the escalation criteria were written before the second support tier existed. Prompt ageing is caught by production sampling: a weekly review of a random sample of conversations plus every flagged failure, clustered by cause. Each cluster becomes a prompt change, and each prompt change runs against the eval suite before release. Treat the prompt like code, with versions, diffs and a rollback, as Prompt versioning and evaluation sets out.
The evaluation regression loop
Everything above depends on one asset: a golden evaluation set that reflects what the system must do, scored automatically where possible and by people where necessary. The loop is simple. On every change, whether yours (prompt, retrieval, code) or theirs (model release), run the suite and compare with the last accepted score. A regression blocks release. New failures found in production are added to the set, so it grows with the product. Once a quarter, the business owner reviews the set to remove obsolete cases and confirm that the expected answers are still right. Without this loop, maintenance is guesswork; with it, drift becomes a number on a chart. Evals: the practice that separates AI demos from AI products explains how to build the set.
Retraining loops for classical ML
Trained models such as fraud scorers, demand forecasts and document classifiers decay differently: the pattern they learned stops holding as behaviour changes. The maintenance loop is monitoring of precision and recall on freshly labelled data, retraining on a cadence or on a threshold breach, and a holdout comparison before the new model replaces the old. Feature pipelines drift too, and a silent schema change upstream is the most common cause of a model that looks fine in the dashboard and is wrong in practice. MLOps for mid-size companies covers the tooling.
A worked example
An NBFC's document intelligence system extracted fields from KYC documents with human review of low-confidence cases. Over several months, the review queue grew without any change to the code. Diagnosis found two causes: the provider had updated the vision model behind an alias, shifting confidence calibration, and a regulator's format change had introduced a new document layout the golden set did not contain. The fix was to pin the model version, add the new layout to the eval set with labelled samples, re-tune the extraction prompts against it, and move the alias-to-snapshot switch behind the routing layer so the next update would be a decision. The KYC document intelligence case study describes the system; the maintenance loop is what keeps it accurate.
Team and timeline
AI system maintenance is a standing responsibility, not a project. Under a Care Plan, a named engineer runs the weekly production sample, the eval suite on every change and provider release, the index refresh, cost and latency dashboards, and the deprecation calendar. Essential at $1,000 per month (₹68,000) covers business hours and 10 hours a month, enough for a single application with a stable provider; Standard at $2,500 (₹1,60,000) adds 24×5 cover, 4-hour response and 25 hours; Enterprise at $5,250 (₹3,40,000) gives 24×7, 1-hour response, 60 hours and a named engineer, which is the level for systems handling customers or money. Retraining loops for ML models are scoped separately under AI/ML development. The pricing page lists the plans.
If the system was built by someone else and has no eval set, the first month is spent building one from production logs before any maintenance can be meaningful.
Before you start: a checklist
- Confirm whether the system calls model aliases or pinned versions, and pin them
- Locate or build the golden evaluation set and record the last accepted scores
- Set up a weekly production sample review with a business owner
- Put the index refresh on a schedule and run evals after each refresh
- Subscribe to provider deprecation notices and keep a migration calendar
- Version prompts with diffs and a rollback path
- Add cost and latency to the same dashboard as quality
- Name the person accountable for the system's behaviour this quarter
Glossary
- Model drift: any gap between current behaviour and the behaviour that was evaluated and accepted
- Alias: a provider model name that can point at different underlying versions over time
- Pinned version: a dated or versioned model snapshot that does not change until retired
- Golden set: the curated evaluation cases with expected outcomes that every change is scored against
- Regression: a fall in eval score compared with the last accepted run
- Production sampling: regular human review of a random sample of live outputs plus flagged failures
Related reading
LLMOps for small teams describes the minimum operations that make this loop possible, and What a Care Plan should cost and what it should include covers the commercial side. OpenAI's model deprecations page is the primary source for one provider's retirement schedule and is worth bookmarking whichever providers you use.
AI systems drift because the world and the providers move; an eval loop, pinned versions and a named owner are what keep them where you left them.
Frequently asked questions
What causes AI model drift in production?
▾
Provider model updates and deprecations behind an API alias, data drift as documents and users change, prompts and rules that age as the product evolves, and, for trained ML models, decay in the patterns they learned. Each has a different diagnosis and fix.
How do you detect LLM drift?
▾
Run a golden evaluation set on every change and every provider release and compare with the last accepted score, and review a weekly sample of production outputs with a business owner. Drift becomes a number rather than a complaint.
What does AI system maintenance cost?
▾
Eazyware Care Plans run from $1,000 per month for business-hours cover with 10 hours, to $5,250 for 24×7 with a named engineer and 60 hours. The pricing page lists the three plans and what each includes.