azyware
Technology

Handling model deprecations: when your provider retires a model

EZ
Eazyware
· 7 min read
Quick answer

What should a model deprecation plan include when your LLM provider retires a model?

Pin model versions, keep the eval suite ready and migrate behind routing with a regression run before switching traffic. A model deprecation plan turns a retirement notice into a scheduled change: evaluate the candidate, fix regressions, shift traffic gradually and keep the old version as fallback until shutdown.

A model deprecation plan has four parts: pin the model versions you depend on so a retirement is a dated event rather than a silent change, keep an evaluation suite ready so a candidate can be scored in an afternoon, route every call through one layer so the switch is configuration, and run a regression comparison before any production traffic moves. Providers retire models on a published schedule, usually with months of notice, and teams that have these four things in place treat the notice as a calendar item. Teams that do not treat it as an incident, often discovered when the API starts returning errors. Deprecation handling is a standing part of our maintenance and support service, and this article sets out the method.

The frequency matters. Every major provider now releases new model versions several times a year and retires older ones on a rolling basis. An application that lives three years will migrate models several times. That is a process to design, not a surprise to absorb.

What a model deprecation actually changes

An LLM version change is not like a library upgrade. The replacement model may be better on average and worse on your specific tasks. It may format output differently, follow instructions more or less literally, refuse different things, use a different tokeniser so your cost and context limits shift, and respond with different latency. The only way to know is to measure, which is why the eval suite is the centre of the plan.

The deprecation playbook

StepWhenWhat is doneExit criterion
Notice loggedOn the provider's announcementRecord the shutdown date, affected versions and recommended replacement in the deprecation calendarA dated ticket with an owner
Candidate evaluatedWithin two weeks of noticeRun the full eval suite on the recommended replacement and one alternative, on the same casesScores compared per task against the last accepted run
Regressions fixedFollowing weeksAdjust prompts, output parsing, token budgets; re-run evals until at or above baselineNo task below baseline; cost and latency within budget
Shadow runBefore any traffic movesRoute a sample of live traffic to both models; compare outputs and outcomesNo unexplained divergence on sampled cases
Gradual switchWeeks before shutdownMove traffic by task and by percentage through the routing layer; watch dashboardsFull traffic on the new version with metrics stable
Fallback retiredAt shutdownRemove the old version from routing config; archive its eval resultsConfig has no reference to the retired version

Pin versions, or the deprecation will happen without you

Providers offer aliases that always resolve to the newest model and snapshots that stay fixed until retired. Calling the alias means your behaviour can change on the provider's schedule with no notice to you; calling a snapshot means it changes only when you decide, and the provider tells you when the snapshot will be withdrawn. Pinning is the difference between a planned migration and a Monday-morning mystery. It also creates the obligation to migrate, which is what the rest of the plan handles. OpenAI's deprecations page is the canonical example of the schedule you are pinning against; other providers publish equivalents.

Route through one layer

Every call in the application names a task, not a model. A routing layer maps each task to a pinned model version, an optional fallback and any task-specific parameters. Migration then means changing the mapping for one task at a time, which can be done for a percentage of traffic, reverted in seconds, and observed in the traces. The layer also turns a deprecation into an opportunity: if the replacement scores worse on a task than another provider's model or an open-weight model, the mapping can point there instead. Model-agnostic routing across OpenAI, Anthropic, Google and open-weight models is how we build, and Multi-model routing explains the mechanics.

Evaluate before you believe the release notes

The replacement will be described as better. On your tasks, it may be. Run the full golden set on the candidate and compare per task, not on the aggregate: a summarisation task might improve while an extraction task with strict formatting regresses. Include cost and latency per case, because a model that is more accurate and twice as slow may not fit a voice agent, and a model with a larger tokeniser vocabulary changes the bill. Score refusals too; a replacement that declines requests the old model handled is a regression users notice first. Model and vendor selection: a benchmark-first approach describes the comparison discipline.

Fix regressions in the prompt, not in the code

Most regressions on a version change are formatting and instruction-following differences: the new model adds preamble, wraps JSON in prose, interprets "brief" differently, or stops using a tool the old one used reliably. These are fixed in the prompt, with the change versioned and the suite re-run. Where output parsing is brittle, the migration is the time to move to structured outputs or function calling, which are more stable across versions; Structured outputs and function calling covers the approach. Only when a regression survives prompt work does the question of a different model arise.

Shadow, then shift gradually

With evals at baseline, route a sample of real traffic to both models and compare outputs and outcomes without showing the candidate's answers to users. Divergences that the eval set did not predict become new eval cases. Then move traffic by task and by percentage, watching quality signals, escalations, cost and latency in the dashboards, with the old version still mapped as the fallback. The shift should be complete well before the shutdown date, so the last weeks are quiet rather than a race.

Self-hosted models deprecate too

Open-weight models do not get switched off by a provider, but they do get superseded, their serving frameworks change, and the hardware they were sized for ages. A self-hosted deployment needs the same plan with a different trigger: the team decides when to migrate rather than the provider. Self-hosted LLMs: when running your own model beats an API covers the trade-off, and our private agentic AI service builds the routing layer to span both hosted and self-hosted models so a migration in either direction is the same procedure.

A worked example

A D2C brand's WhatsApp agent handled order queries and personalised recommendations on a pinned model version. The provider announced the version's retirement with a few months' notice. The eval suite, built from real conversations during launch, was run on the recommended replacement the same week: the recommendation task improved, but the order-status task regressed because the new model padded its structured replies with explanations that broke the template. Two prompt revisions and a move to structured output fixed it. A shadow run on sampled traffic found one divergence in Hinglish handling that became a new eval case. Traffic moved task by task over a fortnight, with the old version as fallback until shutdown. The personalisation and WhatsApp agent case study describes the system.

Team and timeline

A deprecation migration for a single application typically takes two to four weeks of part-time work spread over the notice period: an engineer runs evals and fixes prompts, a business owner reviews divergences, and whoever owns the routing config shifts traffic. It is covered under a Care Plan: Essential at $1,000 per month (₹68,000) with 10 hours is usually enough for one application on one provider; Standard at $2,500 (₹1,60,000) with 25 hours suits several applications or multiple providers; Enterprise at $5,250 (₹3,40,000) with 60 hours and a named engineer suits systems where a migration must be rehearsed and signed off. Applications without an eval suite or routing layer need those built first, which is a short project under LLM applications. Plans are listed on the pricing page.

Before you start: a checklist

  • List every model version the application calls and confirm each is pinned, not an alias
  • Subscribe to deprecation notices from every provider and keep a dated calendar
  • Confirm the eval suite covers every task and records cost and latency per case
  • Make sure all model calls go through one routing layer with per-task mappings
  • Decide who owns a migration and how many hours it is allowed to take
  • Move brittle output parsing to structured outputs before the next migration
  • Plan a shadow run and a gradual shift as standard steps, not optional ones
  • Archive eval results per version so the next comparison has a baseline

Questions clients ask

  • How much notice do providers give? Usually months, published on a deprecations page; the plan assumes the notice is used from day one rather than the last fortnight.
  • Can we just switch to the alias and stop worrying? You would trade scheduled migrations for unscheduled behaviour changes; pinning is the safer of the two obligations.
  • What if the replacement is worse on our task? Route that task to a different provider or an open-weight model; the routing layer exists so this is a config change.
  • Does fine-tuning change the picture? A fine-tuned model is tied to a base version and must be re-tuned on the replacement; budget for it in the plan.
  • Should the migration be announced to users? Only if behaviour changes visibly; done well, it should not.

Why AI systems degrade after launch puts deprecation alongside the other causes of drift, and LLMOps for small teams sets out the tooling that makes a migration routine. Anthropic's model deprecation documentation is a second primary source for how providers announce and schedule retirements.

Pin, evaluate, route and shift gradually, and a provider's retirement notice becomes a fortnight of routine work rather than an outage.

Frequently asked questions

What is a model deprecation plan?

▾

A standing procedure for when a provider retires a model version: pinned versions and a deprecation calendar, an eval suite to score the replacement per task, prompt fixes for regressions, a shadow run on real traffic, and a gradual switch through a routing layer with the old version as fallback.

How long does an LLM version change take?

▾

Typically two to four weeks of part-time work spread across the notice period for one application, most of it evaluation and prompt adjustment. Applications without an eval suite or routing layer need those built first.

Who handles model deprecations after a project ends?

▾

Whoever holds the maintenance responsibility. Eazyware Care Plans from $1,000 per month include deprecation tracking and migrations, with hours scaled by plan; the pricing page lists them.