Multi-model routing: cutting LLM costs without cutting quality
How does multi-model routing reduce LLM costs without reducing quality?
Route simple steps to small models and judgement steps to frontier models; most systems cut inference cost sharply with no quality loss. LLM cost optimisation is an architecture choice, not a vendor negotiation: decompose the workflow, measure each step on a golden set, and send each to the cheapest model that passes.
LLM cost optimisation starts from one observation: most of the tokens in a production workflow are spent on steps that do not need a frontier model. Classifying an intent, extracting fields from a document, deciding whether a message is in scope, formatting a result: small, fast models do these as well as the largest ones, at a fraction of the price. Judgement steps, where the model must reason across ambiguous context and write for a person, are where the frontier models earn their cost. Multi-model routing is the practice of separating the two and sending each step to the cheapest model that passes the eval. This article sets out how we do it under the LLM applications service, and what it takes to do it without losing quality.
Why LLM inference cost behaves differently from other cloud spend
Inference cost scales with tokens, and tokens scale with usage, prompt length, retries and the number of steps in a workflow. A feature that costs a few cents per request in a pilot can cost a meaningful share of revenue at scale, and the bill arrives after the usage. Unlike compute, there is no reserved-instance discount that fixes it; the levers are architectural. The three biggest are routing (which model), caching (whether the model is called at all) and prompt discipline (how many tokens each call carries). Routing is usually the largest, because pricing differences between model tiers are wide, as the published price lists from OpenAI and Anthropic show for any given month.
LLM routing strategy: the four levels
| Level | How the route is decided | Typical saving | Risk if done badly |
|---|---|---|---|
| Static by step | Each workflow step is assigned a model at design time | Largest and most predictable | A step is misjudged as simple; caught by per-step evals |
| Rule-based | Input length, language, tenant tier or feature flag picks the model | Moderate | Rules drift from reality; review quarterly |
| Classifier-based | A small model scores difficulty and routes accordingly | Moderate to large | Classifier errors send hard cases to weak models; needs its own eval |
| Cascade | Try the small model; escalate when confidence or validation fails | Large on easy-heavy traffic | Latency doubles on escalation; bound the retries |
Most systems get most of the benefit from the first level. Decompose the workflow into steps, evaluate each step separately, and assign the cheapest model that meets the threshold. The other levels are refinements for mixed traffic where one step sees both trivial and hard inputs.
Step 1: decompose the workflow
A single "do everything" prompt to a frontier model is the most expensive and least testable architecture. Break it into steps with clear inputs and outputs: normalise the input, classify the intent, retrieve context, extract the structured facts, decide, compose the response, check the response against policy. Each step is now a candidate for a different model, and each has its own eval. This is also how structured outputs become possible, because a step that extracts fields can be validated against a schema in a way that a paragraph cannot.
Step 2: evaluate each step, not the whole
Build a golden set per step. For the classification step, a few hundred labelled inputs. For extraction, documents with verified fields. For composition, a rubric. Run every candidate model against each step's set and record accuracy, latency and cost per thousand requests. The result is a table, not an opinion, and it usually shows that the small models tie the frontier models on the mechanical steps and lose on the judgement ones. The method is the same one described in prompt versioning and evaluation, applied per step.
Step 3: route through one interface
The application calls a routing layer, not a provider. The layer holds the mapping from step to model, handles authentication for each provider, normalises request and response formats, and records the model used on every trace. Behind it sit OpenAI, Anthropic, Google and open-weight models served from your own infrastructure where data residency requires it. This model-agnostic stance is one we take on every build: it lets a new model be benchmarked in an afternoon by changing a configuration entry, and it means a provider outage or price change is a routing change rather than an emergency. The comparison across providers per task is covered in how we choose per task.
Step 4: add failover and caps
The same layer that routes for cost also routes for resilience. If a provider returns errors or latency exceeds a threshold, the request falls over to a second model for that step, which the eval has already shown to be acceptable. Token caps per request and per step prevent a runaway prompt from multiplying the bill, and a bounded retry policy stops a validation loop from calling the model twenty times. Budgets per feature and per tenant, with alerts when spend departs from forecast, close the loop; this is where the observability data earns its keep.
The other two levers: caching and prompt discipline
Caching removes calls entirely. Exact-match caching returns the stored response for identical requests, which is common for classification of repeated inputs and for retrieval queries. Provider-side prompt caching reduces the cost of long, repeated system prompts and retrieved context across calls. Semantic caching, which returns a stored response for a sufficiently similar question, saves more but needs its own eval to be safe.
Prompt discipline is unglamorous and effective: trim the system prompt to what the step needs, retrieve fewer and better passages rather than more, stop including the whole conversation history when the last three turns suffice, and ask for concise outputs where the consumer is code rather than a person. A review of the longest prompts in the traces frequently finds a third of the tokens doing no work.
What quality loss looks like, and how to see it
Routing goes wrong in one way: a step is sent to a model that fails on inputs the eval did not contain. The defence is a golden set that reflects real traffic, per-category scores so a weak category is visible, and a sample of production outputs reviewed each week. Where a step is borderline, run the small model in shadow mode alongside the frontier model for a period and compare before switching; shadow mode before autonomy is the same principle we apply to agents. Any drop in a category score reverts that step to the previous model until the cause is understood.
A worked example
A customer-service platform ran every incoming message through a single frontier-model prompt that classified the intent, decided whether to answer or escalate, drafted the reply and checked it against policy. Quality was good. The monthly bill was growing faster than the customer base, and finance wanted a forecast.
The workflow was decomposed into five steps with a golden set for each. Intent classification and policy checking moved to a small model that tied the frontier model on the eval. Retrieval queries were cached. The reply draft stayed on a frontier model, because the eval showed a clear gap on nuanced complaints. The system prompt was cut by half after a review of the traces. A routing layer put a second provider behind each step for failover, and per-tenant cost attribution let finance model the bill per customer. The bill fell substantially while the category scores on the golden set stayed level, and the next quarter's forecast held. The customer service agent builds we ship use this decomposition as standard, and the KYC document-intelligence case study applies the same routing to document extraction, where small models handle most pages and a stronger model handles the ambiguous ones.
Team and timeline
Introducing routing to an existing feature is typically one AI engineer for two to four weeks: a week to decompose the workflow and build per-step golden sets from traces, a week to benchmark candidate models per step, and one to two weeks to put the routing layer, caching, caps and failover in place and run the borderline steps in shadow mode. On a new build the routing layer is part of the skeleton from the first week. The work sits within the LLM applications service, from $21,000 / ₹13.6L for a scoped feature; a Care Plan keeps the cost alerts and per-step evals maintained as models change. Current figures are on the pricing page, and the routing configuration, eval sets and benchmarks are owned by the client.
Before you start: a checklist
- Pull the last month of traces and attribute cost per feature and per step
- Decompose each workflow into steps with clear inputs and outputs
- Build a golden set per step from real traffic, with per-category scores
- Benchmark at least three models per step on accuracy, latency and cost
- Put a routing layer between the application and every provider
- Set token caps, bounded retries and a budget per feature and tenant
- Run borderline steps in shadow mode before switching models
- Review the longest prompts and remove tokens that do no work
Questions clients ask
- Will users notice a cheaper model? Not on steps where the eval shows a tie; the point of per-step evals is to know which steps those are.
- Can we route to open-weight models we host? Yes; the routing layer treats a self-hosted endpoint like any provider, which matters where data must stay in the VPC.
- Does routing add latency? The layer itself adds negligible overhead; cascades add latency on escalation, which is why we bound them.
- How often should the routing be revisited? Whenever a provider changes prices or releases a model, and at least quarterly; the benchmark takes an afternoon.
Related reading
See LLM inference costs: how to forecast your monthly bill for the budgeting side, what makes an LLM application production-ready for where routing sits among the six requirements, and self-hosted LLMs for when a private model belongs in the mix.
Decompose, measure each step, and let the numbers assign the model; the bill falls and the quality score stays where it was.
Frequently asked questions
What is multi-model routing?
▾
A layer between an application and LLM providers that sends each workflow step to the model best suited to it: small models for classification and extraction, frontier models for judgement and writing, self-hosted models where data must not leave the VPC, with failover between them.
How much can routing reduce LLM costs?
▾
It depends on how much of the workflow is mechanical. Systems where most tokens go to classification, extraction and formatting see the largest reductions; systems dominated by long-form reasoning see less. Per-step evals tell you before you commit.
Does routing to smaller models reduce quality?
▾
Not when each step is evaluated separately and the small model ties the frontier model on that step's golden set. Where the eval shows a gap, the step stays on the stronger model. Our LLM application builds run borderline steps in shadow mode first.