The real cost of running an LLM in production
What does it actually cost to run an LLM in production?
Tokens are one line of seven. A production LLM run-rate also carries retrieval infrastructure, evaluation runs, observability, guardrail and safety calls, re-indexing, and the engineering time to keep all of it correct as models change underneath you.
Token charges are one line of seven. A production LLM run-rate also carries retrieval infrastructure, evaluation runs on every prompt and model change, observability storage, guardrail and classification calls, re-indexing of your content, and the engineering time to keep the system correct when a provider retires a model. Budget the token line alone and you will underestimate the bill.
This is a breakdown of the whole run-rate rather than a token forecast. It covers what sits on each line, what drives it up, which levers genuinely reduce it, what Eazyware charges to run and maintain these systems, and the situation where optimising cost is the wrong thing to be doing.
The seven lines in a production LLM run-rate
1. Inference tokens
Input and output tokens charged by your model provider, paid through your own accounts. Output tokens generally cost more than input tokens, and the input side grows quietly: every retrieved passage, every system prompt revision and every conversation turn you carry forward is billed again on the next call. Definitions sit in the inference cost glossary entry.
2. Retrieval infrastructure
Embedding generation, a vector index, and the database or search cluster that holds it. This line is charged by stored vectors and query volume rather than by conversation, so it scales with your corpus, not your users. A large document estate with few users can make retrieval the dominant infrastructure cost.
3. Evaluation runs
Every prompt change, model change and retrieval change should be re-run against a golden question set before it ships. That is real inference spend that produces no user-facing output, and teams that skip it are not saving money, they are deferring a regression. The practice is covered in evals: the practice that separates AI demos from AI products.
4. Observability and tracing
Storing the prompt, the retrieved context, the response, the latency and the cost for every request, with enough retention to investigate a complaint weeks later. Trace payloads are large because they contain the full context window. The mechanics are in LLM observability: tracing every request from prompt to cost.
5. Guardrails and classification
Input classification, personal-data redaction, output checking and routing decisions are frequently themselves model calls, usually to a smaller model. They add calls per user request rather than per conversation, and they are easy to leave out of a forecast because they are invisible in the product.
6. Re-indexing and content maintenance
Documents change, so embeddings have to be regenerated. A knowledge base that turns over often carries a recurring embedding cost and an operational job that has to run, be monitored and be corrected when a source system changes format.
7. Engineering time
The line nobody puts in the spreadsheet. Model deprecations, provider rate-limit changes, prompt regressions, cost spikes caused by a longer context, and dependency updates all consume engineer hours whether or not anyone budgeted them. This is the line a support contract converts from unpredictable to fixed.
What drives each line, and what controls it
| Cost line | What drives it up | The lever that works |
|---|---|---|
| Inference tokens | Long system prompts, large retrieved context, chatty outputs | Route simple calls to a smaller model; cap retrieved passages |
| Retrieval infrastructure | Corpus size, vector dimensions, replica count | Chunk properly; drop stale content; right-size the index |
| Evaluation runs | Frequent releases, very large golden sets | Tiered evals: a fast subset per change, the full suite per release |
| Observability | Full context stored on every request, long retention | Sample verbose traces; keep summaries at full rate |
| Guardrails | A model call per check, several checks per request | Use a small model, and skip checks a rule can do |
| Re-indexing | Whole-corpus rebuilds on a schedule | Incremental re-indexing keyed on document change |
| Engineering time | Model deprecations, silent quality drift | A support plan with a named cadence |
Why the token line is the easiest to forecast and the hardest to hold
Token cost is arithmetic: calls per day, tokens per call, price per token. The LLM inference cost calculator will do it for you in a few minutes, and the forecasting method is set out in LLM inference costs: how to forecast your monthly bill.
What makes it hard to hold is that the inputs move without anyone deciding they should. Someone adds three sentences to the system prompt and every request in the system gets more expensive. A retrieval setting changes from five passages to ten and input tokens double. A new feature carries conversation history forward and cost per turn starts growing with the length of the conversation. None of these arrive as a budget decision; they arrive as a pull request.
The control that works is attribution. If every request is traced with its cost and tagged by feature, tenant and model, a cost change has an owner within a day. Without attribution, you discover it on the invoice a month later and cannot say what caused it.
Does self-hosting make it cheaper?
Self-hosting changes the shape of the bill more reliably than it changes the size. You stop paying per token and start paying per GPU hour, which means cost becomes fixed and utilisation becomes the thing that decides whether you win. At steady high volume, a well-utilised server is cheaper. At spiky or low volume, you are paying for idle silicon.
Throughput on that hardware is an engineering problem rather than a purchasing one. The vLLM documentation describes continuous batching and paged attention as the mechanisms that raise serving throughput on a given GPU, which is exactly the dimension that determines whether self-hosting pays. The commercial decision is worked through in self-hosted LLMs: when running your own model beats an API.
Data residency often decides this before economics does. If a regulator or a client contract requires processing inside your own infrastructure, self-hosting is a requirement and its cost is the price of being allowed to operate, not an optimisation.
The levers that actually reduce the bill
- Route by task, not by habit. Classification, extraction and routing rarely need your most capable model. Multi-model routing is usually the largest single saving available.
- Cap the context. Retrieve fewer, better passages. Reranking a small set beats stuffing a large one, and it improves answers as well as cost.
- Cache aggressively. Repeated questions are common in support and internal knowledge tools, and semantic caching serves them without a model call.
- Trim the system prompt. It is charged on every single request. A prompt that has grown by accretion over six months is a recurring tax.
- Sample your traces. Keep cost and latency for everything, keep full payloads for a sample and for every error.
- Re-index incrementally. Rebuild what changed, not the corpus.
- Set budget alerts per feature. An alert that fires the same day is worth more than a dashboard nobody opens.
What Eazyware charges to run and maintain this
Model and infrastructure usage is paid through your own provider accounts, so you see the raw bill and we do not mark it up. What we charge for is keeping the system correct. Care Plans run at $1,000 or ₹68,000 a month for Essential with business-hours IST cover and eight-hour response, $2,500 or ₹1,60,000 for Standard with 24x5 cover and four-hour response, and $5,250 or ₹3,40,000 for Enterprise with 24x7 cover, one-hour response and a named engineer. An AI system add-on at $750 or ₹40,000 a month covers evals, cost monitoring, prompt regression and re-indexing specifically.
On the build side, LLM application development starts at $21,000 or ₹13,60,000 and retrieval and knowledge engineering at $14,000 or ₹8,80,000, with every published figure on the pricing page. Indian clients are invoiced in INR with GST.
The shape of the bill differs by system
Which line dominates depends on what you built, and this is worth working out before you optimise anything. An internal knowledge assistant over a large document estate with a few hundred users is usually retrieval-heavy: the index and the re-indexing job cost more than the conversations. A customer-facing support agent at volume is token-heavy, because every request carries retrieved context and several guardrail calls. An agent that completes multi-step tasks is heaviest on tokens per completed task rather than per message, because a single task may involve a dozen model calls and tool round-trips.
The in-app copilot we built for a field-service SaaS company sat in that third category, where the metric that mattered was cost per completed task rather than cost per message. The system is described in the in-app copilot case study. Measuring the wrong unit is how teams conclude that an agent is expensive when what they have actually measured is that it does more per interaction.
When cost optimisation is the wrong priority
Optimising the run-rate before the system is correct is the most common waste we see. A cheap system that gives wrong answers has a cost of zero value, and routing to a smaller model to save money before you have an evaluation suite means you cannot tell what the saving cost you in quality.
The order that works is: make it correct, measure it, then make it cheap. Cost work is also the wrong priority when the total run-rate is small relative to the salary of the person doing the optimisation. If the monthly bill is a few hundred dollars, an engineer spending two weeks on caching has spent more than they will save in a year. Revisit at volume, not at launch.
Related reading
Cutting inference costs by a third: routing, caching, batching goes deeper on the engineering levers, handling model deprecations covers the line item nobody forecasts, and ongoing cover is described under software maintenance and support.
Forecast all seven lines, attribute every request, and treat the token bill as the symptom rather than the system.
Frequently asked questions
What is the biggest hidden cost of running an LLM in production?
▾
Engineering time. Model deprecations, prompt regressions, retrieval drift and provider changes all consume hours that rarely appear in a forecast. The second largest is evaluation inference: re-running a golden question set on every change produces no user-facing output but is what keeps quality measurable after each release.
Is self-hosting an LLM cheaper than using an API?
▾
It changes the bill from per-token to per-GPU-hour, so it wins at steady high volume with good utilisation and loses at spiky or low volume. Serving throughput decides the outcome, which is an engineering problem. Data residency requirements often make self-hosting mandatory regardless of the cost comparison.
How do I forecast my monthly LLM bill before building?
▾
Estimate calls per day, input and output tokens per call including retrieved context, then add retrieval infrastructure, evaluation runs, guardrail calls and observability storage. Eazyware's LLM inference cost calculator covers the token side; add roughly the same discipline to the other six lines before presenting a number.