The hidden costs of LLM application development that quotes leave out
What are the hidden costs of LLM application development?
The hidden costs of LLM application development are inference that scales with usage rather than seats, integration work behind each connector, the evaluation suite and its upkeep, migration when a model is deprecated, human review during rollout, and the monthly care that stops quality drifting.
The hidden costs of LLM application development are inference that scales with usage rather than seats, the integration work behind each connector, the evaluation suite and its upkeep, prompt and model migration when a vendor deprecates a model, human review during rollout, and the monthly care that keeps quality from drifting.
None of these are dishonest omissions. They are simply the lines that live after the build, in a different budget, owned by a different person. This article opens the ledger line by line, gives the figures we quote, and shows how to ask for a number that survives contact with production.
Why the quote and the invoice disagree
A build quote answers one question: what does it cost to make this thing exist. An LLM application, once it exists, keeps spending money in a way traditional software does not. A CRUD feature costs almost nothing per request. An LLM feature costs real money every time somebody uses it, and more money when they use it badly.
That single difference reshapes the budget. In classic software, roughly eighty per cent of lifetime spend is engineering and twenty per cent is hosting. In an LLM application with real volume, usage, evaluation and care can reach or exceed the original build cost within the first eighteen months. The framing is set out in total cost of ownership for AI systems.
The fix is not a bigger build budget. It is a second budget line, owned by whoever owns the product, reviewed monthly, with a named person who can turn a feature off when its cost per use stops making sense. Teams that set this up in week one rarely get a nasty invoice; teams that discover it in month four almost always do.
The cost lines a quote usually leaves out
| Cost line | Who usually owns it | When it first appears | Typical shape |
|---|---|---|---|
| Model API usage | Client, on their own accounts | First week of the pilot | Per million tokens, in and out priced differently |
| Retries, evals and regression runs | Engineering | During the build, then every release | Often 15 to 30 per cent on top of production tokens |
| Connector build and credentials | Client IT plus vendor | Integration phase | Per system, and it is the schedule risk |
| Golden set creation and upkeep | Client subject-matter expert | Before the first release | Days of expert time, then hours per month |
| Human review during rollout | Operations | Shadow mode onwards | Real staff hours, falling as confidence rises |
| Model deprecation and migration | Engineering | Six to eighteen months in | Re-run evals, re-tune prompts, re-measure cost |
| Observability and tracing stack | Engineering | Before any real traffic | A per-seat or per-event SaaS line, or self-hosted |
| Content and knowledge hygiene | Whoever owns the documents | Continuously | Stale documents look exactly like model failure |
Inference is a usage cost, not a licence
Output tokens dominate, not input
Most teams budget for the prompt and forget the answer. Output tokens are priced several times higher than input tokens across the major vendors, so a system that writes long replies costs far more than one that writes short ones at the same request volume. Capping response length and asking for structured output rather than prose is a cost control, not just an engineering preference. Vendor rate cards are published openly; OpenAI's pricing page lists per-million input and output rates per model.
Retrieval multiplies the prompt
A retrieval-augmented application sends the question plus six to twenty retrieved passages. The retrieved context is often ten times the size of the user's question, and it is charged every single turn. Halving chunk size or reranking down to the best four passages can cut the bill by a third without touching answer quality. Anthropic's prompt caching documentation explains how a stable prefix can be cached and re-read at a fraction of the input rate, which matters most when the system prompt is long.
The invisible traffic
Evaluation runs, retries after a malformed response, guardrail classifier calls and background summarisation all consume tokens and none of them appear in a user-facing metric. Our rule of thumb is to add a quarter on top of the modelled production figure and then measure the real ratio in week two. You can model the headline number with the LLM inference cost calculator and the method is in LLM inference costs: how to forecast your monthly bill.
The five-figure lines people forget
- Golden set creation. Two hundred questions with known correct answers takes a subject-matter expert several days. Without it you cannot tell a regression from a bad day, so this is the cheapest insurance in the whole programme.
- Credential and environment access. Every connector needs a service account, a test environment and an owner. In larger organisations this is measured in weeks of waiting, and waiting is billed as calendar time even when nobody is coding.
- Shadow mode staffing. Running the system alongside the humans for four to six weeks means someone accepts or corrects its proposals. That time is real and it belongs in the business case, not in the engineering line.
- Model deprecation. Vendors retire models. When yours goes, you re-run evals, re-tune prompts and re-measure cost, which is typically one to two engineer-weeks per application.
- Document hygiene. An assistant answering from a knowledge base inherits every contradiction in it. Somebody has to own the content, and if nobody does, the application will be blamed for the library's faults.
- Security review cycles. Penetration tests, questionnaires and customer security reviews arrive after launch and consume engineering attention you have already spent in your head.
What does an LLM application cost to run each month?
The honest answer is that usage is yours and support is priced. At Eazyware you pay model providers directly through your own accounts, so there is no margin on tokens; we set budgets, routing and dashboards so the number stays predictable. Support sits in a Care Plan: Essential at $1,000 or ₹68,000 a month with business-hours cover in IST, Standard at $2,500 or ₹1,60,000 with 24x5 cover and a four-hour response, and Enterprise at $5,250 or ₹3,40,000 with 24x7 cover, a one-hour response and a named engineer.
For AI systems specifically, add the $750 or ₹40,000 AI add-on, which covers evals, cost monitoring, prompt regression and re-indexing. That add-on is the line that prevents the slow decline nobody notices until a customer complains. What a plan should and should not include is spelled out in what a care plan should cost, and the plans themselves are on the maintenance and support page.
The build itself is separate. LLM application development starts at $21,000 or ₹13,60,000 and runs to $84,000 or ₹56,00,000 depending on integration count and evaluation depth, with all starting figures listed on the pricing page.
A worked budget for a mid-sized build
Take an internal knowledge assistant for four hundred staff, built over ten weeks with four integrations. The build lands in the low forties of thousands of dollars. Model usage in month one, at perhaps eight thousand questions with retrieval, is a low four-figure monthly number, and evaluation traffic adds a quarter on top of that. A Standard Care Plan with the AI add-on is $3,250 or ₹2,00,000 a month.
Two lines usually surprise the sponsor. The first is the subject-matter expert who spends four days building the golden set and two hours a month maintaining it. The second is the operations lead who reviews escalations weekly for the first quarter. Neither appears in any vendor quote, both are unavoidable, and both are cheap next to the cost of an assistant nobody trusts.
There is a third line that only shows up in year two. Usage grows because the application is good, and the cost per question stays flat unless somebody works on it. A knowledge assistant that doubles its adoption doubles its bill on the same architecture. Routing the easy two thirds of questions to a smaller model, caching the stable system prompt and reranking retrieval down to four passages typically recovers a third of that increase, but only if an engineer is given the time to do it. Put that work in the roadmap before the finance team finds it first.
When paying these costs is the wrong choice
If your expected volume is a few hundred requests a month, the running costs are trivial and the build cost is the whole story. Do not over-engineer routing, caching or a tracing stack for traffic that will never justify them; a simple application with a single model and a weekly manual check is the correct answer at that scale.
If the workflow is high volume but genuinely deterministic, an LLM is the wrong tool and every one of these running costs is avoidable. Rules, a search index or a small classifier will be cheaper, faster and easier to audit. We say no to this regularly. The clearest signal is that your subject-matter expert can write the decision table in an afternoon; if they can, write the table.
How to make a vendor quote the real number
- Ask for expected monthly token spend at your stated volume, with the input and output split shown separately
- Ask what percentage of tokens will be evaluation and retry traffic
- Ask which model tiers each step routes to, and what happens when the cheap tier is not good enough
- Ask who pays for usage and whose accounts the keys live in
- Ask for the cost of a model migration, in engineer-days, before you sign
- Ask what the care plan covers and what counts as a change request
Related reading
Cutting inference costs by a third covers routing, caching and batching in detail, multi-model routing explains how to send cheap work to cheap models without losing quality, and LLM observability shows how to attribute every rupee to a request.
Budget the build once and the running cost every month, because the second number is the one that decides whether the application survives its first year.
Frequently asked questions
What is the biggest hidden cost in an LLM application?
▾
Model usage that scales with requests rather than seats. Output tokens cost several times more than input tokens, and retrieval multiplies the prompt on every turn. Evaluation runs and retries typically add another fifteen to thirty per cent on top of production traffic, and none of that appears in a build quote.
Who pays for the model API usage?
▾
At Eazyware the client does, through their own provider accounts, so there is no vendor margin on tokens and no surprise mark-up. We configure budgets, routing rules and cost dashboards during the build so the monthly figure is predictable and attributable to specific features from the first week of the pilot.
How much should I budget for ongoing support?
▾
Eazyware Care Plans run from $1,000 or ₹68,000 a month for business-hours cover to $5,250 or ₹3,40,000 for 24x7 cover with a named engineer. For an LLM application add the $750 or ₹40,000 AI add-on, which covers evals, cost monitoring, prompt regression and re-indexing.