azyware
Technology

What makes an LLM application production-ready?

EZ
Eazyware
· 7 min read
Quick answer

What makes an LLM application production-ready?

Production LLM apps have versioned prompts, evals, multi-model routing, structured outputs, observability and cost control from the first commit. A demo has none of these and still looks the same on screen, which is why so many pilots stall. Here is what each requirement means, how it is built and what it costs.

A production ready LLM application is one you can change without fear. That means a prompt edit runs against a golden set before it ships, a model outage fails over to another provider, a malformed response is rejected before it reaches your database, any bad answer can be traced to its prompt version and inputs in minutes, and the monthly bill is a number you forecast rather than discover. A demo has none of these and looks identical on screen. The six requirements below are what we build into every application under the LLM applications service, and they are cheaper to add on day one than on day ninety.

Why LLM app architecture differs from ordinary software

Ordinary code is deterministic: the same input gives the same output, so a passing test stays passed. An LLM is not. The same prompt can produce different outputs, a provider can change a model under the same name, and a small wording change can shift behaviour across thousands of unseen inputs. LLM engineering therefore replaces "does the test pass?" with "does the distribution of outputs stay acceptable?" and every practice below follows from that. The stance is the one we describe across our work: evals over demos, shadow mode before autonomy, and the client owning every prompt, model choice and eval set.

Demo versus production, requirement by requirement

RequirementWhat a demo hasWhat production has
PromptsStrings in the code, edited in placeVersioned, reviewed, tied to eval results, rolled back like code
EvaluationSomeone tries five questionsGolden set of 100–300 cases, scored on every change, release blocked on regression
Model accessOne provider, one model, hard-codedRouting layer across OpenAI, Anthropic, Google and open-weight; failover and per-task choice
OutputsProse parsed with regex and hopeSchema-validated structured outputs; invalid responses retried or rejected
ObservabilityConsole logsTrace per request: prompt version, model, inputs, tokens, latency, cost, feedback
CostDiscovered on the invoicePer-feature budgets, caching, routing, alerts on anomalies

Requirement 1: versioned prompts

A prompt is code that happens to be in English. It gets a version, a change description, a review and a link to the eval run that justified it. Prompts live in the repository or a prompt registry, never in a database field someone edits in production. Every logged request records which prompt version produced it, so a complaint from Tuesday can be reproduced against Tuesday's prompt even if it changed on Wednesday. The practice is set out in prompt versioning and evaluation.

Requirement 2: evals that gate releases

An eval set is one to three hundred real inputs with expected outputs or grading rubrics, built from actual usage rather than imagined cases. Every prompt, model or retrieval change runs against it, and a drop in score blocks the release the way a failing unit test does. Scoring combines exact checks where the output is structured, rubric grading by a stronger model where it is prose, and human review for a sample. Without this, every change is a guess and every improvement is an anecdote; with it, a team can ship prompt changes daily. We cover the method in evals: the practice that separates AI demos from AI products.

Requirement 3: multi-model routing

Hard-coding one model couples your product to one vendor's pricing, outages and deprecation schedule. A routing layer puts a thin interface between the application and the providers so that each task can use the model that suits it: a small, fast model for classification and extraction, a frontier model for judgement and synthesis, an open-weight model where data must not leave the VPC. The same layer gives failover when a provider degrades and makes benchmarking a new model a configuration change rather than a rewrite. The cost side is covered in multi-model routing.

Requirement 4: structured outputs

Downstream code should never parse prose. Where the model's job is to produce data (a classification, an extracted record, a decision with a reason), the request specifies a schema and the provider's structured-output or function-calling mode enforces it. The application validates the response against the schema anyway, retries with the validation error on failure, and rejects after a bounded number of attempts. This turns a class of silent corruption into a visible, countable error. The OpenAI structured outputs guide and Anthropic's tool-use documentation describe the provider mechanisms; the validation layer is yours.

Requirement 5: observability

Every request produces a trace: prompt version, model and provider, the inputs (redacted as policy requires), retrieved context if any, the raw and parsed output, token counts, latency per step, cost, and any user feedback attached later. Traces are searchable, so "show me every answer from this prompt version that got a thumbs-down" is a query, not an investigation. Open-source tools such as Langfuse and the OpenTelemetry conventions for generative AI make this a few days of work rather than a project. This is the layer that makes the other five auditable; see LLM observability.

Requirement 6: cost control

Inference cost scales with usage in a way that most software costs do not, and a single verbose prompt or a retry loop can multiply the bill. Production applications set a budget per feature, cache repeated requests where the inputs are identical, route to smaller models where the eval shows no quality loss, cap tokens per request, and alert when spend per user or per day departs from the forecast. The trace data from requirement five is what makes this possible: cost is attributed to the prompt, feature and tenant, so the expensive path is visible. For a SaaS product this is also the basis of per-tenant metering, covered in multi-tenant LLM architecture.

What is deliberately not on the list

Choosing the best model is not on the list, because the answer changes quarterly and a routed architecture makes the choice cheap to revisit. Fine-tuning is not on the list, because most applications never need it once prompts, retrieval and evals are in place. A specific framework is not on the list; we use them where they help and write plain code where they do not. What is on the list is the set of practices that let the application survive the first six months of real users and provider changes.

A worked example

A B2B SaaS company had shipped an AI feature that summarised customer accounts for sales reps. It was popular for a month, then complaints started: summaries missed recent activity, occasionally invented a contact name, and the provider bill had tripled. Nobody could say which prompt was live, because it had been edited in the admin panel several times, and there were no traces to compare a good summary with a bad one.

The rebuild was not a new model. Prompts were moved into the repository with versions. A golden set of two hundred real accounts with reviewed summaries was built, and a groundedness check flagged any name not present in the source data. Output moved to a schema with typed fields rather than free prose. A routing layer sent the extraction step to a small model and the summary step to a frontier model, with failover. Traces went to an observability tool, and cost was attributed per tenant. The invented names stopped because the schema and the check caught them; the bill fell because most of the tokens had been in the extraction step; and the next prompt change shipped with an eval score attached. The in-app copilot case study shows the same set of practices in a field-service product.

Team and timeline

The six requirements add roughly two to three weeks to a first build and save far more afterwards. A typical production LLM application is an AI engineer, a backend engineer and a part-time architect over six to twelve weeks: weeks one and two set up the routing layer, prompt versioning, eval harness and tracing as the skeleton; the feature is then built inside that skeleton rather than retrofitted. The six-week Launch 6 program at $26,500–45,500 ships an MVP with all six in place; the LLM applications service starts at $21,000 / ₹13.6L for a scoped feature, and a Care Plan keeps evals and cost alerts running after launch. Current figures are on the pricing page. Clients own the prompts, eval sets, routing configuration and traces.

Before you start: a checklist

  • Put prompts in the repository with versions before writing the feature
  • Collect 100–300 real inputs and define how each output will be scored
  • Choose a routing layer and confirm at least two providers work behind it
  • Define a schema for every output that downstream code consumes
  • Set up tracing with prompt version, model, tokens, latency and cost per request
  • Set a monthly budget per feature and an alert threshold
  • Decide what is redacted from traces and who may see them
  • Agree that a regression on the eval set blocks a release

Questions clients ask

  • Can we add these to an existing feature? Yes; tracing and prompt versioning first, then the eval set from the traces, then routing and schemas.
  • Do we need all six for an internal tool? Evals, versioning and tracing at minimum; cost control matters as soon as usage is unbounded.
  • Which framework should we use? Whichever the team can debug; the six requirements are framework-independent.
  • How do we know the eval set is good enough? When a change that users notice is also a change the score notices; grow it from production traces.

See prompt versioning and evaluation, structured outputs and function calling and how to handle hallucinations in production for the detail behind three of the six requirements.

Build the skeleton first, then the feature inside it, and the application will still be changeable when the model, the provider and the users have all moved on.

Frequently asked questions

What is the difference between an LLM demo and a production LLM application?

▾

A demo has a prompt and a model. A production application adds versioned prompts, a golden eval set that gates releases, multi-model routing with failover, schema-validated outputs, per-request tracing and cost control. The screen looks the same; the ability to change safely does not.

How long does it take to make an LLM application production-ready?

▾

Two to three weeks of skeleton work at the start of a six-to-twelve-week build. Retrofitting later takes longer because prompts, outputs and costs have to be untangled from the feature. Our LLM applications work starts with the skeleton.

Do we need to fine-tune a model for production?

▾

Rarely. Once prompts are versioned and evaluated, retrieval is sound and outputs are structured, most applications reach their quality target without fine-tuning. Fine-tune only when the eval set shows a gap that prompting and retrieval cannot close.