LLM observability: tracing every request from prompt to cost
What is LLM observability and what should every request trace contain?
Observability traces each request with prompt version, model, latency and token cost so any bad answer can be located in minutes. It makes evals, routing and cost control auditable, and it is a few days of work with tools such as Langfuse or OpenTelemetry. Here is what to record and where teams go wrong.
LLM observability is the practice of recording, for every request an LLM application handles, enough to reproduce it: the prompt version, the model and provider, the rendered inputs, any retrieved context, each tool call, the raw and parsed output, tokens in and out, latency per step, cost, and whatever feedback the user gave afterwards. With that record, "the assistant told a customer the wrong renewal date on Tuesday" becomes a search that returns the exact trace in a minute. Without it, the same complaint becomes an afternoon of guessing. This article sets out what we record on every build under the LLM applications service, the tooling that makes it cheap, and how the traces feed everything else.
Why LLM monitoring is different from application monitoring
Ordinary monitoring answers "is the service up and fast?" LLM monitoring must also answer "is the service right?", and rightness is not a status code. A request can return 200 in two seconds with an answer that is confidently wrong. The only way to know is to keep the inputs and outputs, attach human and automated judgements, and look at the distribution over time. That makes LLM tracing closer to product analytics than to uptime monitoring: the unit is the request and its content, not the server and its CPU.
A second difference is that the interesting failures are multi-step. A wrong answer may come from a retrieval step that returned a stale passage, a routing decision that sent a hard case to a small model, a tool call that returned an error the model ignored, or a prompt version that changed yesterday. The trace has to capture the whole chain as one tree, not as unrelated log lines.
What a request trace should contain
| Field | Why it matters | Common omission |
|---|---|---|
| Prompt version and template ID | Ties the answer to the exact prompt; enables reproduction and rollback | Logging the rendered text but not the version |
| Model, provider and parameters | Explains behaviour changes across routing and provider updates | Recording the logical name but not the snapshot |
| Rendered inputs (redacted per policy) | Reproduction; eval-set growth | Redacting so aggressively that the case cannot be reproduced |
| Retrieved context with source IDs and versions | Distinguishes retrieval failures from generation failures | Storing only the answer |
| Tool calls, arguments and results | Shows what the model asked for and what it got back | Logging the call but not the result |
| Raw and parsed output, validation result | Catches schema failures and retries | Keeping only the final parsed object |
| Tokens in and out, latency per step, cost | Attributes spend and slowness to a step, feature and tenant | Total latency only |
| User, tenant, session and feature | Enables per-tenant metering and per-feature analysis | Anonymous traces that cannot be grouped |
| Feedback and review outcome | Turns complaints into eval cases | Feedback stored in a separate system with no trace ID |
Tooling: Langfuse, OpenTelemetry and your own stack
Open-source tools have made this a few days of work. Langfuse records nested traces, links them to prompt versions it also manages, attaches scores from automated and human evaluation, and reports cost per model and per user; it can be self-hosted, which matters for clients whose traces contain regulated data. The OpenTelemetry semantic conventions for generative AI define standard attribute names for model, tokens and prompts, so teams already running OpenTelemetry can emit LLM spans into the same backend as their application traces and correlate a slow answer with a slow database query in one view.
A team with an APM already in place usually adds LLM spans through OpenTelemetry and uses an LLM-specific tool for the eval and prompt side; a team starting fresh often runs Langfuse alone. Either way, emit traces through a thin wrapper so the backend can change without touching the feature.
What observability makes possible
Locating a bad answer in minutes
Search by user, time and feature; open the trace; read the prompt version, the retrieved passages and the output. In most cases the cause is visible immediately: the passage was from an old policy version, the extraction step returned a null the composer ignored, or the request hit the fallback model during a provider incident. The fix is targeted and the case goes into the golden set.
Growing the eval set from production
Every thumbs-down, every reviewer correction and every schema-validation failure is a trace with a known problem. Sampling these into the golden set, with the correct output added by a reviewer, is how the eval stays representative of real traffic rather than of the cases the team imagined at the start. This is the loop described in prompt versioning and evaluation.
Attributing cost and latency
With tokens and cost on every step, the expensive path is a query: which feature, which prompt, which tenant, which step. Usually one step is responsible for most of the spend, and it is a candidate for a smaller model or a shorter prompt, as set out in multi-model routing. Latency per step shows whether the wait is in retrieval, in the model or in a tool, which decides whether the fix is caching, streaming or a faster model.
Per-tenant metering
For a SaaS product, the same trace data becomes the meter: tokens and cost per tenant per feature per month, which is what billing an AI feature requires and what a customer's finance team will ask to see. The architecture is covered in multi-tenant LLM architecture.
Privacy, retention and who may look
Traces contain what users typed and what documents were retrieved, which may include personal data, customer records or regulated content. Decide before launch what is redacted at capture (payment details, government identifiers, health data), how long traces are kept, who may search them, and whether the backend must be self-hosted. Redaction should preserve enough to reproduce a case, which usually means masking identifiers rather than deleting whole fields. In India the DPDP Act shapes these choices; in regulated sectors the trace store is itself part of the audit trail and should be treated as such.
Where teams go wrong
- Logging only the final answer, so retrieval and tool failures are indistinguishable from generation failures
- Recording the model's logical name but not the snapshot, so a provider update is invisible
- Keeping feedback in a separate system with no trace ID, so complaints cannot be reproduced
- Redacting so much that a trace cannot be replayed, or so little that the store is a liability
- Building dashboards of averages that hide a category failing while the mean looks fine
- Treating observability as a launch-day add-on rather than part of the first week's skeleton
A worked example
A hospital network ran a multilingual voice agent for appointment booking. A department reported that some callers were being offered the wrong clinic. The team had call recordings and application logs, but the logs did not link the transcription, the intent classification, the slot lookup and the spoken response to one another, so each complaint took an engineer most of a day to reconstruct.
Tracing was added so that each call produced one tree: the speech-to-text output with its confidence, the language detected, the intent step with prompt version and model, the tool calls to the scheduling system with arguments and results, and the response text. Within a day of the first traces, the pattern was clear: for one language, a department name was being transcribed as a similar-sounding word, the intent step passed it through, and the lookup returned the nearest match rather than failing. The fix was a validation step with an explicit "did you mean" confirmation for low-confidence department names. The cases went into the golden set for that language. The build is described in the multilingual voice agent case study, and the same tracing underpins the voice agents we ship since.
Team and timeline
On a new build, tracing is part of the first week's skeleton and costs one AI engineer two or three days, plus a day for the redaction and retention policy. Retrofitting an existing feature is typically one engineer for one to two weeks, most of it instrumenting each step and agreeing the privacy rules. Dashboards for cost, latency and quality per category follow in the second week. It is included in every LLM applications build from $21,000 / ₹13.6L and in every agent build, and a Care Plan keeps the alerts and the eval loop running after launch. Current figures are on the pricing page; the traces, dashboards and tooling configuration belong to the client.
Before you start: a checklist
- List every step in the workflow and decide what each span records
- Log prompt version, model snapshot and parameters on every request
- Store retrieved context with source IDs and versions, and tool calls with results
- Record tokens, latency and cost per step, tagged with feature, tenant and user
- Attach feedback and review outcomes to the trace ID
- Write the redaction and retention policy and decide who may search traces
- Choose self-hosted or managed tooling based on where trace data may live
- Build per-category dashboards, not averages, and set cost and quality alerts
Glossary
- Trace: the full record of one request as a tree of spans, from input to final output.
- Span: one step within a trace, such as a retrieval, a model call or a tool execution.
- Score: a quality judgement attached to a trace by an automated check, a model judge or a person.
- Semantic conventions: OpenTelemetry's standard attribute names, including those for generative AI.
- Model snapshot: the specific dated version of a model behind a logical name.
Related reading
See what makes an LLM application production-ready for where observability sits among the six requirements, how to handle hallucinations for what traces should help you catch, and AI audit trails for the regulatory side of the same data.
Record the whole chain for every request, attach what people think of the result, and every other practice, evals, routing, cost control, becomes a query instead of an investigation.
Frequently asked questions
What should an LLM trace record?
▾
Prompt version, model snapshot and parameters, redacted inputs, retrieved context with source versions, tool calls and results, raw and parsed output with validation status, tokens, latency and cost per step, and the user, tenant, feature and any feedback, all linked under one trace ID.
Is Langfuse enough for LLM observability?
▾
For many teams, yes: it records nested traces, manages prompt versions, attaches evaluation scores and reports cost, and can be self-hosted. Teams with an existing OpenTelemetry stack often emit LLM spans there and use Langfuse or similar for the eval and prompt side.
How do we handle personal data in traces?
▾
Redact identifiers at capture while keeping enough to reproduce a case, set a retention period, restrict who can search, and self-host the store where regulation requires. We set this policy in the first week of every LLM application build.