azyware
Technology

Evals: the practice that separates AI demos from AI products

EZ
Eazyware
· 7 min read
Quick answer

What should you know about LLM evals before shipping an AI product?

An eval suite is a scored set of real inputs and expected outputs run on every change; it is the difference between a demo and a product. It replaces opinion with a number, catches regressions before users do, and lets you swap models and prompts safely. Here is how to build one and what it costs.

LLM evals are the practice of running a fixed, scored set of real inputs through your system every time anything changes, and refusing to ship when the score drops. A demo is judged by whether the founder liked the last five answers. A product is judged by whether it still answers the same three hundred questions correctly after you changed the prompt, upgraded the model or added a new tool. That distinction is the whole article.

We build evals into every LLM application we deliver, and we treat them the way a software team treats a test suite: they are not optional, they run in CI, and a red result blocks the release. This article explains what an eval suite contains, how to build one from scratch, which scoring methods actually work, and what it means for your team and timeline.

What an LLM eval suite is and why it matters

An eval suite has three parts: a dataset of inputs, an expected output or rubric for each, and a scorer that compares what the system produced with what was expected. Run it on version one, record the score, then run it on every subsequent version. If the score rises, ship. If it falls, find out why.

Language models are non-deterministic and their behaviour shifts with every change to a prompt, a retrieved document, a tool definition or a model version. Without a suite you cannot tell whether a change helped or whether a vendor's silent update broke something. Teams without evals discover regressions from customer complaints. Teams with evals discover them in a pull request.

Demo versus product: the difference in one table

QuestionDemoProduct with evals
How do you know it works?Someone tried it and liked itScored suite, tracked per version
What happens when a prompt changes?HopeSuite runs in CI; regression blocks merge
Can you swap models?Risky and manualRun the suite against the candidate; compare
How do you handle a bad answer in production?Patch the promptAdd the case to the suite, fix, verify no regression
Who decides quality?Loudest stakeholderThe score, agreed in advance
Cost of a changeUnknown until users complainKnown before release

Building the dataset: where the inputs come from

The dataset is the most valuable asset in the suite and the one most teams get wrong. Synthetic questions written by the engineer are nearly useless, because they reflect what the engineer imagines users will ask. Real inputs come from four places.

  • Production logs, once the system has been in shadow mode or beta for a few weeks; sample by intent so rare cases are represented
  • Support tickets and search logs from the existing product, which show what users struggle with before any AI exists
  • Subject-matter experts writing the twenty hardest cases they can think of, with the answers they would accept
  • Failures: every bad answer reported by a user becomes a new case, so the suite grows with the product

A useful first suite is one hundred to three hundred cases. Fewer and the score moves on noise; many more and you lose the ability to inspect failures by hand, which is where most of the learning happens. Tag each case with intent, difficulty and source, so you can report per slice rather than as one average that hides a failing category.

Expected outputs: exact, rubric or reference

How you express the expected output depends on the task. Classification, extraction and structured outputs have exact answers: the invoice total is a number, the ticket category is one of twelve labels, the function call has these arguments. Score these with exact match or field-level accuracy and nothing fancier.

Free-text tasks such as summaries, drafts and explanations do not have a single correct answer. For these you write a rubric: the summary must mention the deadline, must not invent a party that is not in the source, must be under 120 words. A rubric is scored by a model acting as a judge, checked against a sample of human ratings so you know the judge agrees with people most of the time.

Retrieval tasks have a reference passage. Score whether the right passage was retrieved (recall), how much irrelevant material came with it (precision) and whether the answer stayed inside what was retrieved (groundedness). We cover this in how to measure RAG quality.

Scoring methods that actually work

Deterministic checks first

Anything that can be checked by code should be. Is the output valid JSON? Does it match the schema? Is the number within range? Does the SQL parse and run read-only? Did the function call name a real tool? These checks are free, fast and never wrong, and they catch a surprising share of regressions on their own.

Model-as-judge, with a calibration set

For rubric scoring a second model reads the output and the rubric and returns a pass or fail with a reason. This works well when the rubric is specific and badly when it is vague ("is this helpful?"). Calibrate it: have two people rate fifty outputs, run the judge on the same fifty, and measure agreement. If the judge disagrees with people more than occasionally, tighten the rubric. Anthropic's guidance on building evals and OpenAI's evaluation documentation both describe this pattern; the mechanics are similar across providers.

Human review on a sample

No automated scorer replaces a person reading twenty outputs a week. Human review catches the failures nobody wrote a rubric for: tone drift, subtle policy violations, answers that are correct and useless. Keep it small, scheduled and owned by the product side.

Eval-driven development in practice

Eval-driven development means the suite runs before every merge, the same way unit tests do. An engineer changes a prompt, opens a pull request, CI runs the suite and posts the score per slice next to the previous score. If any slice dropped beyond the agreed tolerance, the review is about why. If everything held or improved, the change ships.

The same workflow makes model changes safe. When a provider releases a new version, or you want to test a cheaper model for a subtask, run the suite against the candidate and compare. The decision is a number, not a debate. This is what lets us route across OpenAI, Anthropic, Google and open-weight models per task; see multi-model routing and how we choose a model per task.

Where the suite lives and who owns it

The suite is code and data in your repository, alongside the prompts it tests, which are versioned like code; see prompt versioning and evaluation. Record cost and latency per case too, so a small accuracy gain that doubles cost is visible as the trade-off it is. Ownership is shared: engineering owns the runner and scorers, product owns the dataset and rubrics, and both agree the thresholds. A suite owned only by engineering drifts from what users care about; one owned only by product never gets run.

A worked example

A field-service SaaS company added an in-app copilot that drafted job summaries and suggested next actions from ticket history. Before launch, we built a suite from three hundred real tickets: for each, the summary the dispatcher had actually written and the action they had actually taken. Deterministic checks covered format and length; a calibrated judge scored whether the summary contained the facts a dispatcher needed; exact match scored the suggested action.

The first run showed strong summaries and weak action suggestions on one slice: tickets with a parts delay. Inspecting the failures showed the copilot did not see inventory status. Adding that context raised the slice without touching the others, and the suite proved it. Three model upgrades later the suite is still the release gate, and each upgrade was accepted or rejected on the score. The in-app copilot case study describes the product side.

Team and timeline

Building a first suite takes an AI engineer and a product owner about two weeks: one to assemble and tag the dataset, one to write scorers, calibrate the judge and wire it into CI. It is part of every LLM application build, which starts at $21,000 / ₹13.6L, and of every SaaS copilot. If you already have a system and want to know whether it is safe to ship, a three-week ProofRun builds the suite and reports the score per slice. Ongoing eval maintenance sits inside a Care Plan. Full figures are on the pricing page.

Before you start: a checklist

  • Collect 100–300 real inputs from logs, tickets and experts, tagged by intent
  • Decide per task whether the expected output is exact, rubric or reference
  • Write deterministic checks for format, schema and range before anything else
  • Calibrate any model judge against fifty human-rated outputs
  • Record cost and latency per case alongside accuracy
  • Wire the suite into CI so a regression blocks the merge
  • Agree thresholds per slice with product, in writing
  • Schedule a weekly human review of twenty sampled outputs

Glossary

  • Eval suite: a fixed dataset of inputs and expected outputs with scorers, run on every change
  • Slice: a subset of cases sharing an intent or attribute, scored separately
  • Rubric: a written list of what a free-text output must and must not contain
  • Model-as-judge: using a second model to score outputs against a rubric
  • Calibration: measuring how often the judge agrees with human raters
  • Groundedness: whether an answer stayed inside the retrieved or provided context

What makes an LLM application production-ready, LLM observability: tracing every request and AI proof of concept vs demo cover the surrounding practices. For a vendor-side reference, OpenAI's evals documentation describes the same structure.

If you can point to a score that went up on the last change and would have blocked the release had it gone down, you have a product; if you cannot, you have a demo.

Frequently asked questions

How many eval cases do we need?

▾

One hundred to three hundred real cases for a first suite, tagged by intent so scores can be reported per slice. Grow it from production failures; each bad answer becomes a new case.

Can a model judge its own outputs reliably?

▾

A second model can score outputs against a specific rubric, but only after calibration against human ratings. Vague rubrics produce unreliable judges; deterministic checks should cover everything code can verify.

Do evals slow down development?

▾

They speed it up. Without a suite every change needs manual testing and every regression reaches users. With one, a prompt or model change is judged by a number in CI within minutes.