azyware
Technology

How to measure whether RAG development services is working

EZ
Eazyware
· 7 min read
Quick answer

How do you measure RAG development services?

Measure four layers: retrieval, generation, product and business. Retrieval tells you whether the right passage was found, generation whether the answer stayed inside it, product whether anyone uses the system, and business whether a number your CFO tracks moved. A programme failing at one layer is failing.

Measure four layers and report them together: retrieval, whether the right passage was found; generation, whether the answer stayed inside that passage; product, whether people actually use the system; and business, whether a number your finance team already tracks has moved. RAG development services metrics that stop at the first layer will flatter a system nobody trusts.

This article sets out what to instrument at each layer, when each becomes meaningful, the numbers that consistently mislead buyers, and how to turn the whole thing into a release gate rather than a slide.

Why one number is never enough

Buyers usually arrive with a single question: is it accurate? The word covers at least four distinct properties, and the layers below exist to pull them apart. Treating them as one number is how a system reaches launch with excellent internal scores and a user base that has quietly gone back to asking a colleague.

A retrieval-augmented system can fail in three independent places. It can fetch the wrong passages. It can fetch the right passages and answer from memory anyway. Or it can do both correctly and still be ignored because the answer arrives eight seconds late without a citation anyone can click. Each failure has a different owner and a different fix, so a single accuracy percentage hides the only information you need.

The metric definitions themselves are covered in how to measure RAG quality: recall, precision and groundedness. This post is about the layer above that: turning those definitions into a programme scorecard with owners, cadences and thresholds, so you can tell in month three whether the engagement is working.

An evaluation suite is a fixed set of questions with known correct answers and known correct sources, run automatically on every change to prompts, chunking, embeddings or model. It is the artefact that separates a product from a demonstration, a stance we set out in evals over demos.

The four-layer scorecard

LayerWhat it answersMetricMeasured byUseful from
RetrievalDid we find the passage that contains the answer?Recall at k, mean reciprocal rankGolden question set with labelled source passagesWeek two
GenerationDid the answer stay inside what we retrieved?Groundedness, citation accuracy, refusal rateAutomated judging plus weekly human review of a sampleWeek four
ProductDo people come back and act on the answer?Weekly active askers, repeat rate, click-through on citations, escalation rateApplication analyticsWeek two after launch
BusinessDid a number the company already tracks move?Handling time, ticket deflection, hours saved per role, first-response timeYour existing operational reportingMonth three
Cost and reliabilityIs it affordable and up?Cost per answered question, p95 latency, index freshness lagTracing and billing dashboardsWeek one

What to build before the system: the golden question set

The first deliverable of a serious retrieval engagement is not a pipeline; it is a golden question set. One hundred to three hundred real questions, taken from support tickets, search logs and the inbox of the person everyone asks. Each carries the correct answer and the document that contains it. Building it is unglamorous and the client has to do most of it, because only your people know what correct looks like.

  • Draw from real traffic. Invented questions are always easier than the ones users ask, so scores on them are always higher.
  • Include the questions that should be refused. A system that never says it does not know will confidently answer about a policy that was withdrawn in 2023.
  • Cover every document type. If half your corpus is scanned PDFs, half the set should test them.
  • Label the source, not just the answer. Without a labelled passage you cannot separate a retrieval failure from a generation failure.
  • Include near-duplicates. Two versions of the same policy is the most common real-world trap.
  • Version it like code. Questions get added as the corpus grows; scores across versions are only comparable if you say which version ran.
  • Keep a held-out slice. Tuning against every question you own produces a system that is excellent at your test and mediocre at work.

Which numbers mislead buyers

Three numbers appear in almost every vendor demonstration and none of them survive contact with production.

Answer accuracy with no source labelling

If a reviewer marks answers right or wrong without checking which document they came from, a system that guesses well scores the same as one that retrieves well. The two behave completely differently the moment your corpus changes.

Benchmark scores on public datasets

A vendor's score on a public question-answering benchmark tells you about that benchmark. Your corpus has your acronyms, your document structure and your contradictions. Ask for a score on fifty of your own questions instead, which is exactly what a three-week ProofRun produces.

Volume of questions asked

Usage rises for six weeks after any launch because people are curious. The number that matters is the repeat rate in week eight among people who used it in week two. If that falls below a third, the answers are not good enough, and no amount of internal marketing or training will fix it. Cohort the usage data from the first day rather than reporting a single rising line, because the rising line is the one number that is always available and almost never informative.

How the scorecard becomes a release gate

Evaluation only changes behaviour when it can block something. We run the suite in continuous integration on every prompt, chunking or model change, and a pull request that drops groundedness below the agreed threshold does not merge. The open-source Ragas framework, whose documentation defines faithfulness and context precision as separate measures, is a reasonable starting point for the automated half; the human-reviewed sample is what keeps the automated judge honest.

Model deprecations make this non-negotiable. Providers retire model versions on their own schedule, and a system tuned against a version that no longer exists will quietly change behaviour overnight. The same suite that gates your releases is what turns a forced model migration into a two-day task. Tracing every request from prompt to cost, covered in LLM observability, gives you the other half of the picture: what each answer cost and where the latency went.

When measurement is the wrong first move

There is a stage at which building an evaluation suite is premature. If you have not decided which ten questions the system exists to answer, a three-hundred-question golden set is procrastination dressed as rigour. Start with ten, get them right, and grow the set as the corpus grows.

Measurement is also wrong when it is used to litigate rather than to improve. A scorecard whose purpose is to prove a vendor wrong produces a vendor who optimises the scorecard. Agree the thresholds before the build starts, agree who reviews the weekly sample, and treat a failed gate as information rather than as a breach.

What the numbers cost to produce

Evaluation is roughly a fifth of a retrieval engagement, not a line item you can drop. Retrieval and knowledge engineering starts at $14,000 or ₹8.8 lakh and runs to $49,000 or ₹32 lakh, and the golden set, the automated suite and the weekly review are inside that. Starting prices for every programme are on the pricing page.

After launch, someone has to keep running it. A care plan starts at $1,000 or ₹68,000 a month, and the AI add-on at $750 or ₹40,000 covers evals, cost monitoring, prompt regression and re-indexing. Ongoing maintenance and support is where index freshness and model migrations actually get handled; most clients stay on for six to twelve months.

A worked example

A field-service SaaS company had a copilot its users ignored. Query logs showed the retrieval layer was finding the right documentation most of the time, so the instinct was to tune retrieval further. The product layer told a different story: people clicked a citation, landed on a page that did not obviously contain the answer, and stopped asking. The fix was citation granularity and answer formatting, not embeddings. The full engagement is written up in the in-app copilot case study, and the lesson generalises: instrument the product layer or you will spend a quarter optimising the layer you can already see.

A measurement checklist

  • Agree the ten questions that must never be answered wrongly, before any code is written
  • Label source passages, not just correct answers, in the golden set
  • Set thresholds for recall, groundedness and p95 latency, and write them into the contract
  • Wire the suite into continuous integration so a regression blocks a merge
  • Review a random sample of twenty real answers weekly, with a named human owner
  • Track cost per answered question alongside quality, not separately
  • Define the one business number this system is supposed to move, and who reports it
  • Re-run the whole suite on every model change, planned or forced

Evals: the practice that separates AI demos from AI products covers the discipline in general, why basic RAG fails in production lists the failure modes the scorecard is designed to catch, and prompt versioning and evaluation explains how to keep prompts under the same control as code.

A retrieval programme is working when a threshold you agreed in advance is being met on questions your own people wrote, and it is not working the moment anyone proposes lowering the threshold.

Frequently asked questions

What is a good recall score for a RAG system?

▾

Recall at five above 0.9 on a golden question set drawn from real traffic is a reasonable production target for most corpora. The number matters less than its stability: a score that swings five points between runs usually means the question set is too small or the corpus contains contradictory documents.

How many questions should a RAG evaluation suite contain?

▾

Start with ten questions you cannot afford to get wrong, grow to one hundred before launch, and three hundred within a few months of live traffic. Keep a held-out slice you never tune against, and version the set so scores from different months remain comparable.

How soon can we expect business results from a RAG project?

▾

Retrieval and cost metrics are readable within two weeks, product metrics within a fortnight of launch, and business metrics such as handling time or ticket deflection from month three. Anyone promising a measurable business number in week four is measuring enthusiasm, not outcomes.