How to measure RAG quality: recall, precision and groundedness
How do you evaluate RAG quality, and which metrics tell you whether retrieval and answers are actually good?
Measure RAG with a golden question set: retrieval recall and precision, answer groundedness and citation accuracy, re-run on every change. The golden set is built from real user questions with verified answers and sources, and metrics are reported per question type so a failing category cannot hide in an average.
RAG evaluation is the practice of measuring a retrieval-augmented system against a fixed set of real questions with known answers, every time anything changes. It separates the retrieval question (did the right passages come back?) from the answer question (did the model use them faithfully and cite them?), because the two fail for different reasons and are fixed by different people. Without it, every improvement is a guess and every complaint is an anecdote. This article explains how to build the golden set, which metrics to compute, how to run them in a pipeline, and how to read the results so the next engineering week goes to the right problem.
Why RAG evaluation comes before tuning
Teams usually tune RAG by changing something (chunk size, embedding model, prompt) and asking a few questions to see if it feels better. It feels better for the questions they asked and worse for the ones they did not. A golden set replaces feel with a number, and the number tells you which stage to fix: low recall means parsing, chunking or search; low precision means too much noise reaching the model; low groundedness means the answer layer is improvising. The six production failures each show up as a distinct pattern in these metrics, which is what makes evaluation a diagnostic rather than a scorecard.
RAG metrics: what each one measures
| Metric | Question it answers | Computed from | Points to |
|---|---|---|---|
| Retrieval recall | Did the passage that answers the question appear in the retrieved set? | Retrieved chunks versus the golden source passage | Parsing, chunking, search method, filters |
| Retrieval precision | How much of the retrieved set was relevant? | Relevant chunks divided by retrieved chunks | Candidate count, re-ranking, metadata filters |
| Rank of the right passage | Was it first, or buried at position nine? | Position of the golden passage | Re-ranker quality, fusion weights |
| Answer correctness | Does the answer match the verified answer? | Judge model or human against the golden answer | Everything upstream, plus the prompt |
| Groundedness | Is every claim in the answer supported by the retrieved passages? | Claim-by-claim check against the context | Answer prompt, refusal behaviour |
| Citation accuracy | Do the citations point to the passages that support the claims? | Cited passage versus supporting passage | Citation logic, chunk boundaries |
| Refusal accuracy | When nothing relevant exists, did it say so? | Unanswerable questions in the set | Answer prompt, confidence threshold |
| Permission compliance | Did any restricted passage reach a user who may not see it? | Negative test cases | Entitlement filtering |
Building a golden set RAG teams can trust
Start from real questions
Take questions from support logs, search logs, chat transcripts, or a week of asking staff what they wanted to know. Invented questions test what the builders imagined; real ones test what users do. One to three hundred is enough to start, provided they are spread across the intents and document types the system covers.
Record the source passage, not just the answer
For each question, a person finds the passage in the corpus that answers it and records its location. This is the step teams skip, and it is the one that makes retrieval recall measurable at all. It is tedious and is best done by the people who know the documents: the support lead, the policy owner, the paralegal. A model can propose candidate passages to speed it up; a human confirms.
Include the hard cases
Questions with exact identifiers, questions phrased nothing like the document, multi-part questions, questions whose answer changed between document versions, questions in each supported language, and questions with no answer in the corpus at all. The last group measures refusal, and a system evaluated only on answerable questions will never learn to say "I could not find that".
Tag every question
Document type, intent, language, difficulty, whether it needs an exact match, whether it is multi-hop. Tags are what turn one average into a per-category report, and the per-category report is what tells you the contract questions are failing while the help-centre questions are fine.
Retrieval evaluation in practice
Run every golden question through the retrieval stage alone and record the top-k chunks. Recall at k is the share of questions whose golden passage appears in the top k; report it at k equals five and ten, because the model sees roughly that many. Precision is the share of retrieved chunks that are relevant, which needs relevance labels beyond the single golden passage; a judge model can label the rest with human spot-checks. Rank of the golden passage shows whether re-ranking is working. Run this on every change to parsing, chunking, embedding model, index configuration or search method, and keep the history so a regression is visible the day it happens. Hybrid search is usually the first change this evaluation justifies.
Groundedness evaluation: the answer layer
Groundedness asks whether each claim in the answer is supported by the retrieved context. It is computed by splitting the answer into claims and checking each against the passages, with a judge model doing the bulk and humans checking a sample of the judge's verdicts. A high groundedness score with a low correctness score means retrieval brought the wrong passages and the model faithfully summarised them; the fix is retrieval. A low groundedness score with high recall means the right passages were there and the model went beyond them; the fix is the answer prompt and the refusal threshold. Citation accuracy is checked the same way: does each citation point to the passage that actually supports the claim it is attached to. Frameworks such as Ragas implement these metrics and are a reasonable starting point, provided the judge prompts are checked against human labels for your domain.
Using judge models honestly
A judge model scoring another model's answers is convenient and can be wrong in consistent ways: lenient on fluent answers, harsh on terse correct ones, blind to domain errors. Calibrate it: have humans score a sample of one hundred, compare, and adjust the judge prompt until agreement is acceptable. Re-calibrate when the judge model changes, and report human-verified numbers in any external claim.
Running evaluation on every change
The golden set runs in the deployment pipeline. A change to a prompt, a model, a chunking rule or a connector triggers the full suite, and a regression beyond an agreed threshold blocks the release. Results go to a dashboard with history per metric per category. Production monitoring complements it: user feedback, reopened conversations and unanswerable-question logs feed new golden questions monthly, so the set tracks what users actually ask. The broader discipline is described in Evals: the practice that separates AI demos from AI products; RAG evaluation is that practice applied to retrieval.
Reading the results
- Recall low, everything else fine: fix parsing, chunking or search before touching the prompt
- Recall high, rank poor: add or tune the re-ranker
- Precision low: reduce candidates, tighten metadata filters, re-rank
- Groundedness low with recall high: tighten the answer prompt, raise the refusal threshold
- Correctness low with groundedness high: the retrieved passages are wrong or stale; check freshness and filters
- Refusal accuracy low: the system answers unanswerable questions; add a confidence gate
- One category failing while the average looks fine: that category's document type needs its own chunking or parsing
A worked example
A hospital network built an assistant for front-desk and call-centre staff over clinical scheduling policies, insurance procedures and department directories, feeding both a chat interface and a multilingual voice agent. The first golden set was two hundred real questions from call logs, tagged by department, language and whether they needed an exact match on a doctor's name or a procedure code. Retrieval evaluation showed recall on name and code questions well below policy questions, which hybrid search fixed. Groundedness evaluation then showed the answer layer inventing appointment durations when the policy did not state one; a stricter prompt and a refusal path fixed that. The set now runs on every change, grows monthly from unanswered questions, and the per-language report is what decides when a new language is ready to launch.
Team and timeline
Building the first golden set is one to two weeks of a retrieval-focused AI engineer working with a domain owner on your side who labels passages; the evaluation pipeline, judge calibration and dashboard take a further week. It is included in every retrieval and knowledge engineering build, from $14,000 / ₹8.8L, and is the main deliverable of a three-week ProofRun from $6,250 when a client wants to know how well retrieval works on their own corpus before committing. Under a Care Plan we run the suite on every change and grow the set monthly. Prices are on the pricing page.
Before you start: a checklist
- Collect 100–300 real questions from logs and transcripts
- Have a domain owner record the source passage for each
- Add unanswerable questions and negative permission cases
- Tag questions by type, intent, language and difficulty
- Decide the metrics and the regression threshold that blocks a release
- Calibrate the judge model against human scores on a sample
- Wire the suite into the deployment pipeline with a history dashboard
- Agree who adds new questions monthly and from which sources
Glossary
- Golden set: real questions with verified answers and source passages, used for every evaluation run
- Recall at k: share of questions whose golden passage appears in the top k retrieved chunks
- Precision: share of retrieved chunks that are relevant
- Groundedness: whether every claim in an answer is supported by the retrieved context
- Citation accuracy: whether citations point to the passages that support the attached claims
- Judge model: a model used to score answers, calibrated against human labels
- Regression threshold: the metric drop that blocks a release
Related reading
What is RAG? for the foundation, Golden question sets for the structured-data equivalent, and the pricing page.
Build the golden set first, measure retrieval and answers separately, report per category, and run it on every change; the system that is measured is the one that improves.
Frequently asked questions
How many questions does a golden set need?
▾
One to three hundred to start, spread across document types, intents and languages, with unanswerable questions included. Grow it monthly from real unanswered questions.
Can a model grade the answers instead of people?
▾
Yes for scale, after calibrating it against human scores on a sample and re-checking when the judge changes. Report human-verified numbers for any external claim.
What is a good RAG recall score?
▾
It depends on the corpus and question mix, which is why it is measured per category rather than promised. The useful question is whether recall on your hardest category is acceptable and rising.