azyware
RAG & knowledge engineeringMetric

Recall, precision and groundedness

Also: RAG evaluation metrics, retrieval quality metrics

In one sentence

What is Recall, precision and groundedness?

Recall, precision and groundedness are the three core measures of a RAG system: whether retrieval found the right passages, whether it avoided wrong ones, and whether the answer is supported by what was retrieved.

What Recall, precision and groundedness means

Recall asks: of the passages that contain the answer, how many did retrieval return? Precision asks: of the passages returned, how many were actually relevant? Groundedness (also called faithfulness) asks a different question of the generation step: is every claim in the final answer supported by the retrieved passages, or did the model add something? A fourth, answer relevance, checks whether the answer addresses the question at all.

Measuring them needs a golden dataset: real questions, the passages that answer them and a reference answer, built with the people who know the domain. Retrieval metrics are computed mechanically against the labelled passages; groundedness is usually scored by a judge model with human spot checks. The metrics are tracked per release so a change to chunking, reranking or the prompt can be shown to help or hurt.

These are not user-satisfaction scores and not a substitute for them, but they are the diagnostic layer underneath: a low-satisfaction assistant with high recall and low groundedness has a generation problem; with low recall it has an indexing problem. Without the split you are guessing.

Who it really matters to

  • CTO / Head of Engineering: These numbers turn RAG tuning from opinion into engineering; without them every change is a gamble.
  • Product manager: Groundedness is the metric behind "can we trust what it says", which is the adoption question.
  • Compliance officer: A tracked groundedness score and a golden set are evidence of due diligence when a regulator asks how the system is controlled.
  • Data lead: Owning the golden dataset and the eval pipeline is a durable, high-value responsibility that outlives any single model.

Why it exists

A RAG demo looks good on the ten questions the builder tried. Production users ask thousands of others, and failures are quiet: the answer is fluent and wrong. Recall, precision and groundedness exist to make quality visible and attributable before and after launch, and to separate retrieval failures from generation failures so the right component gets fixed. The trade-off is upfront effort building and maintaining a labelled question set and the discipline to run it on every change. That is the cost of the stance we take: evals over demos.

Where it is applied

  • Release gating for a SaaS support copilot: no deploy if recall or groundedness drops on the golden set
  • Regulatory evidence for a bank's policy assistant, showing answers are traceable to approved circulars
  • Comparing embedding models and chunk strategies for a hospital's protocol assistant on real clinician questions
  • Monitoring drift in an education tutor as the curriculum is updated each term
  • Vendor comparison during a ProofRun, scoring candidate models on the client's own questions rather than public benchmarks

Is Recall, precision and groundedness a skill?

MetricA set of numbers tracked per release against a golden question set. Eazyware builds the eval harness as part of every Retrieval & Knowledge Engineering engagement and reports the scores rather than showing a demo.

Eazyware service that covers it: Retrieval & Knowledge Engineering. Starting prices are on the pricing page.

Frequently asked questions

How many questions does a golden set need?

Enough to cover the real distribution of question types and document areas; a few hundred well-labelled questions usually beats thousands of synthetic ones. Start with fifty from real logs and grow it as failures are found.

Who scores groundedness?

A judge model does the bulk scoring, checking each claim against the retrieved passages. Domain experts review a sample every release to keep the judge honest, and disagreements go back into the golden set.

Related reading

Need Recall, precision and groundedness built, not just explained?

PRJECT IN MIND?