azyware
Business

Evaluation suites vs manual QA: which to choose and when

EZ
Eazyware
· 7 min read
Quick answer

What is the difference between an evaluation suite and manual QA for AI systems?

Both are quality control, but they answer different questions. An evaluation suite runs fixed inputs against known-good outputs and scores every change automatically. Manual QA puts a person in front of the system to judge what a script cannot. Evals catch regressions; manual QA finds what nobody encoded.

Both are quality control, but they answer different questions. An evaluation suite runs fixed inputs against known-good outputs and scores every change automatically. Manual QA puts a person in front of the system to judge what a script cannot. Evals catch regressions cheaply at scale; manual QA finds what nobody thought to encode.

This article is about the trade-off rather than the technique: what each method actually detects, what each costs per release, how the ratio between them should shift as a system matures, and the specific situations where building an evaluation suite is a waste of money. If you want the case for evals in general, we have made it elsewhere; this is the comparison.

The scenario that creates the argument

A provider ships a new model version on a Tuesday. It is cheaper and faster, and the team wants it. Someone opens the product, tries six questions, sees good answers and approves the swap. Three weeks later, support notices that the system has quietly stopped citing sources on one document type, because the new model handles the instruction differently. Nobody tested that document type, because the person doing the testing did not know it mattered.

That is the whole argument in one story. Manual QA tested what a human thought to test. An evaluation suite would have tested what a previous release proved was important, which is a different and less flattering list. The failure was not laziness; it was that six questions cannot cover a system whose behaviour changes on every model, prompt and index update.

Definitions, precisely

An evaluation suite is a versioned set of test cases, each with an input, the context it should retrieve, an expected outcome and a scoring rule, run automatically against a build and reported as numbers you can compare across releases. Scoring may be exact match, rule-based checks such as required citation presence or schema validity, similarity against a reference answer, or a model grading against a written rubric. The suite lives in the repository beside the code. Its definition is in the evaluation suite glossary entry.

Manual QA is a person exercising the system against a plan or their own judgement, recording defects in prose. In AI systems it usually takes three forms: exploratory testing by an engineer, structured acceptance testing by the domain expert who owns the process, and review of real production transcripts after launch. The third is the most valuable and the most neglected.

What each method catches

DimensionEvaluation suiteManual QA
Primary jobDetect regression against known-good behaviourDiscover unknown failure modes and judge quality
CoverageHundreds of cases per run, identical every timeWhatever a person has time and imagination for
Feedback timeMinutes, in the pipeline, on every changeHours to days, scheduled around people
Marginal cost per runCompute and grading tokens onlyA person's day, every release
Setup costHigh: the golden set is the real workLow: start this afternoon
Blind spotAnything not in the set; drifts from reality if unmaintainedInconsistent, unrepeatable, biased towards recent memory
Best signalDid this change break something that used to workWould a real user accept this answer
Who owns itEngineering, with domain input on the rubricThe domain expert who owns the process
Scales withNumber of releases and model changesHeadcount, and nothing else

The honest case for manual QA

Manual QA is not the amateur option, and teams that automate too early ship systems that pass every test and annoy every user. Four things only a person will find.

  • Tone and register. An answer can be factually correct, properly cited and still read as dismissive to a customer chasing a refund. No automated grader we have deployed judges this reliably.
  • Unknown unknowns. Your golden set contains the questions you already knew about. Real users ask questions in the wrong order, in three languages, with typos in the invoice number.
  • Interface failure. The model is right and the product is still broken: the citation link goes nowhere, streaming stalls, the copy button loses formatting.
  • Domain judgement at the edge. Whether a particular clinical note is safe to summarise, or a particular loan case should be escalated, is a call a clinician or a credit officer makes, not a rubric.

The honest case for an evaluation suite

The suite earns its keep the second time you change something. Its advantage is not intelligence but consistency: it asks the same three hundred questions on every build, and it does not get tired at question forty. That makes four things possible that manual QA cannot deliver at any reasonable cost: swapping models without a leap of faith, tuning a prompt and seeing the effect on every case rather than the three you remember, catching a retrieval regression after a re-index, and giving a governance committee a number that moves. We treat prompts as versioned artefacts for the same reason, as described in prompt versioning and evaluation.

For retrieval systems the metrics are well established and worth adopting rather than inventing: recall, precision and groundedness, measured separately so you know whether the retriever or the generator failed. The method is set out in how to measure RAG quality.

The ratio that actually works

The answer is never one or the other, and the useful question is how the mix shifts over time. Before launch, the ratio is roughly even: you need a person exploring because you do not yet know what to encode, and you need the first fifty eval cases to stop the obvious regressions. During shadow mode, when the system proposes and humans approve or correct, manual review is the dominant activity and every correction becomes a new eval case at almost no extra cost. That is one reason we run shadow mode on nearly everything. After launch, the balance inverts: the suite runs on every change, and human time moves from testing to sampling production transcripts, which feeds the suite again.

A practical rule: every production incident becomes a permanent eval case within a week, and every eval case that has passed for a year without variation is a candidate for deletion. A suite that only grows becomes slow, expensive and ignored. Budget the human half honestly too: an hour a week spent reading real transcripts, by the person who owns the process rather than the person who built the system, is the cheapest source of new test cases you will ever find, and it is the first thing teams quietly stop doing when they get busy.

How to build the first suite in a week

  • Take fifty real inputs from logs or the existing process, not invented examples.
  • Have the domain expert write the answer they would accept, and note why.
  • Encode three cheap rule-based checks first: schema valid, citation present, no prohibited content.
  • Add similarity or rubric grading only for the cases where rules cannot express the requirement.
  • Run it in the pipeline from day one, even when the suite is tiny, so the habit forms.
  • Record cost and latency alongside quality, because a build that is right and unaffordable is still a failed build.
  • Review the set monthly against real traffic and retire cases that no longer represent anything.

When an evaluation suite is the wrong investment

Three situations. If you are still testing whether the idea works at all, a suite slows you down; run a proof of concept, learn what good looks like, then encode it. If the system changes once a quarter and a competent person can cover it in two hours, a suite costs more than it saves. And if the output has no defensible notion of correct, such as open-ended creative drafting where the brief lives in the reviewer's head, an automated grader will measure the wrong thing confidently. Be suspicious, too, of a suite built entirely from synthetic questions: it measures the model's agreement with its own generator, not with your users. A golden dataset from real traffic is the only version worth maintaining.

What this costs

Building the first suite alongside a system is usually five to ten working days of engineering plus two to three days of a domain expert's time, and it sits inside the build rather than as a line item. Where it becomes an ongoing cost is after launch. Our Care Plan tiers start at $1,000 or ₹68,000 a month for Essential, $2,500 or ₹1,60,000 for Standard and $5,250 or ₹3,40,000 for Enterprise with a named engineer, and the AI system add-on at $750 or ₹40,000 a month covers exactly this work: re-running evals, prompt regression, cost monitoring and re-indexing. Details are on the software maintenance and support page and the pricing page. A three-week AI POC Sprint at $6,250 or ₹4,00,000 is the right vehicle when you need to establish what correct means before you can encode it. Our other head-to-head guides are collected on the comparison hub.

The open-source Ragas documentation is a good primary source on metric definitions for retrieval-augmented systems, including faithfulness and context recall, and it will save you inventing your own scoring badly. On our side, evals: the practice that separates AI demos from AI products explains the discipline in full, and evals over demos sets out why we refuse to ship on a demo.

Use manual QA to discover what matters and an evaluation suite to make sure it keeps mattering after the next model ships.

Frequently asked questions

Can an evaluation suite replace manual QA entirely?

▾

No. A suite only tests what somebody encoded, so it cannot find failure modes nobody has imagined, judge tone, or catch interface problems around a correct answer. The suite handles regression at scale; people handle discovery and judgement. Teams that drop manual review stop learning what their users actually ask.

How many test cases does a useful evaluation suite need?

▾

Start with fifty real cases drawn from logs rather than invented ones, and grow towards two or three hundred as production incidents and shadow-mode corrections arrive. Size matters less than provenance: fifty genuine cases beat five hundred synthetic ones, because synthetic sets measure agreement with their own generator.

Who should write the expected answers?

▾

The domain expert who owns the process, not the engineer building the system. Engineers write answers the system can produce; credit officers, clinicians and support leads write answers the business will accept. Engineering then encodes those judgements as rules, similarity checks or a written rubric for model grading.