How to measure whether AI proof of concept development is working
How do you measure AI proof of concept development?
Measure an AI proof of concept on four numbers against a golden set fixed before tuning: task quality, latency at the worst realistic payload, cost per completed task, and the shape of the failures. Everything else is commentary, and a demo is not a measurement.
Measure an AI proof of concept on four numbers, taken against a golden set that was fixed before any tuning began: task quality on your own data, latency at the worst realistic payload, cost per completed task, and the shape of the failures. Everything else is commentary. A demo that impresses a room is not a measurement.
What follows is the evaluation practice we run inside a three-week ProofRun: how to build the instrument, which four metric families to report, which popular numbers quietly mislead, and where to set thresholds so the result actually decides something.
A proof of concept has one job, so it has one scoreboard
The job is to tell you whether a model clears your bar on your data, cheaply enough and fast enough, before you commit a build budget. That makes evaluation the deliverable rather than an afterthought, which is the opposite of how most pilots are run. If the evaluation suite is built after the model looks good, it will be built to agree with the model.
An evaluation suite is a fixed set of inputs with agreed correct outputs, plus a harness that runs every candidate configuration against them and reports the same numbers each time. It is ordinary software engineering applied to a probabilistic component, and its value comes entirely from being written first. The argument is made at length in evals: the practice that separates AI demos from AI products.
One discipline matters more than any metric choice: the numbers must be comparable across runs. Same inputs, same scoring rules, same harness, recorded per prompt version and per model version. A 6 per cent improvement that came from quietly changing the test set is not an improvement, and it is the most common form of self-deception in AI work.
How do you measure AI proof of concept development?
Report four families of number, each with a threshold agreed in week one. The table gives the metric, what it actually tells you, how it is produced and the kind of threshold worth writing down.
| Metric family | What it tells you | How it is measured | Threshold worth agreeing |
|---|---|---|---|
| Task quality | Whether output is correct on your data | Golden set scored by exact match, field-level accuracy or human adjudication | A target percentage on the full set, plus a floor on the hardest subset |
| Groundedness | Whether claims are supported by retrieved sources | Per-claim checks against the cited passages | No unsupported claims above an agreed rate on the eval set |
| Latency | Whether the system fits the interaction | p50 and p95 at the largest realistic payload, not the average one | A p95 ceiling in seconds for the real worst case |
| Cost per completed task | Whether the economics survive volume | Total tokens and tool calls per task, priced at published rates | A ceiling per document, ticket or call at forecast volume |
| Failure distribution | Whether the errors are fixable or structural | Every failure categorised by cause, not just counted | No single cause above an agreed share of failures |
| Escalation or abstention rate | Whether the system knows when to stop | Share of inputs handed to a human with context | A band, since both zero and very high are warning signs |
The golden set is the instrument, so build it first
A golden dataset is one to three hundred real inputs with correct outputs agreed by someone qualified to judge them. Size matters less than composition. A set of a hundred examples that deliberately includes the awkward tail tells you more than ten thousand records of the happy path, because the happy path is not what will decide whether this reaches production.
Build it from three sources in roughly equal measure: representative everyday cases, known hard cases your team already argues about, and cases where two reasonable experts might disagree. That third group is the most valuable, because it forces your organisation to write down what a correct answer is, which is work you would otherwise discover halfway through a build.
Hold the set out. Nothing in the golden set may be used as a prompt example or a tuning input, because a system evaluated on its own training material reports its memory rather than its ability. Keep a separate development set for iteration and touch the golden set only to score.
The four metric families in practice
Task quality, scored the way the business scores it
Score at the level the work is done. For document extraction that means field-level accuracy, because a document with nineteen correct fields and a wrong account number is a failure, not a 95 per cent success. For retrieval-based answering, recall, precision and groundedness are the right triple, as described in how to measure RAG quality. For an agent, it is completion rate on scenario suites, covered in how to measure an AI agent.
Latency at the payload that actually arrives
Measure p95 on the largest realistic input, not the mean on the sample that fits comfortably. A forty-page scanned statement behaves nothing like a two-page form, and a voice interaction has a turn-taking budget measured in hundreds of milliseconds rather than seconds. A proof of concept that reports average latency on small inputs has measured the wrong thing politely.
Cost per completed task
Price the whole task, including retries, tool calls and any reranking step, rather than the single model call. Then multiply by forecast volume and check whether the business case survives. Use the LLM inference cost calculator for the forecast, and instrument traces so cost per request is visible from day one, as described in LLM observability.
The shape of the failures
Categorise every failure by cause: retrieval missed the document, the model misread a correct document, the schema was wrong, the input was genuinely ambiguous. A concentrated failure mode is usually fixable in days. Failures spread thinly across a dozen unrelated causes signal a structural problem that more prompt work will not solve, and that is exactly the finding a proof of concept exists to surface early.
Metrics that mislead
Each of these appears in pilot reports and none of them supports a build decision on its own.
- Demo success rate. Queries chosen by the person running the demo measure the demo, not the system.
- Aggregate accuracy with no subset breakdown. A strong headline number can hide near-total failure on the 10 per cent of cases that carry the commercial risk.
- Model-graded quality with no human anchor. Letting a model score its own family of outputs is cheap and drifts in flattering directions unless calibrated against human judgement on a sample.
- Benchmark scores from vendor cards. Public benchmarks measure public tasks. They do not predict performance on your document formats or your regional-language inputs.
- User satisfaction during a two-week pilot. Novelty inflates it, and the people testing are volunteers rather than the median user.
- Tokens saved or time saved, estimated. An estimate multiplied by headcount is a business case, not a measurement, and it should never appear in the same table as measured numbers.
What thresholds to set, and what measuring costs
Set thresholds in week one and record them in the statement of work. Good thresholds are specific and falsifiable: field-level accuracy above an agreed percentage on the full set and above a lower floor on the hard subset, p95 latency under an agreed number of seconds at the largest payload, and cost per task under an agreed figure at forecast volume. Vague targets produce a report nobody can act on.
Eazyware's three-week AI POC Sprint includes the harness, the evaluation report and a production readiness assessment, from $6,250 or ₹4,00,000 up to $10,500 or ₹6,80,000 where the data is complex or several models are benchmarked. After launch, keeping the suite green as models change is what the AI system add-on to a Care Plan covers, at $750 or ₹40,000 a month on top of a plan from $1,000 or ₹68,000. All starting figures are on the pricing page.
When heavy measurement is the wrong focus
There is such a thing as over-instrumenting a small question. If the decision is reversible, the volume is low and the cost of a wrong answer is a mild inconvenience, a hundred-example set and a single accuracy number is proportionate, and building a six-metric dashboard for it is displacement activity.
Measurement is also the wrong focus when the blocker is adoption rather than accuracy. A system that scores well and is not used has a change problem, and no additional metric on the model will reveal it. In that case instrument usage and outcomes in production instead, and accept that the proof of concept already answered its question.
What this looked like on a real system
For a hospital network's multilingual voice agent, the numbers that decided the design were not accuracy in the abstract. They were transcription quality per language, turn-taking latency under real telephony conditions, and the escalation rate to a human when confidence dropped. Measuring those three early changed the architecture rather than merely grading it. The system is described in the multilingual voice agent case study.
Related reading
From POC to production: the checklist covers what an evaluation suite must become before launch, and prompt versioning and evaluation explains how to keep results comparable across changes. Ragas, an open-source evaluation framework, documents metrics such as faithfulness and context recall for retrieval-based applications and is a reasonable starting point for a harness you intend to own.
Write the thresholds before the code, score against a set nobody was allowed to tune on, and the proof of concept will tell you something you did not already believe.
Frequently asked questions
What metrics should an AI proof of concept report?
▾
Task quality on a held-out golden set, groundedness where answers cite sources, p95 latency at the largest realistic payload, cost per completed task at forecast volume, and a categorised breakdown of failures. Each needs a threshold agreed before tuning starts, otherwise the result cannot settle a build decision.
How large should a golden evaluation set be?
▾
One to three hundred examples is enough for most proofs of concept. Composition matters more than size: include everyday cases, known hard cases and cases where two experts might disagree. Keep the set held out from tuning, because a system scored on its own examples reports memory rather than capability.
Can a language model grade its own outputs during evaluation?
▾
Model-graded scoring is useful for scale but must be calibrated against human judgement on a sample, otherwise it drifts in flattering directions. Use it to triage large sets, and have a qualified human adjudicate the hard subset and any case where the commercial consequence of an error is significant.