How to measure whether AI copilot development is working
How do you measure AI copilot development?
Measure an AI copilot on four levels: offline eval pass rate, task completion in production, weekly repeat use by the same people, and the business number the copilot was meant to move. Message volume and satisfaction scores are the two metrics that mislead most often.
Measure an AI copilot on four levels: offline evaluation pass rate before release, task completion rate in production, weekly repeat use by the same people, and the business number the copilot was built to move. AI copilot development metrics only mean something when all four are read together, because each one alone can be gamed.
This article sets out the four levels in order, names the metrics that flatter a copilot without telling you anything, and shows how the evaluation suite becomes a release gate rather than a report nobody opens.
Level one: offline evaluation, before anyone sees it
An evaluation suite is a fixed set of realistic inputs with known correct outcomes, run automatically against every prompt, model or retrieval change. For a copilot it has three parts: retrieval cases that check the right records come back, answer cases that check groundedness and citation, and action cases that check the copilot proposes the correct API call with the correct arguments.
Two hundred graded cases is a workable starting size for a product copilot, built from real support tickets and real session transcripts rather than invented questions. The practice is described in evals: the practice that separates AI demos from AI products, and it is the only measurement that runs before users are exposed to a change.
Grade three things separately. Retrieval recall tells you whether the answer was findable. Groundedness tells you whether the answer was supported by what was retrieved. Action correctness tells you whether the proposed write was the one a competent human would have made. A copilot can score well on the second and badly on the first, which means it is confidently answering from the wrong record.
Level two to four: what to watch once it ships
Each level answers a different question and has a different owner. Put all four on one page and review them weekly.
| Level | Metric | What it proves | Healthy direction |
|---|---|---|---|
| Offline | Eval pass rate by category | A change did not regress quality | Stable or rising, never shipped below baseline |
| Offline | Action correctness on scenario set | The copilot proposes the right write | Above the threshold you set for autonomy |
| Production | Task completion rate | Users finished the job without leaving | Rising over the first eight weeks |
| Production | Escalation and correction rate | How often a human overrides the copilot | Falling, with reasons categorised |
| Adoption | Weekly active copilot users over weekly active users | The feature became habit, not novelty | Rising then plateauing above your target |
| Adoption | Repeat use in week four by week-one users | Retention of the feature itself | Above half for a copilot doing real work |
| Business | Time to complete the target workflow | The copilot removed work, not added it | Falling against the pre-launch baseline |
| Business | Cost per completed task | The economics hold at volume | Falling as routing and caching improve |
Read the four levels as a chain of evidence. Offline pass rate says the change is safe to release. Task completion says the copilot works on traffic you did not anticipate. Adoption says users believe it. The business number says the company got something for the money. A break anywhere in that chain tells you exactly where to look: high eval scores with low completion means your test set is unrepresentative, and high completion with low adoption means the copilot is good at a job nobody needed done.
Which metrics mislead?
Three numbers look like progress and are not. Watch them, but never report them as success on their own.
- Message volume. A copilot that generates many messages per task is usually failing to understand the first one. Volume rises when the copilot is confusing as readily as when it is useful.
- Thumbs-up rate. Fewer than one in twenty users rate anything, and those who do are the delighted and the furious. Use ratings to find examples to inspect, not to measure quality.
- Deflection. Counting conversations that did not become tickets rewards a copilot for stonewalling. Measure resolution instead, as argued in AI ticket deflection is the wrong metric.
- Total users who tried it. First-week curiosity is not adoption. The honest number is how many of those people came back in week four.
- Average latency. Averages hide the tail that users remember. Report the ninety-fifth percentile for first token and for full response.
The discipline here is the same one site reliability engineering applies to services: pick a small number of indicators that track the user's actual experience and set explicit objectives against them, as Google's service level objectives chapter describes. A copilot with twenty dashboards and no agreed threshold has no measurement at all.
How to make evals a release gate
Version the prompts like code
Prompts, retrieval configuration and tool definitions live in the repository, not in a console. Every change is a pull request that runs the suite. The mechanics are in prompt versioning and evaluation.
Trace every request
You cannot debug a copilot from its output. Store the retrieved chunks, the tool calls, the arguments, the model used and the token cost for every request, and sample them weekly. LLM observability turns a vague complaint into a reproducible case you can add to the eval set.
Set the shadow threshold in advance
Before launch, agree the action correctness number at which the copilot may act without approval, per action type. Reassigning a record might pass at a lower bar than issuing a credit. Writing the threshold down before you see the results is what stops the bar moving to meet the system.
Grow the suite from failures
Every escalation, correction and complaint becomes a new graded case within the week. A suite that does not grow is a suite that stopped reflecting your product.
Sample traces by hand as well. Twenty randomly chosen conversations a week, read end to end by an engineer and the product owner, will surface problems no aggregate catches: a retrieval source that has gone stale, a refusal rule that fires too often, a tenant whose vocabulary the copilot has never seen. We treat that reading session as part of the release process, not as an optional extra, because it is where new eval cases come from.
What measurement costs
Building the first evaluation suite is roughly a fifth of a copilot build, and it is included in our AI copilot development for SaaS programme, which runs from $19,500 or ₹12,80,000 to $63,000 or ₹41,60,000 depending on how many jobs write to your systems. Starting prices for every programme are on the pricing page.
Keeping it alive is the part teams forget to budget. Models are deprecated, your product changes, and an unmaintained suite quietly stops gating anything. The AI add-on to a Care Plan is $750 or ₹40,000 a month on top of a plan from $1,000 or ₹68,000, and it covers eval runs, prompt regression, cost monitoring and re-indexing. That is the cheapest insurance in the whole programme.
When this measurement stack is the wrong choice
If the copilot is a two-week experiment for fifty internal users, four levels of measurement is overhead. Run twenty eval cases, watch weekly repeat use, and talk to five users. Formal suites earn their keep when the copilot is in front of customers or touching money.
If you have no pre-launch baseline for the workflow, do not fabricate one. Measure the workflow manually for two weeks first, with a stopwatch and a spreadsheet, or your business-level number will be an argument rather than evidence. A copilot that cannot be compared to anything is impossible to defend at the next budget review.
And if nobody owns the weekly review, stop. Metrics without an owner become decoration within a month, which is the failure pattern in copilot adoption: why most AI features die in a month.
A worked example
For a field-service SaaS copilot, the jobs were reassignment, scheduling and status changes. The eval set was built from a year of support transcripts, with action cases graded against what the dispatcher actually did. The copilot ran in shadow mode while dispatchers accepted or corrected proposals, and the acceptance rate by action type decided which actions were released first. The engagement is described in the in-app copilot case study.
The number that changed the roadmap was not accuracy. It was the share of dispatchers who used the copilot again in week four, which stayed low until the copilot was moved out of a side panel and into the job record itself. The model had not changed at all. Placement and default behaviour moved adoption more than any prompt revision in the entire programme, and only the adoption metric made that visible.
Before you instrument anything
- Write down the one business number the copilot exists to move
- Measure that number for two weeks before launch
- Build two hundred eval cases from real transcripts, not imagined ones
- Split grading into retrieval, groundedness and action correctness
- Agree the autonomy threshold per action type, in writing
- Instrument traces on day one, including retrieved chunks and token cost
- Name the person who reads the weekly review and can change the roadmap
Related reading
How to measure an AI agent: the six metrics that matter covers the acting side in more depth, evals over demos explains why we refuse to ship on a demo, and reporting on AI cost per account shows how to turn token spend into a number your finance team recognises.
A copilot is working when the same people use it again next week to finish the same job faster, and everything else is a proxy for that.
Frequently asked questions
What is the single most important AI copilot metric?
▾
Weekly repeat use by the same people, measured as the share of week-one users still using the copilot in week four. It is the hardest number to game: users only return to a feature that saves them time on work they have to do anyway. Everything else supports or explains it.
How many evaluation cases does a copilot need?
▾
Around two hundred graded cases is a workable starting suite for a product copilot, drawn from real transcripts and support tickets and split across retrieval, groundedness and action correctness. The suite then grows from production failures, with every escalation or correction added as a new case within the week.
How often should copilot evals run?
▾
On every change to prompts, retrieval configuration, tool definitions or model choice, as an automated step in continuous integration, plus a scheduled run weekly to catch drift from data and index changes. A suite that runs only before major releases will not catch the regressions that matter.