How to measure whether AI customer service agent is working
How do you measure AI customer service agent?
Measure an AI customer service agent on resolution rate, seven-day reopen rate, cost per resolved conversation, escalation quality and CSAT on agent-handled contacts. Deflection is not a metric: it counts conversations that ended, not problems solved. Track the five together, because each can be gamed alone.
Measure an AI customer service agent on five numbers: resolution rate, reopen rate within seven days, cost per resolved conversation, escalation quality and CSAT on agent-handled contacts. Deflection is not one of them; it counts conversations that ended, not problems that were solved. Track the five together, because every one of them can be gamed on its own.
This article sets out what each of those AI customer service agent metrics actually measures, how to build the evaluation suite that gates releases, what to instrument on day one, and how the numbers should move across the first ninety days.
The five metrics that matter
A useful support metric has three properties: it is defined without reference to the agent's own opinion, it can be audited against a sample of real conversations, and it gets worse when the agent behaves badly. Most dashboard metrics fail the third test.
| Metric | Definition | How it gets gamed | Sensible first target |
|---|---|---|---|
| Resolution rate | Share of conversations where the customer's stated issue was completed without a human | Counting any conversation the customer abandoned | Start at the intents you shadowed, not the whole inbox |
| Reopen rate (7 days) | Share of resolved conversations the same customer reopens within a week | Closing and reopening under a new ticket ID | Below your human team's current reopen rate |
| Cost per resolved conversation | Model spend plus tool calls divided by conversations actually resolved | Dividing by all conversations handled | Track the trend, not an absolute, for the first quarter |
| Escalation quality | Share of handovers that arrive with a correct summary and no repeated questions | Escalating early so the agent never gets anything wrong | Above ninety per cent from the first week |
| CSAT on agent-handled contacts | Satisfaction scored only on conversations the agent handled end to end | Surveying only successful paths | Within a point of your human baseline |
Two of these have glossary entries worth reading before you set targets: cost per resolved conversation and reopen rate. They are the pair that catches an agent which looks good on volume and bad on outcomes.
Notice what is absent from the list. There is no containment rate, no average handling time for the agent, and no intent-recognition accuracy. Containment and handling time measure the machine's convenience rather than the customer's outcome, and intent accuracy is an internal diagnostic that belongs in the eval suite rather than on a board slide. Report them if they help your engineers debug; do not judge the investment on them.
Why deflection misleads
Deflection counts conversations that did not reach a human. A customer who gave up, a customer who found the answer, and a customer whose refund was processed correctly all increment the same counter. That is why platform dashboards report high deflection while your reopen rate and your phone queue tell a different story.
The replacement is simple to state and harder to instrument: a conversation is resolved when the customer's stated issue was completed and they did not come back about it. We argued this case in full in AI ticket deflection is the wrong metric, and it is the standard we hold our own AI customer service agents to before any rollout widens.
How do you build an evaluation suite for a support agent?
You build it from your own history, before the agent exists. An evaluation suite is a fixed set of scenarios with known correct outcomes that the agent is run against on every prompt, model or retrieval change.
Sample from real tickets, not imagined ones
Pull one to three months of conversations, cluster them by reason, and take a stratified sample: the high-volume intents in proportion, plus deliberate over-sampling of the edge cases that caused complaints. Two hundred scenarios is usually enough to be decisive and small enough to maintain. This is the golden dataset, and it is an asset you keep long after any particular model is retired.
Score three things separately
Score the retrieval, the answer and the action independently, because they fail independently. Retrieval is judged on whether the correct policy or record was fetched. The answer is judged on groundedness, meaning every claim traces to something retrieved. The action is judged on whether the right tool was called with the right arguments. Open-source frameworks formalise the middle one: Ragas documents faithfulness and context-recall metrics for retrieval-grounded answers, and the same decomposition works whether or not you use that library.
Gate releases on the suite, not on a demo
Set a threshold per category and refuse to ship below it. When a provider deprecates a model, you re-run the suite and see the damage in an hour rather than in customer complaints over a fortnight. This is the practice we describe in evals: the practice that separates AI demos from AI products.
What to instrument on day one
Measurement is mostly a logging problem. If these are not captured from the first conversation, you will be reconstructing them from memory three months later.
- Full trace per conversation: every model call, prompt version, retrieved chunk, tool call and response, with token counts and latency.
- Intent label: assigned at the start and corrected at the end, so you can report resolution by intent rather than in aggregate.
- Action ledger: every write the agent made, with the threshold it was checked against and whether a human approved it.
- Escalation reason: a short enumerated code, not free text, so escalation patterns are countable.
- Customer follow-up window: a seven-day flag on every conversation, to compute reopen rate without manual joins.
- Cost attribution: spend tagged to conversation and intent, so an expensive intent is visible before the monthly bill arrives.
- Prompt and retrieval version: stamped on every conversation, so a regression can be traced to the change that caused it.
Most of this is standard LLM tracing, covered in LLM observability: tracing every request from prompt to cost. The support-specific additions are the action ledger and the follow-up window.
How the numbers should move in ninety days
In shadow mode, the only number that matters is proposal acceptance: how often a human agent accepts the agent's drafted response or action without material edit. Below seventy per cent, the agent is not ready for any intent. Above ninety per cent on a specific intent for two consecutive weeks, that intent is a candidate to go live.
After go-live, expect resolution rate to climb intent by intent, not overall. Expect reopen rate to spike briefly when a new intent is enabled and settle within a fortnight. Expect cost per resolved conversation to fall as routing and caching are tuned, typically by a third over the first quarter. If resolution rises while reopen rate also rises, the agent is closing conversations it has not solved, and you should narrow its scope immediately.
Where measurement goes wrong
The commonest failure is measuring the agent instead of the service. If the agent resolves forty per cent of contacts but your human queue is now slower because it holds only the hard cases, total customer experience may have worsened. Report the whole funnel: contacts, agent resolutions, escalations, human handling time and end-to-end time to resolution.
The second failure is a stale eval set. A suite built in January and never updated stops representing your traffic by June, and passing it means nothing. Refresh a fifth of it every quarter from recent tickets, and feed the failures into a knowledge-gap report so content fixes follow.
The third is measuring too soon. Statistical noise on a few hundred conversations will show you patterns that are not there. Wait for volume before you act on a weekly change, and use the eval suite, which is deterministic, for release decisions.
A fourth failure is quieter and more expensive: nobody owns the review. Metrics that no named person reads every week become decoration. We ask clients to name an escalation reviewer before go-live, someone with the authority to narrow the agent's scope on Monday if Friday's numbers were bad. Where that role exists, agents improve steadily. Where it does not, the dashboard is still green six months later and the support team has quietly stopped routing anything interesting to the agent.
One reporting habit is worth adopting from the first week. Publish a single weekly table with resolution, reopen, escalation quality, CSAT and cost per resolved conversation broken out by intent, and keep every historical week in it. Intent-level reporting is what lets you switch off one misbehaving intent instead of pausing the whole agent, and a running history is what turns an argument about whether last month was better into a two-second glance at a column.
What measurement costs
The eval suite and instrumentation are typically three to five days of work inside a build that runs from $12,500 or ₹8,00,000, and we do not quote an agent without them. Ongoing, the AI add-on to a care plan at $750 or ₹40,000 a month covers eval runs on every model change, prompt regression testing and cost monitoring, on top of a plan from $1,000 or ₹68,000 a month. The details sit on the maintenance and support page and the pricing page.
Related reading
How to measure an AI agent: the six metrics that matter generalises this to agents outside support, and how to measure RAG quality goes deeper on the retrieval half of the scoring.
An agent you cannot re-run against two hundred known cases is not measured; it is merely watched.
Frequently asked questions
What is the single best metric for an AI customer service agent?
▾
Resolution rate paired with seven-day reopen rate. Resolution alone rewards an agent that closes conversations without solving them, and reopen rate alone says nothing about volume. Read together they show whether customers' problems actually went away, which is the only outcome support exists to produce.
How many scenarios should an evaluation suite contain?
▾
Around two hundred, sampled from one to three months of real tickets in proportion to intent volume, with deliberate over-sampling of edge cases that previously caused complaints. That size is decisive enough to gate releases and small enough that a fifth can be refreshed each quarter without becoming a project.
How soon should an AI customer service agent show results?
▾
Proposal acceptance in shadow mode is readable within two to three weeks. Resolution rate becomes meaningful once an intent has been live for a fortnight at volume. Cost per resolved conversation usually falls over the first quarter as routing, caching and retrieval are tuned against real traffic.