azyware
Business

AI ticket deflection is the wrong metric: measure resolution

EZ
Eazyware
· 6 min read
Quick answer

Why is ticket deflection the wrong metric for AI support?

Deflection counts customers who gave up before reaching a person; resolution counts problems solved. Measure auto-resolution by intent, CSAT on AI-handled conversations and escalation reasons, and the number that goes on the board will be one the CFO and the customer both believe.

"Our bot deflects 60% of tickets" is a sentence that should worry you. It means six in ten customers did not reach a person; it says nothing about whether they got what they came for. Many left. Some tried again on another channel. Some told a friend. Deflection was the metric of the FAQ-bot era because it was easy to count. It is not the metric of the agent era, because agents can actually resolve, and resolution can be measured. This article sets out the measurement framework we install with every support agent, why each number matters, and how to present the result to finance and to the support team.

Why deflection misleads

  • It counts abandonment as success
  • It rewards bots that make it hard to reach a person
  • It hides which intents fail, because it is a single average
  • It ignores the second contact on another channel
  • It tells the CFO a cost story that the churn number later contradicts

The resolution framework

MetricDefinitionWhy it matters
Auto-resolution rate by intentConversations of that intent closed by the agent with no reopen within 7 days and no human touchThe only true measure of automation; per intent, never averaged
CSAT on AI-handled conversationsOne question asked after the conversation endsProves resolution felt like resolution
Escalation rate and reasonsShare handed to people, clustered weeklyThe roadmap: which reasons become new intents, which stay human
Reopen rateConversations reopened within 7 daysCatches false resolution
Actions within policyActions executed, and any outside policy (target zero)Safety, and the trust of the support manager
Knowledge gaps closedUnanswerable questions turned into help-centre content per weekCompounding improvement
Cost per resolved conversationRunning cost divided by resolved conversationsThe finance number
Hours returnedHuman handling time before minus afterWhere the value shows up

Per intent, or it is not a metric

A 70% auto-resolution average can hide "where is my order" at high resolution and "damaged item" at zero, which is the correct shape: the second should never be automated. Reported per intent, the number tells you what to expand and what to leave with people. Reported as an average, it tells you nothing and invites bad decisions, like pushing the agent into intents it should escalate.

How to measure CSAT without annoying customers

One question, after the conversation is closed, on the same channel, with an optional comment. Ask it on a sample if volume is high. Segment by whether the agent resolved or escalated; the escalated score tells you about the hand-off, which is the part most teams neglect. Never ask mid-conversation, and never let the agent ask for a good score.

Reopen rate: the honesty check

If the customer comes back within a week on the same issue, the first conversation was not resolved, whatever the agent logged. Reopens are subtracted from resolution automatically in our dashboard. A rising reopen rate on an intent is the earliest signal that a policy changed or a model update degraded behaviour, and it triggers an evaluation run.

Presenting to finance

Finance wants cost per resolved conversation before and after, hours returned, and the running cost line (inference, channel fees, care). Present the intents that are automated and the ones deliberately kept human, so the number is credible rather than inflated. The build and running costs are in How much does an AI agent cost?.

Presenting to the support team

The team wants to know the agent is safe and that it makes their day better: actions within policy (zero outside), escalations arriving with context, hours returned to complex work, and their own influence on the roadmap through the weekly escalation review. A support manager who owns that review is the single strongest predictor of a successful deployment.

A worked example

A subscription business measured its old FAQ bot at 55% deflection and was pleased until reopen and channel-switch analysis showed most "deflected" customers had emailed instead. The replacement agent, with account data and gated actions, reported per intent: billing questions and plan changes at high auto-resolution, cancellations deliberately routed to people. CSAT on AI-handled conversations was measured after each one; reopens were subtracted. Cost per resolved conversation and hours returned replaced deflection in the monthly report, and the retention team stopped hearing about the bot. The customer service agent page describes the build.

Dashboards and evaluation together

Production metrics tell you what happened; the evaluation suite tells you why. Every prompt or model change is run against golden conversations before release, so a drop in resolution on the dashboard can be matched to a regression in the eval. The two together are what let a team expand autonomy with confidence rather than hope. See What is an AI agent? for how evals fit.

Team and timeline

The measurement framework is installed in the first week of a support-agent build, before shadow mode, so the baseline exists. It is part of the customer service agent scope and maintained under a Care Plan. Your side owns the weekly review; we own the dashboard and the evals.

Before you start: a checklist

  • Define resolution: no reopen in 7 days, no human touch
  • List intents and mark which will stay human
  • Agree the CSAT question and when it is asked
  • Set the policy-violation target at zero and alert on it
  • Baseline current handling time and cost per contact
  • Name the support manager who owns the weekly review

Setting targets per intent

Targets are set after two weeks of shadow mode, from the agent's draft accuracy per intent, not from a vendor's benchmark. A read-only intent with clean data might target high auto-resolution from month one; a write intent starts lower and rises as the policy gate proves itself; a judgement intent targets zero automation and a fast, well-summarised hand-off. Revisit targets quarterly with the support manager and finance together, so the number on the board is one both believe.

The one-page monthly report

Page one: resolution by intent with reopens subtracted, CSAT for AI-handled and human-handled conversations, escalation reasons top five, policy violations (zero), knowledge gaps closed, cost per resolved conversation, hours returned. Page two: changes made this month (prompts, models, policies) with their evaluation results. The report is the same every month so trends are visible, and it goes to finance and the support team together, because they need the same numbers to trust the system.

Glossary

  • Deflection: contacts that did not reach a person; abandonment counts as success
  • Resolution: problems solved, with no reopen and no human touch
  • Reopen window: days within which a return on the same issue cancels the resolution
  • CSAT: satisfaction score asked after a conversation
  • Channel switch: the customer trying another channel after a failed one
  • Golden conversations: evaluation set of real chats with expected outcomes

Mistakes we see

Measurement goes wrong when the bot reports on itself, when averages replace per-intent numbers, when CSAT is asked mid-conversation, when reopens and channel switches are not tracked, and when the support manager is not part of the review. The framework in this article is designed so none of those can hide a failing intent.

Questions clients ask

  • How do we get a baseline? From the helpdesk: handling time, cost per contact and CSAT for the three months before launch.
  • Should we report to the board monthly? Yes: resolution by intent, CSAT, cost per resolved conversation and hours returned, with the intents kept human listed.
  • What if resolution drops after a change? The evaluation run identifies the regression; roll back the prompt or model and rerun.
  • Can we compare vendors on this? Yes, and you should: ask any vendor for per-intent resolution with reopens subtracted.
  • Does this apply to voice? Identically; voice adds handling time and per-minute cost.

What good looks like after 90 days

At ninety days the report shows per-intent resolution stable or rising, CSAT on AI-handled conversations at or above human-handled, reopens flat, zero policy violations and a knowledge-gap list that is shorter than in month one. That is the evidence to expand autonomy.

AI customer service agents: resolve, don't deflect, How to build a support agent that knows the customer's order, and the pricing page. Zendesk's own customer-experience research is a useful external reference for CSAT benchmarks.

Measure what customers experience, per intent, with reopens subtracted. Deflection flatters the bot; resolution builds the business.

Frequently asked questions

What is a good auto-resolution rate?

▾

It depends on the intent: high for status and simple changes, deliberately near zero for disputes and damage. Judge the mix, not an average.

Should we still track deflection?

▾

Only as a diagnostic alongside reopens and channel switches, never as the headline number.

How often should metrics be reviewed?

▾

Escalation reasons weekly during the first quarter, the full dashboard monthly, and the evaluation suite on every change.