Once an agent is live, the reporting question arrives quickly, usually from someone who has to justify the deployment. The metrics that are easiest to produce are not the ones that tell you whether it is working.
The six to track
- Resolution rate. Share of calls where the caller's request was completed, verified against the transcript, not inferred from the call ending.
- Repeat-call rate within 48 hours. The clearest signal of a bad answer. A caller who calls back did not get what they needed.
- Escalation rate with reason. Escalation is fine. Escalation you cannot categorise is not.
- Unmatched intent rate. How often callers ask for something the knowledge base has nothing on. This is your content backlog.
- Time to first useful moment. Not time to answer, which is always instant. Time until the caller has said what they need and been understood.
- Abandonment against calls offered. The denominator matters more than the metric.
The three that mislead
Containment rate
Containment counts calls that did not reach a human. A caller who hung up in frustration is contained. A caller who got a confidently wrong answer is contained. Optimising for it directly pushes an agent toward refusing to escalate, which is the opposite of what you want.
Average handle time
On a human queue, handle time is a proxy for efficiency. On an agent it is a proxy for almost nothing. Short calls can mean crisp resolution or premature endings. Long ones can mean thorough handling or a loop. Read it split by outcome or do not read it.
Raw transcription accuracy
Word error rate is useful to engineers and misleading to everyone else. An agent can misrecognise several unimportant words and still act correctly, or get one account number wrong in an otherwise perfect transcript and fail the call. Judge on outcomes.
| Metric | What it actually measures | Read weekly? |
|---|---|---|
| Resolution rate | Whether the agent works | Yes |
| Repeat calls, 48h | Whether the answers were right | Yes |
| Unmatched intents | What to build next | Yes |
| Containment | Whether calls avoided humans | No |
| Average handle time | Very little on its own | No |
If you track one thing, track repeat calls. It is hard to game, easy to compute, and it moves when quality moves.
Written by
Lucy Product
Product



