← All notes

Voice AI

Deflection is a vanity metric on its own

A voice agent that handles 70 percent of calls without a human may be succeeding or may be hanging up on people. The measurement has to say which.

  • Invexa Technologies
  • 3 min read

Six weeks after launch, a voice agent deployment reports 72 percent deflection and the project is declared a success. Two months later, repeat call volume is up 18 percent and the human queue is full of people who are annoyed before they say hello. The agent was not resolving calls. It was ending them.

Deflection counts calls that did not reach a person. It says nothing about whether the caller’s problem went away. On its own it rewards exactly the wrong behaviour, because the cheapest way to raise it is to make escalation harder.

The Metrics That Actually Discriminate

A useful scorecard separates what the agent did from what the caller got.

  • Containment. The share of calls fully handled by the agent, with no transfer and no callback. This is the honest version of deflection, and it should be the denominator for everything else.
  • True resolution. Containment minus any caller who calls back on the same issue within 72 hours. Repeat contact is the single best lie detector in the whole set, and a healthy deployment sits under 8 percent.
  • Escalation quality. Of calls that transferred, how many arrived with a correct summary and the customer’s records already open. A transfer that makes the caller repeat everything is worse than no automation.
  • Task completion. For agents that do real work, booking, payment capture, address change, the percentage where the backend write actually succeeded. Saying “your appointment is confirmed” without a confirmed row is the most damaging failure available.
  • Abandonment during automation. Callers who hang up mid-conversation. Above 10 percent, something in the flow is failing people and deflection is quietly counting it as a win.

Read The Transcripts, Weekly

Dashboards tell you the shape of the problem. Transcripts tell you the cause. A weekly review of 30 to 50 calls, sampled deliberately from the failure buckets rather than at random, will surface things no metric names: an intent that the agent recognises but has no tool for, a phrasing customers use that the prompt never anticipated, a confirmation step that sounds ambiguous when spoken aloud.

Build a small labelled evaluation set from those calls, 100 to 200 real utterances with expected outcomes, and run it before every prompt or model change. Without it, every improvement is a guess, and prompt edits that fix one intent quietly break two others.

Set Targets Before Launch

The most common process failure is measuring after the fact. Agree the numbers before go-live: which intents are in scope, what containment is expected for each, what latency is acceptable, and what escalation rate is considered healthy rather than a failure. A first line support agent handling order status and appointment booking should reach 60 to 75 percent containment on in-scope intents within a couple of months. An agent asked to handle billing disputes should not, and forcing it to will produce the repeat-call pattern described above.

Publish the scorecard where the support team can see it. They will spot the discrepancy between the reported number and the calls they are actually receiving faster than any analytics job will.

At Invexa, we agree the measurement framework before we write the first prompt, because an agent optimised for the wrong number gets very good at the wrong thing.

Next step

Have a project in mind?

A 30-minute call is usually enough to know whether we are the right team for it. If we are not, we will say so.

Start a project

Replies within one working day

Start a project