Back to blog
supportmetricsai agentsAugust 8, 2026 · 5 min read

What the percentage of resolved conversations actually measures

An AI agent can close 85% of conversations without escalating to a human and still fix nothing. What that metric measures, and which one actually matters.

Iván Itzcovich

Iván Itzcovich

Co-founder

A company turns on an AI agent for support. The dashboard shows 85% of conversations close without reaching a human. Three weeks later, the same dashboard still reads 85%, and the support team is getting the exact same volume of WhatsApp complaints as before, except now they say "I already asked your bot and it didn't help." The number went up. Support didn't.

What does that 85% actually measure?

It measures how many conversations never reached a human, nothing more. That number, the containment or deflection rate, counts a customer who got the right answer exactly the same as one who got tired and closed the window. To the dashboard, both conversations are an identical success.

The flaw is in how the metric is designed: counting "did it reach a human, yes or no" is easy to instrument. Counting "did the customer's problem actually go away" requires knowing what happened after the conversation ended, and that data almost never lives on the same dashboard.

How much actually gets resolved, versus just contained?

An exact number beats a hunch here. Gartner surveyed 5,728 customers in December 2023 and found that only 14% of customer service issues get fully resolved through self-service. Even in cases the customer themselves described as "very simple," full resolution barely reaches 36%.

The proportion matters more than the number alone: if 14 out of 100 problems get resolved end to end, the other 86 didn't disappear, they just stopped generating a ticket with a name attached. They're still out there, and at some point they come back, almost always through a different channel, with the customer a bit more frustrated than the first time.

If it's not the model failing, where's the problem?

The same Gartner survey gives a clue that will sound familiar to a technical reader: in 43% of the cases where self-service failed, the reason was that the customer couldn't find relevant content for their problem. It wasn't that the model reasoned poorly. It was that the right answer wasn't where the system looked, or didn't exist in a retrievable format.

(If you've worked with RAG for more than a week, this number won't surprise you. It's the same old symptom with a new name.)

An additional 45% of customers felt "the company didn't understand what I was trying to do": the survey-language version of a poorly resolved intent, a step before the model ever generates any text. Both numbers point at the same place, the knowledge base and how the query gets interpreted before searching it, and neither calls for a bigger model.

What metric actually tells you if support improved?

Containment answers "did we avoid human contact?" The one that matters answers a different question: "did the customer write back about the same thing?" That second question is the recontact rate, and it has an advantage the first one doesn't: it's much harder to game. An agent can sound convincing and close a conversation without resolving anything; what it can't do is stop that same customer from writing back two days later if the problem was still there.

Neither metric alone is enough. Containment without the recontact rate is a cost number, not a quality one. Together they start saying something real: high containment plus low recontact signals the agent is resolving things; high containment with high recontact signals the agent is very well trained at saying goodbye.

Containment and recontact, read together

high containmentlow containment
low recontacthigh recontact

Recontact

Closes without a human and the customer doesn't write back about the same thing. It's the only quadrant where high containment means what the dashboard implies.

An illustrative reading of the four combinations. The top two report the same containment number.

Why does this metric change the contract negotiation?

Because it defines what gets paid for. A discount tied to message volume rewards more messages, which is exactly the number the company wants to bring down: if the agent resolves worse and needs six back-and-forths where two used to be enough, the bill goes up. The vendor gets paid more for working worse, and there's no way to defend that at the table.

A discount tied to resolution flips the incentive. The unit is the resolved conversation, so the price per unit drops as the share closed without a human goes up. Both sides end up looking at the same number, which is the only way a price discussion stops being a tug-of-war.

With one caveat, which is this article's whole point: a discount tied only to containment pays for closures, not resolutions. For the incentive to be set up correctly, the condition has to include recontact. High containment with high recontact isn't an outcome that deserves a discount.

How do you measure this without building a whole new dashboard?

You don't need a sophisticated evaluation system to start. It's enough to tag every conversation the agent closes with the topic and the customer, and cross that against new contacts from the same customer on the same topic within a fixed window: 48 or 72 hours is usually enough. That crossing already lives in the logs of any messaging system. It's a query on data that already exists.

What determines the outcome is which number the team looks at first in the weekly review. A team that starts by looking at containment ends up tuning the prompt to close conversations faster. One that starts by looking at recontact ends up tuning the knowledge base and the escalation rules, which is where the problem actually lived.

Whichever metric a team chooses to look at first ends up being the one that team optimizes, for better or worse. Before pushing an agent's containment number up, it's worth asking which of the two problems is actually being solved.


Sources: Gartner, "Gartner Survey Finds Only 14% of Customer Service Issues Are Fully Resolved in Self-Service" (press release, August 19, 2024, gartner.com). Survey of 5,728 customers, December 2023.

Share X LinkedIn WhatsApp
Iván Itzcovich

Iván Itzcovich · Co-founder, StudioChat

Want agents like these working for your team?

Talk to us now