Track four numbers to know if an AI agent is working: resolution rate (problems actually solved, not just handled), reopen rate (repeat contact within 48 hours, target under 10%), cost per resolution against your real human cost, and escalation quality (does a handoff carry context or make the customer start over). Skip containment rate and deflection rate. Both are easy to inflate and neither proves anything got solved.
Why containment rate lies to you
Most AI agent dashboards lead with containment rate: the share of conversations that never reached a human. It is the number vendors put on the homepage because it is always high and always going up.
The problem is what containment counts as a win. A customer asks a question, gets a non-answer, gives up, and closes the tab. That conversation is "contained." No human touched it. The dashboard logs a success. The customer got nothing.
Deflection rate has the same flaw from a different angle. It measures how many conversations got routed away from a human, not how many got solved. A platform can report 80% deflection while a large share of those customers never had their actual problem addressed, according to Notch's analysis of resolution rate benchmarks, which found that vendors "conflate three outcomes: genuine resolution, deflection, and containment" and that a platform reporting 80% by counting deflections is not outperforming one reporting a lower number on workflows that actually finished.
If you run an AI agent for support, scheduling, or intake and the only number you check is containment or deflection, you can watch it climb for months while the thing you actually care about, whether customers got helped, goes nowhere.
The four numbers that actually tell you it is working
Resolution rate. The share of conversations or tasks where the customer's problem was genuinely solved, not routed away or logged as handled. This is harder to measure than containment because it usually needs a follow-up signal (did they come back, did they rate it positively, did the underlying ticket close) rather than a single event. It is also the only one of the four that maps directly to whether the agent is worth what you are paying for it.
Reopen rate. The share of "resolved" conversations where the same person contacts you again about the same issue within 24 to 48 hours. This is the number that catches an inflated resolution rate. Fin AI's KPI framework puts a reopen rate under 10% in the "strong" range; above that, the agent is closing conversations it has not actually finished.
Cost per resolution. Total monthly agent cost divided by the number of conversations genuinely resolved, compared against what a human costs you to handle the same volume. Fin reports AI resolutions running $0.50 to $1.84 per contact against $6 to $8 or more for a human agent handling the same interaction, a roughly tenfold gap on routine requests, per the same KPI framework. Your own numbers will differ, but the comparison is what turns "the agent seems fine" into a dollar figure you can defend in a budget conversation.
Escalation quality. When the agent cannot finish something and hands off to a person, does that handoff carry the conversation history and what has already been tried, or does the customer repeat themselves from zero? This one rarely shows up on a vendor dashboard because it is not a single logged event, it is a design choice in how the handoff is built. A high escalation rate is not automatically bad if every escalation carries context and gets resolved fast. A low escalation rate is not automatically good if the agent is quietly logging failed attempts as contained instead of escalating them.
| Metric | Formula | Healthy range |
|---|---|---|
| Resolution rate | Conversations genuinely solved ÷ total conversations | 55% to 80%, tier-dependent |
| Reopen rate | Repeat contact within 48 hrs ÷ resolved conversations | Under 10% |
| Cost per resolution | Total monthly agent cost ÷ resolved conversations | $0.50 to $2, vs. $6 to $8 human |
| Escalation quality | Handoffs with full context carried over ÷ total escalations | Trending up, no fixed target |
What counts as a good resolution rate in 2026
Resolution rate benchmarks vary enormously by how the agent is built, not just by industry. Notch's platform-tier breakdown puts legacy, intake-only chatbots at 10% to 25% resolution, standard AI assistants that can retrieve information but not take action at 40% to 60%, top-tier AI-native platforms with real system integration at 55% to 70% first-contact resolution, and fully agentic platforms connected to backend systems like billing or scheduling at 70% to 85% end-to-end resolution. Notch also reports 77% autonomous resolution within 12 months across more than 20 million conversations it has processed for clients in insurance, e-commerce, and SaaS.
Fin's own benchmark, drawn from 12,000 customers, lands at 76% resolution, with a monthly improvement trajectory of roughly 1 percentage point as the agent's knowledge base gets tuned. The practical read for a small business: if your agent connects to real systems (a calendar, a CRM, an order database) rather than just answering from a document, expect it to climb toward 70% or better within its first year, not out of the gate.
How to calculate cost per resolution without an analytics platform
You do not need a dedicated AI observability tool to get this number, especially at the volume most small businesses run. Pull three figures from what you already have:
Total monthly agent cost. Your platform fee or your build's hosting and model usage cost, whichever applies. This is a line item you already pay.
Resolved conversations. Count the conversations where the customer's issue closed without a repeat contact within 48 hours. If your platform reports "resolved" or "handled" without that filter, subtract your reopen rate from the total first, or the number will read better than it is.
Your real human cost per contact. Fully-loaded cost, not just wage, meaning wage plus benefits plus the training and turnover cost of the role. If you do not track this, a support or intake role in most markets runs $18 to $30 an hour fully loaded; divide by how many contacts one person handles in an hour to get a per-contact figure.
Divide monthly agent cost by resolved conversations for your cost per resolution, then compare it to your human cost per contact. That comparison, not the raw agent cost, is the number worth putting in front of whoever approved the budget.
How fast should an AI agent pay for itself
Druid's ROI framework calculates AI agent ROI as total benefits minus total costs, divided by total costs. For a single, well-scoped workflow, the framework puts realistic payback at three to six months; a multi-workflow rollout across one department runs six to twelve months; a full department-wide transformation stretches to two to three years.
That timeline only holds if the agent is scoped narrowly. Two industry data points explain why scope matters more than the technology itself: an IBM 2025 study cited in the same framework found that only 25% of AI initiatives achieve their expected returns, while separately 74% of executives report seeing returns within the first year, per Google Cloud's 2025 research, when the deployment stayed narrow enough to measure. The gap between those two numbers is scope. A single agent handling one well-defined job is what shows up in the 74%. An ambitious system meant to handle everything at once is what shows up in the 25%.
If your agent has been live for more than two payback windows for its scope, three months for a single workflow, and none of the four numbers above are moving in the right direction, that is the signal to stop tuning and re-scope rather than keep adjusting prompts.
A simple monthly tracking routine
Check these four numbers on the same day every month, not just at launch. Most of the value of tracking them is the trend line, not any single reading:
- Pull resolution rate and reopen rate from your platform's reporting or your ticketing system's closed-then-reopened flag.
- Divide that month's agent cost by resolved conversations for cost per resolution.
- Spot-check five escalated conversations for whether the human received full context or had to ask the customer to repeat themselves.
- Compare all four against last month. A resolution rate climbing while reopen rate holds steady or drops is the pattern that means the agent is actually getting better, not just running longer.
Our custom AI agent development work includes this kind of measurement plan before an agent ever goes live, and our AI setup service covers the audit that catches a vanity-metrics dashboard before it becomes the thing your team reports on. If an agent you already have has been live for a quarter and nobody can tell you its reopen rate, talk to us about auditing it before you spend more on tuning it blind.
Frequently asked questions
What is a good resolution rate for an AI agent?
It depends on what the agent connects to. Simple, retrieval-only assistants typically land at 40% to 60%. Agents connected to real backend systems like scheduling, billing, or order status reach 70% to 85% once tuned. Fin's benchmark across 12,000 customers is 76%; Notch reports 77% within a year across 20 million conversations. Below 40% on an agent that has been live more than a quarter usually points to a scope or integration problem, not a model problem.
What is the difference between containment rate and resolution rate?
Containment rate measures whether a conversation avoided a human, including cases where the customer gave up without getting help. Resolution rate measures whether the customer's problem actually got solved. A high containment rate with a high reopen rate means the agent is closing conversations, not finishing them. Track resolution and reopen rate together; treat containment as a secondary number at best.
How soon should an AI agent pay for itself?
For a single, well-scoped workflow, three to six months is a realistic payback window, based on Druid's ROI framework. Multi-workflow deployments across a department run six to twelve months. If a narrowly scoped agent has been live longer than six months with no clear cost-per-resolution improvement, that is a sign to re-scope rather than keep tuning.
What is a normal escalation rate for an AI agent?
There is no fixed healthy number, because a high escalation rate on complex, high-stakes requests can be the right outcome. What matters more is the trend and the quality of the handoff: escalation rate should not be climbing month over month, and when it does escalate, the human receiving the case should get full conversation context rather than a customer who has to start over.
Do I need special software to track these four numbers?
No. Resolution rate and reopen rate can come from your existing ticketing or scheduling system if it logs a closed-then-reopened flag. Cost per resolution is arithmetic on numbers you already have: your monthly agent bill and your fully-loaded human cost per contact. A spreadsheet updated once a month is enough for most small businesses; a dedicated analytics platform only earns its cost at much higher conversation volume.
Why do vendor dashboards lead with containment or deflection instead of resolution?
Because containment and deflection are easy to measure automatically and almost always trend upward, which makes for a better-looking dashboard. Resolution rate requires a follow-up signal, like whether the customer came back, which is harder to instrument but is the number that actually correlates with the agent saving you money instead of just moving conversations out of a queue.
