Here's a number that should make you suspicious: Intercom's Fin AI Agent reports an average resolution rate of 66-67% across 6,000-plus customers, yet independent case studies of the same kind of deployment run closer to 42-50% (Intercom, 2025). That gap isn't a rounding error. It's the difference between two metrics that get used interchangeably and shouldn't be: containment (the conversation never reached a human) and resolution (the customer's problem actually got solved). If you buy an AI agent on the containment number and manage it on the containment number, you will think it's working long after your customers have decided it isn't.
This article unpacks why those two metrics diverge, why vendor "resolution" claims tend to be inflated, and exactly how a small business should measure the one that matters.
What is the difference between containment and resolution in AI customer service?
Containment means the conversation ended without a human touching it. Resolution means the customer's actual problem got solved. They sound like the same thing, but a customer who gives up in frustration and closes the chat window is "contained" and not resolved.
Containment is the easier number to hit and the easier number to fake. If your AI agent never escalates, your containment rate approaches 100% by definition, including every case where the customer rage-quit, gave up, or went and emailed you instead. That's why containment was the headline metric in the old deflection-focused chatbot era. It measured how much work you avoided, not how much value you delivered.
Resolution is the honest metric because it's tied to an outcome. Did the order get tracked, the refund get processed, the question get answered correctly and completely? A good way to think about it: containment is about your cost, resolution is about your customer. You can cut cost while quietly destroying the experience, and the containment number will happily hide it. This is the same trap that shows up in the broader AI customer service benchmarks for 2026 — the metric you optimize is the metric you get.
Why are vendor "resolution rate" claims inflated?
Vendor resolution claims are inflated mostly because of how "resolved" gets defined and who gets to define it. When the vendor counts any conversation the AI closed without escalation as a resolution, they've quietly relabeled containment as resolution — and the number jumps.
The clearest evidence is the gap inside a single vendor's own data. Intercom's Fin reports a 66-67% average resolution rate, with more than 20% of customers exceeding 80%, but independent and customer case studies of comparable deployments land at 42-50% (Intercom, 2025). Both figures can be technically true. The headline average pools the most mature, best-tuned accounts, while the realistic early-maturity range for a fresh deployment sits 15-25 points lower. The number you'll personally see in month one is the lower one, not the brochure one.
There's a second reason the marketing number runs hot: definitions drift in the vendor's favor. Some platforms count a conversation as resolved if the customer simply didn't reopen it within a window, which assumes silence equals satisfaction. It often doesn't. Plenty of unhappy customers don't reopen a chat — they just stop using you and tell a friend. To their credit, some vendors are tightening this — Zendesk now uses an LLM to verify resolutions specifically to stop inflated counts. The takeaway isn't that vendors are lying. It's that "resolution rate" is not a standardized term, so a 67% from one tool and a 45% from another may be measuring completely different things.
There's also a maturity-and-survivorship effect worth naming. The accounts that show up in a vendor's headline average are disproportionately the ones that stuck with the product, invested in their knowledge base, and tuned weekly — the 20%-plus of Intercom customers exceeding 80% resolution did real work to get there (Intercom, 2025). The accounts that gave up at 35% in month two quietly drop out of the dataset. So the published average isn't a forecast of your result; it's a snapshot of the people who climbed the curve and stayed. Before you sign anything, ask the vendor pointedly: how do you define a resolution, do you verify it, and what does the median new customer see in their first 90 days — not the all-time average across your best accounts.
What is a good AI resolution rate for a small business?
A realistic AI resolution rate for a well-run SME deployment starts around 30-50% and climbs toward 65-80% only with a mature knowledge base and ongoing tuning. If a vendor promises you 80% out of the box, treat it as a sales claim, not a forecast.
The single biggest determinant of your result is deployment maturity, not which vendor you pick. The supporting market data is consistent here. McKinsey found generative AI can reduce the volume of human-serviced contacts by up to 50% and unlock up to 60% of addressable care volume (McKinsey, 2023; McKinsey, 2025). Salesforce reported that 30% of service cases were resolved by AI in 2025, projected to rise to 50% by 2027 (Salesforce State of Service, 2025, forecast). Notice that even these optimistic, well-resourced benchmarks cluster in the 30-60% range — nowhere near the 80% you'll hear in a demo.
So set staged targets and judge against your own trajectory, not a competitor's billboard. Intercom community guidance suggests aiming for roughly 70% bot CSAT in month one and 75-80% by month three, with resolution starting at 30-50% and improving through weekly QA review. The trend line matters more than the starting point.
It also helps to be honest about the ceiling. Even Gartner's most aggressive forecast — that agentic AI will autonomously resolve 80% of common customer service issues by 2029, cutting operational costs 30% (Gartner, 2025, forecast) — carries the word "common" for a reason. The 80% applies to high-volume, well-structured questions, not to the messy disputes, edge cases, and emotional conversations that make up the long tail of any real support queue. A blended resolution rate that includes those will always sit below the headline. If your structured intents resolve at 75% and your overall number is 50%, you're not failing — you're seeing the long tail honestly, which most vendor dashboards are designed to blur.
Here's a realistic maturity curve to plan against:
- Month 1: 30-50% resolution on your highest-volume, most-structured intents (order status, hours, basic FAQs). Expect to find gaps fast.
- Months 2-3: 50-65% as you curate the knowledge base and fix the answers your customers actually rephrase.
- Months 4-6+: 65-80% is achievable on well-defined intents with disciplined weekly review — but sentiment-heavy and dispute-type questions will stay lower, and that's normal.
How should an SME actually measure AI resolution?
Set a baseline before you deploy, then track resolution by intent — not a single blended containment percentage. Without a "before" snapshot of your current first-response time, resolution rate, CSAT, ticket volume, and cost per contact, any ROI claim you make later is unprovable.
The cost numbers are worth grounding in real benchmarks so you know what a contained-and-resolved conversation is worth. Gartner pegs the median cost per contact at $1.84 for self-service versus $13.50 for assisted channels (Gartner). A 2020 Forrester Total Economic Impact study of IBM watsonx Assistant put savings at $5.50 per contained conversation, with a 337% three-year ROI and payback under six months (Forrester Consulting, commissioned by IBM, 2020) — directional and not SME-specific, but useful for sizing. The point: the dollar value is real, which is exactly why you can't afford to count frustrated drop-offs as wins.
A practical measurement framework that won't fool you:
- Track resolution, not just deflection. Verify that the problem was solved, ideally with a post-chat "did this answer your question?" or an LLM-based resolution check.
- Watch re-contact rate — customers who return within ~72 hours. A rising re-contact rate is the smoking gun behind an artificially high "resolved" number.
- Compare AI CSAT to your own human-agent CSAT, not industry averages. Intercom notes customers score AI roughly 5-10 points harder, so a small gap is normal, not a failure.
- Segment by topic. High-structure intents (authentication, order status, refunds) resolve far better than sentiment-heavy or dispute intents. A blended number hides both your wins and your problems.
- Review weekly. The McKinsey-style gains (a 5-10% CSAT lift and 25-30% efficiency improvement) come from iteration, not installation (McKinsey, 2024).
This is the same discipline behind a proper 30-60-90 day KPI plan for AI agents: measure the right thing, on a schedule, against your own baseline.
Containment vs resolution: a side-by-side comparison
Containment tells you how much human labor you avoided; resolution tells you whether the customer left satisfied. The table below makes the trade-off explicit so you can see why optimizing the wrong one quietly backfires.
| Dimension | Containment rate | Resolution rate |
|---|---|---|
| What it measures | Conversation didn't reach a human | Customer's problem actually got solved |
| Whose interest it serves | Your cost line | Your customer's outcome |
| Easy to inflate? | Yes — never escalate and it nears 100% | Harder — tied to a verified outcome |
| Counts a frustrated drop-off as success? | Yes | No |
| Typical vendor headline | High (often the marketing number) | Lower; 66-67% reported avg vs 42-50% independent |
| Best guardrail metric to pair with it | Escalation rate | Re-contact rate within ~72h |
| Right way to use it | Capacity/cost planning only | The primary success metric |
The honest read of this table: containment is a useful operational stat for capacity planning, but it should never be your headline success metric. If your containment is 90% and your resolution is 45%, you don't have a great AI agent — you have an AI agent that's good at not escalating.
Why does optimizing containment backfire for SMEs?
Optimizing containment backfires because the fastest way to raise it is to stop escalating, and the fastest way to stop escalating is to let the AI guess instead of handing off. That trades a visible cost (a human reply) for an invisible one (a customer who quietly leaves).
This isn't hypothetical. Forrester predicted that in 2026, a third of companies will harm customer experiences with frustrating AI self-service, driven by pressure to cut costs by deploying customer-facing AI prematurely (Forrester, 2026). Containment-rate marketing is the mechanism: it rewards exactly the behavior — never hand off, always answer — that produces those bad experiences. And the reputational risk compounds, because 84% of consumers already believe human agents are more accurate than AI, with just 8% preferring AI over humans (SurveyMonkey, 2025). You're starting from a trust deficit; a contained-but-unresolved conversation widens it.
The fix is to design for clean escalation and measure resolution as your north star. An AI agent that knows when to say "let me get a person on this" will show a lower containment rate and a higher resolution rate — and that's the trade you want. This is where a platform's design philosophy matters more than its feature list. Omago, an AI agent platform that helps SMEs automate customer conversations across WhatsApp, Telegram, and web chat, is built around grounded answers and clean handoffs rather than maximizing how many conversations it can keep away from your team — because the goal isn't to avoid your customers, it's to serve them and capture the lead when a human is actually needed.
How do you stop your AI agent from gaming its own numbers?
Force the AI to abstain when it isn't sure, and verify resolutions with an outcome signal rather than trusting silence. An AI agent instructed to say "I don't know" and escalate when context is missing will report a lower, more honest number — and that honesty is the entire point.
The mechanics matter because AI can be confidently wrong. Peer-reviewed research found that large language models "can hallucinate with high certainty even when they have the correct knowledge" — they often sound most confident exactly when they're wrong (Simhi et al., Technion/Oxford/Hebrew University, 2025). A model that's optimized to always answer will produce a beautiful containment rate built on fluent, authoritative, incorrect answers. Grounding every reply in your verified knowledge base via retrieval-augmented generation, and curating that base to remove stale or conflicting articles, is what keeps "resolved" honest.
A few concrete anti-gaming controls:
- Verify, don't assume. Use a post-conversation confirmation or an LLM-based resolution check instead of counting "didn't reopen" as solved.
- Make escalation a feature, not a failure. Track escalation rate alongside resolution; a healthy deployment escalates the cases it should.
- Treat re-contact rate as a lie detector. If "resolved" goes up but customers keep coming back, the resolution number is fiction.
- Pick tools that take real actions. A genuine resolution often means doing something — capturing and routing a lead, running a guided multi-step flow, triggering an action through an integration like Airtable — not just returning text that sounds like an answer.
Frequently Asked Questions
What is the difference between deflection and resolution rate?
Deflection (often used interchangeably with containment) counts conversations that didn't reach a human agent. Resolution counts conversations where the customer's problem was actually solved. A customer who gives up in frustration is deflected but not resolved, which is why deflection can hide failure. Always treat resolution as the primary success metric and use deflection only for capacity and cost planning.
Why is Intercom's 66-67% resolution rate different from independent reports of 42-50%?
The 66-67% figure is a vendor-reported average across 6,000-plus customers, pooled toward the most mature, best-tuned accounts (Intercom, 2025). Independent and customer case studies of comparable deployments report 42-50%, which reflects the realistic early-maturity range a new deployment will actually see. Both can be true; they're measuring different populations at different maturity levels.
What is a realistic AI resolution rate to expect in the first month?
Plan for 30-50% resolution in month one on your highest-volume, most-structured intents, improving toward 65-80% over several months with knowledge-base curation and weekly review. Intercom community guidance suggests targeting roughly 70% bot CSAT in month one and 75-80% by month three. If a vendor promises 80% resolution immediately, treat it as marketing.
How do I know if my AI agent is faking its resolution numbers?
Watch your re-contact rate — customers returning within about 72 hours. If your reported resolution rate climbs while re-contacts also climb, the resolution number is inflated. Pair this with a post-chat confirmation or an LLM-based resolution check, and compare AI CSAT against your own human-agent CSAT rather than industry averages.
Should I worry if my containment rate drops after deploying clean escalation?
No — a lower containment rate paired with a higher resolution rate is usually a healthy sign. It means your AI agent is correctly handing off the cases it shouldn't attempt, instead of guessing and producing confidently wrong answers. Containment is a cost stat; resolution and customer satisfaction are the outcomes that actually grow the business.
Sources: Intercom (2025); McKinsey (2023, 2024, 2025); Salesforce State of Service (2025); Gartner; Forrester Consulting commissioned by IBM (2020); Forrester (2026); SurveyMonkey (2025); Simhi et al., Technion/Oxford/Hebrew University (2025).
