Someone above you has asked what "autonomous" would mean for your support headcount, and you need an answer that survives the follow-up questions. The vendors are not much help, because each of them defines the word to match what they sell.
The reading is harder than it should be. The vendor pages ranking for the term define it in ways that do not line up, and one of them describes AI that helps agents write replies faster. The forecast that gets quoted everywhere, that agentic AI will resolve 80% of common customer service issues by 2029, is an analyst prediction published without a methodology, and the pages quoting it rarely say so.
Underneath the vocabulary there is a capability with clear edges. Software can now hold a support conversation, look things up, take an action in another system and close the ticket, with no person reviewing the steps. It works for a bounded set of problems, fails on others, and the failures are specific enough to design around.
Pluno publishes this guide and sells an AI support agent. Every figure below carries a named source and a date, and where no credible figure exists the guide says so instead of filling the hole.
TL;DR
-
The definition is contested. The ranking vendor pages disagree on whether escalation is part of it and whether the action has to complete, and one describes AI-assisted service while calling it autonomous.
-
Deflection is not resolution. A contained conversation and a solved problem are different events, and published rates rarely say which one they counted.
-
Silence can produce a billable resolution, and the mechanisms differ. Fin scores 24 hours of customer disengagement as an assumed resolution. Zendesk uses 72 hours as a session cutoff and then verifies the conversation with a model. Pluno uses a 72-hour window as well. Ask every vendor, ours included, what triggers a charge.
-
Gartner's 80% is a forecast, published with no methodology, and so is the pessimistic forecast that 40% of agentic AI projects will be cancelled.
-
Gartner's own later research points the other way. Half of the companies that blamed headcount cuts on AI are predicted to rehire under different job titles by 2027, and full automation is now described by Gartner as prohibitively expensive.
-
Measure resolution. Containment is the easier number and the less useful one. Get the vendor's definition of a billable event in writing.
No trustworthy public figure exists for how often these systems are confidently wrong in production. Benchmarks measure task failure. Nothing measures the live queue.
What is autonomous customer service?
Autonomous customer service is customer support handled end to end by software, from the customer's first message to a resolved issue, without a person reviewing or approving each step.
The definition hides a real disagreement. The vendor pages ranking for this term describe it in four incompatible ways, and the differences change what you would be buying.
Salesforce leaves escalation out. Its definition is AI, natural language processing and machine learning performing "customer service tasks without human intervention," full stop. The page mentions handing off to a rep elsewhere, so the point covers its definition only.
Klaviyo builds escalation in. Its agents resolve issues autonomously, "then escalate to human agents when necessary." Escalation is part of what the thing is.
Pega frames it around the back office. Pega publishes no definition of the term. It describes straight-through processing resolving inquiries in the channel, which points at what the others leave out. The work behind the reply, meaning the refund issued and the claim processed. That is the most demanding version of the idea and the least often met in practice.
ContactPoint360 defines something else entirely. Its page calls autonomous customer service an AI-driven model where machine learning and speech analytics are "used to improve customer interactions." Improving interactions describes AI-assisted service. Autonomy is a stronger claim.
The useful definition takes Klaviyo's shape and Pega's scope. Autonomy means no person in the loop per step, which is a different claim from no person ever, and it means the action completes rather than stopping at the reply. If the agent promises a refund it cannot issue, the ticket has been postponed.
Conversational AI for customer service covers the interface layer underneath.
Autonomous, automated, assisted, agentic: four words vendors use interchangeably
These four words describe different things. The difference is who stays in the loop and what you end up buying.
| Term | What it means | Who is in the loop | What you buy |
|---|---|---|---|
| Automated | A rule fires. Deterministic, no judgment. | Nobody, and nobody was needed. | A workflow |
| Assisted | AI drafts, a person sends. | A person, on every reply. | Faster replies |
| Autonomous | AI handles the conversation and the action end to end. | A person, only on escalation. | Resolved tickets |
| Agentic | An architecture: plan, call tools, act, observe, repeat. | Depends entirely on scoping. | An architecture |
The last row carries the distinction worth holding. Agentic describes how the software is built. Autonomous describes what it is allowed to do. A tool can be agentic in architecture and barely autonomous in deployment, because every action it takes is scoped to read-only lookups. Vendors blur the two because "agentic" sells and autonomy has to be defended with numbers.
How autonomous customer service works
The mechanism runs in four stages, and every vendor page that describes one describes some version of these.
1. Understand. Classify the intent, pull out the entities that matter (order number, account, error code), and retrieve the context on the customer: prior contacts, account state, entitlements.
2. Ground. Retrieve the knowledge the answer will come from. Most systems ground in help centre articles and documentation, which is why most are good at documented questions. A smaller set also ingests the resolved support tickets a team has accumulated, reaching issues nobody wrote up. The grounding choice decides most of what a system can handle and is rarely on the pricing page.
3. Act. Call the systems that hold the answer or perform the change. Reset the entitlement, re-run the failed import, issue the credit. This is the stage Pega's framing points at, and many deployments stop short of it. An agent that cannot act can only tell the customer what to do next.
4. Decide and hand off. Score confidence, respect the scope of actions it is allowed to take, and escalate when it should. What the escalation carries matters as much as when it fires.
AI ticketing systems covers the routing and triage layer underneath this.
Deflection is not resolution, and the difference is what you are paying for
Deflection means the conversation ended inside the AI. Resolution means the customer's problem is solved. Those are different events, the first is much easier to produce, and the gap between them is what a per-resolution price gets charged against. What separates vendors is how a resolution is verified before it is billed.
What the two largest vendors publish about themselves.
Zendesk's AI agents page carries an FAQ line, about its paid AI Expert service, saying customers "unlock up to 80%+ automation across their support operations," with no source attached.
Zendesk's documentation is more precise than its marketing. An automated resolution is counted when a customer's issue is resolved without live-agent intervention. Conversations are evaluated at the end of the session, by default 72 hours after the first message, and those flagged resolved are verified by a large language model.
Intercom publishes a real average. Its CTO Darragh Curran wrote in March 2026 that Fin's "average resolution rate across customers has increased every month and now stands at 76%." Fin's public benchmark page shows 85%, and that figure is labelled as an average of the top 10 performers in each industry. The two numbers get quoted interchangeably. Only the 76% describes the customer base.
Fin's own documentation splits resolutions in two. A confirmed resolution is one where the customer gives affirmative feedback. An assumed resolution is one where the customer disengages for 24 hours after Fin's last answer. A customer who gave up and left is scored the same as one who was helped.
The two mechanisms differ and the difference is worth getting right. At Fin, silence is the proof. At Zendesk, silence starts the evaluation, a model then checks the conversation, and failed checks do not consume a resolution.
The same post announced Intercom is moving its pricing metric from resolutions to outcomes, meaning Fin completing the action it was configured to perform.
Where Pluno sits on this. Pluno marks a ticket resolved 72 hours after its last reply when the customer does not respond, which is the same rule Fin applies at 24 hours. The axis worth comparing on is what triggers a charge. Pluno never counts an escalation, so a ticket the AI worked on and handed over costs nothing. Ask every vendor the same question, ours included.
How Fin's per-outcome pricing works goes deeper on the billing.
How far autonomous customer service goes today
Autonomous agents go further than sceptics expect on documented questions, and nowhere near the headline forecasts on everything else.
The forecast everyone quotes. On 5 March 2025, Gartner predicted that "by 2029, agentic AI will autonomously resolve 80% of common customer service issues without human intervention, leading to a 30% reduction in operational costs." The named analyst is Daniel O'Sullivan, Senior Director Analyst.
The release publishes no methodology. No survey, no sample, no model. An analyst forecast is a legitimate thing for an analyst firm to publish and a different thing from a measurement, so pages citing the 80% as evidence are citing a forecast as data.
Gartner's own later work reads differently. Four releases since then, each with a named analyst:
-
10 September 2025, Kathy Ross: by 2028, none of the Fortune 500 will have fully eliminated human customer service, and by 2027 half of the organisations expecting significant AI-driven workforce reductions will drop those plans. AI "often struggles with exceptions and high-risk scenarios."
-
26 January 2026, Patrick Quinlan: by 2030, GenAI cost per resolution will exceed $3, higher than many offshore human agents. "Full automation will be prohibitively expensive for most organizations."
-
3 February 2026, Kathy Ross and Emily Potosky: by 2027, 50% of the companies that attributed headcount reductions to AI will rehire staff for similar functions under different job titles. Published alongside a survey of 321 customer service leaders in October 2025, which found that only 20% had reduced staffing.
-
4 August 2026, Eric Keller: 87% of customers say it is essential that a company using GenAI provides an option to reach a human agent, from a survey of 3,566 B2B and B2C customers in February and March 2026. Gartner's guidance on the same release is that service leaders should not use GenAI as a mandatory first step for every issue.
The forecasts are compatible on automation and divergent on cost. Resolving 80% of common issues is a different claim from eliminating human service, so O'Sullivan and Ross can both be right. The cost story moved inside the same firm in ten months.
Apply the same standard to the number that helps the sceptics. In June 2025, Gartner's Anushree Verma predicted that over 40% of agentic AI projects will be cancelled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls. That release publishes no methodology either. The category's two most-quoted numbers point in opposite directions and neither one shows its workings.
The one structured maturity model comes from Intercom, published in June 2026 and credited to the Fin team. Five levels by share of volume resolved by AI: manual at zero, experimenting at 10 to 25%, integrating at 40 to 60%, scaling at 60 to 80%, transforming above 80%. Useful for locating your own team, and published by a vendor that sells the destination.
What autonomous agents handle well today, and what they do not
Autonomous agents handle documented, repetitive questions and any lookup or action with a clean API. They fail on novel problems, commercial judgment calls, and anything where being confidently wrong is expensive. The dividing line is whether the answer already exists somewhere and whether the action has a clean interface.
Handled well. Documented, repetitive questions. Status and account lookups where an API exists. Password resets, entitlement changes, plan and seat adjustments. Reprocessing a failed job. And increasingly, issues a team has solved before in a ticket even when nobody wrote them up, which in B2B software tends to be a large share of the queue.
Handled badly. Problems nobody has seen before. Anything needing a commercial judgment, like a goodwill credit or a contract exception. Issues where the diagnosis needs system access the agent lacks. Regulated decisions. And anything where being confidently wrong is expensive, which is a question of consequences before difficulty.
A benchmark reality check, with its caveat attached. τ²-Bench, published in June 2025 by Barres, Dong, Ray, Si and Narasimhan, tested language models on realistic customer service tasks. In the telecom domain, on new tasks, it reports pass rates of 34% for GPT-4.1, 42% for o4-mini and 49% for Claude 3.7 Sonnet. The models tested are a generation behind, so this is not today's ceiling. What it does show is how hard these tasks are under strict scoring, worth holding next to any 80% claim.
What breaks, and how teams design around it
Autonomous agents fail in five ways a team can design around: confidently wrong answers, unbounded actions, weak escalation, missing audit trails, and a hidden route to a human.
Confidently wrong answers. The failure mode is rarely gibberish. It is a fluent, plausible, specific answer that happens to be untrue. The JourneyBench paper from Observe.AI, published in January 2026, names one mechanism precisely, parameter hallucination, where an agent uses example values from a tool's description in place of the customer's input, then calls the tool with them. We could not find a published production error rate from any named source. Benchmarks measure task failure under controlled conditions. If a page quotes a live hallucination rate, ask where it came from.
The company owns what the agent says. In Moffatt v. Air Canada, 2024 BCCRT 149, decided 14 February 2024, Air Canada's chatbot gave a passenger wrong information about bereavement fares. The airline argued it was not responsible for what its chatbot said. Tribunal member Christopher C. Rivers called that "a remarkable submission" and awarded $650.88 in damages, $812.02 with interest and fees. The bot in that case followed a script. The reasoning applies with more force to a system nobody reviews.
Action scoping. The design question is what the agent can do. What it says is the easier half. Read-only lookups are close to free. Bounded writes, like resetting a password or re-running an import, are recoverable. Anything with money or contract consequences needs a hard stop or an approval step, because a wrong action costs more to undo than a wrong sentence.
Escalation triggers. Low confidence, missing system access, and any decision no policy covers. What the escalation carries matters as much as the trigger. A transcript makes the human start from zero. The research so far, its sources and the similar tickets it found save the time the AI was meant to save.
The audit trail. When nobody reviewed the reply, the log is the only record of why the agent said what it said. Retrieval sources, actions taken, confidence at the point of decision. The first incident is a bad time to discover it was never captured.
The customer's exit. A hidden route to a human turns a technically resolved ticket into a churn risk. Gartner's guidance is that AI should not be a mandatory first step for every issue.
Ending the engineering escalation loop goes deeper on handoff design.
What to measure
Measure resolution, and measure the AI separately from the humans. Five numbers do most of the work.
Resolution rate, defined out loud. Containment counts conversations that ended inside the AI. Resolution counts problems that were solved. Write down which one your dashboard shows, because vendors use both words for both.
The billable event, in writing. Ask what silence counts as, after how long, whether a reopened ticket is refunded, and what happens to a conversation the AI worked on and escalated. The answers vary more than the headline rates do.
CSAT on AI-resolved conversations only. A blended score hides the AI's performance inside the humans'. Splitting it is half a day of reporting work and changes what you learn.
Escalation rate, and escalation quality. How often the AI hands over, and whether the human restarts from nothing when it does. The second number decides whether the deployment saves time.
Cost per resolution, all in. Per-resolution fees, seat costs, engineering time on the integrations, and hours spent on knowledge upkeep. Gartner's projection that GenAI cost per resolution will exceed $3 by 2030 is a useful check on your own model.
Where the AI layer sits
Pluno ingests resolved support tickets, which is what lets it handle the issues a help centre never documented. It is an AI support agent for complex technical tickets, and it works inside Zendesk and Intercom.
The decision it belongs to is separate from the help desk decision. The help desk decides where tickets live. The AI layer decides how many of them a person ever opens. On Zendesk and Intercom that second decision comes down to the native AI or Pluno.
The module for autonomous resolution is Deflection AI (module name pending Syed's sign-off before it becomes a link), which ingests resolved support tickets, the help centre, uploaded files and custom API integrations. It responds on email, the web widget, WhatsApp, social and Zendesk Messaging.
Escalation follows the pattern the failure-modes section argues for. When confidence is low, when system access is needed, or when a business decision is required, Pluno hands over with a research summary, ticket references and suggested next steps.
Where the ticket is a real bug, two modules carry the loop. The Troubleshooting Agent investigates across code, logs, session recordings and Sentry. The Escalation Copilot two-way syncs the ticket to Jira and Slack so the answer comes back without anyone chasing it.
Pricing is €0.90 per resolution, roughly $1 for US readers. The monthly base fee scales with ticket volume and is published on the pricing page.
Two limits worth stating. Pluno needs a back catalogue of resolved support tickets, so a team in its first months has little for it to learn from. And Deflection AI on Intercom is scoped per company. AI agents for Zendesk compared is the deeper read on the alternatives.
For the tools themselves, AI agents for customer support covers the landscape.
Frequently asked questions
What is autonomous customer service? Autonomous customer service is customer support handled end to end by software, from the first message to a resolved issue, with no person reviewing or approving each step. It does not mean no person is ever involved. Most working definitions include escalation, because an agent with no way to hand over becomes a liability.
What is an autonomous AI agent? An autonomous AI agent differs from a chatbot on two axes. It takes actions in other systems instead of only producing text, and it decides its own path through a problem instead of following a scripted tree. A chatbot answers. An agent looks something up, does something, then answers.
Is autonomous customer service the same as agentic AI? No. Agentic AI describes how the software is built: it plans, calls tools, acts and observes the result. Autonomous describes what it is permitted to do without a person approving each step. A system can be agentic in architecture and narrowly scoped in practice.
Can AI resolve support tickets without a human? Yes, for a bounded class of tickets, and the published rates need reading carefully. Intercom reports a 76% average resolution rate across its customers, while the 85% often quoted from its benchmark page averages the top 10 performers in each industry. Both vendors treat a customer who stops replying as a resolution candidate, Fin after 24 hours and Zendesk after a 72-hour session window followed by a model check. Pluno applies a 72-hour window too and bills no escalated tickets, so compare on what triggers a charge.
Will autonomous customer service replace support agents? Gartner's own research argues against it. Kathy Ross predicts that by 2028 none of the Fortune 500 will have fully eliminated human customer service, and that by 2027 half of the companies that attributed headcount cuts to AI will rehire under different titles. In Gartner's August 2026 survey of 3,566 customers, 87% said an option to reach a human is essential.
The bottom line
-
Ask what triggers a charge, beyond what the published rate counts. Silence-based resolution windows are close to universal. Whether an escalated ticket gets billed is where vendors differ.
-
Instrument resolution rate and CSAT on AI-resolved conversations separately before you scale past a pilot. A blended number hides what you need to see.
-
Scope the actions before you scope the answers. A wrong sentence is recoverable. A wrong action with money attached is not.
-
Keep the human exit visible. In Gartner's August 2026 survey, 87% of customers called it essential.
If the tickets your team spends longest on are the ones nobody ever wrote up, the AI layer is where that gets solved. See how Pluno resolves those inside Zendesk and Intercom.




