Macha

Why Do AI Support Pilots Get Switched Off? Five Checks That Predict a Rollback

Abbas, Customer Support & AI, Macha

Written by

Ankeet Guha, Co-founder & CTO, Macha

Reviewed by

Published September 27, 2026

Most AI support pilots that get switched off fail on data exposure, wrong answers, or a team that can't explain what went wrong, and a May 2026 Sinch survey found 74% of enterprises had rolled back a live agent. Five checks, each under an hour, tell you whether yours is heading the same way.

Key takeaways

  • AI support pilots get switched off mainly for customer data exposure, hallucination or brand risk at 22%, and inability to diagnose what went wrong at 16%, per Sinch's May 2026 survey.
  • A May 2026 Sinch survey of 2,527 senior decision makers across ten countries found that 74% had already rolled back or shut down a live AI customer communications agent.
  • The Sinch rollback rate rose to 81% among organizations with mature AI governance, even as 98% of the same respondents were increasing AI communications investment in 2026.
  • Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
  • Klarna's AI assistant handled 2.3 million chats in its first month in February 2024, and still handled about two-thirds of inquiries after Klarna began recruiting human agents again in May 2025.
Why Do AI Support Pilots Get Switched Off? Five Checks That Predict a Rollback

AI support pilots get switched off for three measured reasons: customer data exposure, cited by nearly one-third of teams that rolled an agent back, hallucination or brand risk (22%), and not being able to diagnose what went wrong (16%), according to a May 2026 Sinch survey. The headline number from that survey is real, and it's less damning than it reads. In research published on 13 May 2026, Sinch surveyed 2,527 senior decision makers across ten countries and found that 74% had already rolled back or shut down a live AI customer communications agent. The rate went up, not down, among organizations with mature governance frameworks: 81%. The same respondents were not retreating. 98% reported increasing their investment in AI communications in 2026.

Each of the five checks in this post maps to one of those causes, and each takes under an hour:

CheckWhat you doRollback cause it testsTime
1Rebuild one three-week-old conversation end to endInability to diagnose (16%)Under 1 hour
2Read 20 conversations marked resolvedFalse containment, hallucinationUnder 1 hour
3Try to reach a human as a customerNo human route, silent handoffsUnder 1 hour
4List what the agent can read and commitCustomer data exposureUnder 1 hour
5Compare the reported metric with the renewal metricUnclear business valueUnder 1 hour
PR Newswire: Sinch research reporting 74% of enterprises have rolled back live AI agents.
PR Newswire: Sinch research reporting 74% of enterprises have rolled back live AI agents.

Why do the best-governed teams roll back the most?

Read those two numbers together and the obvious interpretation falls apart. If rollback rates were a measure of incompetence, the best-governed organizations would have the lowest ones. They have the highest. Sinch's chief product officer, Daniel Morris, drew the conclusion in the coverage: "The most advanced organizations aren't failing less; they're seeing failures sooner." Gartner's Greg Carlucci made the same point from the other side: "Until you actually see it in use with the customer, there are things you'll need to adjust."

So the thing that separates a team that iterates from a team that switches the agent off permanently is not whether failures happen. It's whether the team can see a failure, attribute it, and fix the specific thing, before the conversation in the leadership meeting becomes "turn it off". Every one of the five checks below is a check on that ability.

The wider forecasts point the same way. Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, and its stated reasons are escalating costs, unclear business value and inadequate risk controls, which are three ways of saying nobody could show the number. In the same release, Anushree Verma noted that "many use cases positioned as agentic today don't require agentic implementations", and Gartner estimated only about 130 of the thousands of vendors marketing agentic AI were real, the rest being rebranded chatbots and RPA.

Gartner press release predicting over 40% of agentic AI projects canceled by end of 2027.
Gartner press release predicting over 40% of agentic AI projects canceled by end of 2027.

The much-repeated MIT figure belongs here with a caveat rather than in a headline. The NANDA report behind "95% of generative AI pilots fail" measured whether pilots produced a measurable P&L return, not whether the software worked, and it has been argued about ever since. It's a useful corrective to vendor case studies and a bad number to plan against.

Why do AI support pilots actually get switched off?

Sinch's respondents gave three leading causes. Customer data exposure was first, cited by nearly one-third, as Customer Experience Dive reported the split. Hallucination or brand risk was second at 22%. Inability to diagnose the issue was third at 16%.

That third one deserves more attention than it gets, because it is the only cause on the list that is purely self-inflicted and purely fixable. A wrong answer is a bug. A wrong answer nobody can explain is a governance problem, and governance problems get resolved by removal.

The operator accounts fill in what the survey categories flatten. A support lead posting on r/SaaS in January 2026 described the version that costs the most and shows up latest: half of tickets resolved without a human, response times down to seconds, leadership happy with the dashboard, and then net retention down 6 points at Q3 renewals. His diagnosis is the sharpest sentence in the whole cluster: "AI optimizes for the metric you give it. If you measure 'tickets closed,' it'll close tickets. It won't care if your customers are struggling."

Reddit r/SaaS: a support lead reports net retention fell 6 points after AI resolved half of tickets.
Reddit r/SaaS: a support lead reports net retention fell 6 points after AI resolved half of tickets.

The same post names a second-order effect that vendor decks leave out. His best support people started leaving, not because they were replaced, but because the job changed: every easy ticket went to the agent and every remaining ticket was an angry edge case. One of them told him, "I used to help people. Now I clean up messes." If your rollback risk assessment covers customer churn and not agent attrition, it is half a risk assessment.

There is a widely shared post on r/micro_saas titled "We lost 56 customers after I tried to save money with AI support", which describes 14 cancellations in a month and a customer who typed "I need to speak to a person" four times without reaching one. It's a good illustration and we're not going to lean on it, because a commenter in that same thread makes a specific, plausible argument that it is an astroturf post written to set up a tool recommendation. We can't verify it either way. Treat it as a story that matches the pattern rather than as data.

The pattern itself is corroborated in places nobody is selling anything. In a September 2026 thread on r/customerexperience arguing that AI-first support has failed as a product, the top reply is a user describing twenty minutes of explaining that the internet connection was broken and not the password, receiving the password-reset link each time, and finally typing AGENT in capitals until the transfer happened.

Reddit r/customerexperience: a thread arguing AI-first support has failed as a product.
Reddit r/customerexperience: a thread arguing AI-first support has failed as a product.

Which five checks predict a rollback?

Run these on a Tuesday afternoon. Each one takes under an hour, and each maps to a documented cause rather than to a general worry.

Check 1: pull one conversation from three weeks ago, end to end

Pick a conversation the agent handled three weeks ago and reconstruct it completely: the customer's question, what the agent retrieved, every tool call it made and what came back, and the reply it sent. Not a summary. The trace.

What failure looks like: you can see the reply and not the reasoning, or the logs have rolled off. This is the 16% cause, and it's the one that turns a small incident into a shutdown, because an incident you can't explain has exactly one available remedy. It is also what turned a single wrong answer into a lost day for the operator who posted "Our chatbot told a customer we do 90 day returns. We do 14": with no retained logs, the first hypothesis was a cross-tenant data leak, and it took 24 hours to rule out.

What to do: fix retention before you fix anything else, and keep query, retrieved sources, tool calls and final reply in one record with a request ID. Why your AI agent gave the wrong answer is the diagnostic that runs off that trace.

Check 2: read twenty conversations the agent marked resolved

Not the escalated ones. The resolved ones. Sample at random, then sample deliberately from your highest-risk categories, which for most teams means refunds, cancellations, billing and anything with a deadline attached.

What failure looks like: a meaningful share are what one r/customerexperience thread calls false containment. The agent answered, the customer stopped replying, the system logged a resolution, and the customer had actually given up. A commenter in that thread proposes the metric that catches it: count a resolution only if the customer did not reopen, did not contact again on the same intent within seven days, and no human had to complete the missing action. His illustrative split is 72% reported against 51% durable.

What to do: make that sampling a standing job with a name attached. A reply on Intercom's community from September 2026 lists the four methods worth copying: random sampling of resolved conversations against policy, focused review of high-risk categories, repeat-contact monitoring within a few days of an AI resolution, and CSAT compared between AI-resolved and human-handled conversations. The same reply makes the point that resolution rate can't substitute, because it only tells you the customer did not escalate. Scoring conversations with an AI judge is how to scale the reading once you know what you are looking for.

Worth naming the incentive here. A vendor billing per resolution or per deflected conversation books revenue on the moment the customer stopped replying, and the billing unit does not distinguish that from a good answer. Nobody on the vendor side is measuring your durable number for you. What deflection rate actually measures goes through how each vendor counts it, and automated resolution against deflection separates the two metrics people use interchangeably.

Check 3: try to reach a human, as a customer, on a Saturday

Open your own widget or email a support address from an address nobody recognizes, and try to get to a person. Then check what happened on the inside.

What failure looks like: two distinct failures share this check. Either the agent will not let you out, which is the complaint behind every "I typed AGENT in capitals" story, or it hands off correctly and nobody arrives. The second one is less visible and at least as damaging. A March 2026 thread on r/CustomerSuccess is titled with the finding directly: the poster looked at 25 AI support tools and found none that detect when a human goes silent after the AI escalates. The customer just sits there. The most useful reply reframes the escalation as an SLA event rather than a routing event: the moment the agent hands off, a timer starts, and if no human sends a first reply within X minutes it pings a lead or bumps priority.

What to do: make "talk to a human" a literal, always-available path, and instrument the far side of it. When to hand off to a human covers the thresholds; the part most teams skip is the alert on the handoff that nobody picked up.

Check 4: list what the agent can read and what it can commit

Write down every system the agent can query, every field it can return into a conversation, and every action it can take that changes something. Then ask what an unauthenticated person in a chat widget can get out of it.

What failure looks like: the top rollback cause in the Sinch data, cited by nearly one-third of respondents, is customer data exposure. The common shapes are an agent that can look up an account from a name or email without verifying the person, a tool that returns a whole record when the answer needed one field, and an agent that can write when it only ever needed to read.

What to do: scope the tools to the narrowest read that answers the question, put identity verification in front of anything account-specific, and route anything that moves money or changes a plan through a confirmation with a server-side rule behind it. Read and write scopes and confirmations is the pattern.

Check 5: compare the metric you report upward with the metric that predicts renewal

Put the number your leadership sees next to the number that actually moves the business: repeat contact rate on the same intent, CSAT on AI-handled conversations against human-handled ones, and churn among accounts that contacted support in the period.

What failure looks like: the r/SaaS case, where containment and response time both improved while net revenue retention fell 6 points, and nobody connected the two until renewal data arrived two quarters later.

What to do: report the durable number alongside the headline number from the start, even while it is worse, because the gap is the argument you will need when somebody proposes switching it off. The automate-first maturity model sets out which metrics belong at which stage, and the economics of AI support covers what the cost side should look like next to them.

When is AI the wrong answer for a support team?

A post about rollbacks that ended with "configure it better" would be dishonest. There are support operations where an AI agent is the wrong tool, and recognizing yours early is cheaper than a rollback.

When the volume is too small to amortize the setup. Below a few hundred tickets a month, the work of writing content, wiring lookups and sampling outputs costs more than the tickets do. The honest advice at that size is to write the five macros you keep retyping and revisit in a year.

When support volume is a product bug report in disguise. If half your tickets are one broken flow, automating the answer removes the pressure to fix the flow and hides the signal. Fix the product, then automate what remains.

When the relationship is the product. Enterprise accounts with named CSMs, high-touch onboarding and contract-specific terms do not want a faster generic answer. An agent can still draft internally, but customer-facing autonomy's the wrong shape.

When nobody will own the content. An agent's answers can't be better than the material behind them, and that material needs a person whose job includes it. If the knowledge base has no owner today, an AI agent does not create one; it just makes the gap visible to customers instead of to staff.

When the answer has to be right by law. Regulated advice, medical or financial guidance, and anything where a wrong answer creates liability belongs behind a human decision. The Air Canada tribunal case, where a company was held to a bereavement-fare policy its own chatbot invented, is the cheap version of that lesson.

What did Klarna actually roll back?

Klarna is the case everyone cites and most people cite wrongly. In February 2024 the company said its AI assistant handled 2.3 million chats in its first month, roughly two-thirds of inquiries, at resolution times under two minutes. By May 2025 it was recruiting human agents again, and the CEO said the cost focus had gone too far: "As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality."

The detail that gets dropped: Klarna's AI still handled about two-thirds of inquiries after the correction. What was rolled back was not AI. It was AI-only, and specifically the absence of a guaranteed human route. Siemiatkowski's framing was that a customer should always be able to reach a person if they want one.

Customer Experience Dive: Klarna recruiting human agents again after its AI-only push.
Customer Experience Dive: Klarna recruiting human agents again after its AI-only push.

Gartner has since put a number on how common that correction is. It expects that by 2027, 50% of companies that attributed headcount reduction to AI will rehire staff for similar functions under different job titles. The same release reports a Gartner survey of 321 customer service and support leaders in October 2025, in which only 20% had actually reduced agent staffing because of AI at all. And a separate Gartner prediction holds that by 2028, none of the Fortune 500 will have fully eliminated human customer service.

What should you do if you're mid-rollback right now?

Switching the agent off is a reasonable emergency action and a bad permanent decision, and the two get conflated in the same meeting. The useful move is to narrow rather than to remove: take the agent off the categories where it failed, keep it on the ones where the trace shows it was right, and publish the durable resolution number for both. That gives the next conversation something to argue with other than a feeling.

An agent layer on the help desk you already run is the shape that survives this, because narrowing is a configuration change rather than a migration. Macha fits teams whose tickets already live in Zendesk, Freshdesk, Gorgias, Front, HubSpot or Intercom and who want the agent working inside the ticket, with its own tool calls and its own trace, per category. It is priced per ticket, from $299 a month for 750 tickets, and setup and monitoring by the Macha team is included on every plan. It's the wrong choice if your support volume is small, if the content has no owner, or if what you need is a product fix.

Customer Experience Dive's analysis of why three-quarters of enterprises have rolled back AI agents.
Customer Experience Dive's analysis of why three-quarters of enterprises have rolled back AI agents.

How we researched this

We read the Sinch study's press release and the Customer Experience Dive analysis of the same research for the cause split, and we used a headless browser to read the three Gartner press releases directly, because gartner.com refuses scripted requests and serves a browser normally. Reddit behaves the same way, so the operator threads were read in full the same day rather than from search snippets. We have not surveyed anyone ourselves, every number here is attributed to whoever produced it, and one widely shared thread is cited with the caveat that its own commenters dispute it. All sources were checked on 2026-09-21, and the Sinch and Customer Experience Dive figures were re-checked on 24 September 2026. The Sinch press release itself doesn't publish the cause percentages; they come from Customer Experience Dive's write-up of the research, which gives data exposure as "nearly one-third" rather than an exact figure.

Frequently asked questions

Is the 74% rollback figure trustworthy? It comes from a vendor-commissioned survey, which is worth knowing, and the methodology's disclosed: 2,527 senior decision makers across ten countries and six industries, surveyed in January and February 2026, at organizations of 1,000 employees and up. Sinch sells communications infrastructure, so it benefits from the conclusion that this is hard. The internal consistency is what makes it credible anyway: a vendor pitching AI would not usually report that its best-governed customers roll back most.

Does a rollback mean the AI failed? Often it means the monitoring worked. The 81% rollback rate among organizations with mature governance is the clearest evidence in the dataset that switching something off is frequently a sign of visibility rather than of a worse agent. The failure case is the pilot that runs for two quarters with nobody able to say whether it is helping.

What single check catches the most? Check 1. If you can't reconstruct a three-week-old conversation completely, every other check produces opinions instead of findings, and you won't be able to defend the agent when somebody questions it.

How long before we should expect the numbers to look good? Longer than a pilot window, which is part of why so many get canceled. Gartner's stated cancellation causes are escalating costs, unclear business value and inadequate risk controls, and all three are what a short evaluation period produces. Decide in advance which categories you are automating and what durable resolution rate would count as success on those categories only.

Are we better off building it ourselves? Only if you're prepared to build the observability too. The clearest first-hand account in this whole cluster is an operator whose home-built stack had a real bug, a concurrency fault in the retrieval layer, that stayed invisible for weeks because nothing logged the retrieved chunks. The model is rarely the hard part. The trace, the eval set and the handoff alerting are.

What should we tell leadership when the agent gets something wrong? The trace, the category, and what changed. An incident with a named cause and a specific fix reads as a system under control. An incident with no explanation reads as a reason to remove the system, which is exactly what the 16% number in the Sinch data describes.

Does hiring back agents mean AI in support was a mistake? Klarna kept about two-thirds of its volume on AI while it rehired, so the two aren't opposites. Gartner's expectation that half of AI-attributed reductions get rehired under different job titles suggests the shape most teams land on is a smaller, more senior human team doing the work the agent hands over, which is a different job than the one that was cut.

If your tickets sit in Zendesk, Freshdesk, Gorgias, Front, HubSpot or Intercom and you want an agent you can narrow by category and trace per conversation, Macha is built for that; the pricing page shows the per-ticket arithmetic, and a trial is $50 of free usage with no card.

Sources:

Macha

About Macha

Macha is an AI agent platform that works on top of the help desk you already use — Zendesk, Freshdesk, Gorgias, or Front — and connects to the rest of your stack, even your own internal systems. Its AI agents resolve tickets and automate entire workflows end to end, all set up in plain English, no code. Learn more about Macha →

Zendesk
5.0 on Zendesk Marketplace

Loved by support teams worldwide

See what support teams are saying about Macha AI.

The application seems excellent to me! We are still testing, and we need support for some details and they were extremely efficient too!

Daniela Costa

Daniela Costa

Head of Support, Seabra

Macha has been a great addition to our support toolkit. It generates clear, well-organized responses that fit naturally into our workflow. One feature we particularly appreciate is its ability to automatically reply in the same language as the ticket.

Marius F

Marius F

Support Head, Zentana

We've been using Macha for a little while now and it's been really great addition so far! It's powerful, convenient, and makes getting work done a lot easier for our agents.

Alexander Wedén

Alexander Wedén

Head of Support

Support team is very helpful and responsive. Really enjoy how lightweight this is within Zendesk itself vs other more intrusive tools.

Cathleen Wright

Cathleen Wright

Zendesk Admin, Cortex IO

So far it's pretty good! Our queries are a little nuanced, so we can't always use it, but it's got enough utility for us. It can even incorporate our bilingual country with greetings in a second language.

Jae Oliver

Jae Oliver

Head of Support, Wise

Really enjoying using Macha, it has made a noticeable difference to our support team in a short amount of time. I really like the ticket summary feature, saves us a lot of time.

Harry Jackson

Harry Jackson

Head of Support, Crumb

Macha AI is a great addition to my workspace! It's powerful, convenient, and it really makes productivity so much easier for our agents!

Dave G

Dave G

Head of Support, Cyber Power Systems

Very impressed! AI integration for Zendesk has certainly come a long way and Macha seems to set the standard for now. This will for sure save lot of time in our support team.

Pauli Juel

Pauli Juel

Head of CS, Dokument24

Macha has been working great for us so far! The auto-responses are accurate and our resolution time has dropped significantly.

Lana T

Lana T

Zendesk Admin, Swotzy

Macha AI is a great addition. The knowledge base feature means our agents always have the right answers at their fingertips.

Mischa Wolf

Mischa Wolf

Head of Support, Topi

We're enjoying this integration so far. It's made our support team more efficient and our customers get faster responses.

Paula G

Paula G

Head of Customer Support, Xly Studio

The team enjoys using it. It saves considerable time on common questions and the integration options are excellent.

Kilian Leister

Kilian Leister

Support Head, Didriksons

Ready to supercharge your team with AI?

Get started in minutes. Connect your tools, configure your agents, and let AI handle the rest.

$50 in free credits · no time limit, no credit card