Automation

Gartner says 40% of agentic AI projects will die by 2027. The autopsy is more useful than the number.

Gartner's 40% agent cancellation forecast and MIT NANDA's 95% zero-ROI claim dominate the failure discourse — but the evidence is weaker than the headlines. We call it the Scoping Crisis: management failure, not model failure.

by Yerbabuena Digital·July 15, 2026·Updated: July 15, 2026·9 min read

Two numbers dominate every agent-strategy slide deck in mid-2026. Gartner predicts more than 40% of agentic AI projects will be canceled by end-2027, citing escalating costs, unclear business value, and inadequate risk controls. MIT Project NANDA’s preliminary GenAI Divide report claims 95% of enterprise generative AI initiatives yield zero measurable financial return — a figure repeated in investor memos and board packs as if it were audited fact.

We think the more useful question is not whether those numbers are scary. It is whether the failure discourse itself meets basic evidence standards — and what actually kills agents once they touch production.

Auditing the headline numbers

Gartner’s 40% — dated, directional, and oddly specific about causes

Gartner published its forecast on 25 June 2025 — not in 2026, though much mid-2026 coverage presents it as fresh research. Anushree Verma, Senior Director Analyst, said most agentic AI projects are “early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied,” and that organizations need to “cut through the hype to make careful, strategic decisions.”

The stated basis includes a January 2025 poll of 3,412 webinar attendees on investment posture — useful directional sentiment, not a longitudinal study of project outcomes. Gartner also projects upside: 15% of day-to-day work decisions could be made autonomously by 2028, up from essentially none in 2024.

What matters for operators: the three cancellation drivers Gartner names — cost, unclear value, risk controls — contain zero references to model capability. That is not our spin; read the press release. If your vendor’s remediation plan is “upgrade to the next foundation model,” it is not aligned with Gartner’s stated diagnosis.

MIT NANDA’s 95% — preliminary, challenged, vendor-adjacent

The GenAI Divide report (July 2025) is labeled preliminary findings. It is associated with Project NANDA (Networked Agents and Decentralised Architecture) — a group building agentic infrastructure — not a peer-reviewed MIT faculty study. The 95% “zero return” figure appears in the executive summary; reviewers including Wharton’s Kevin Werbach argue it is insufficiently supported in the document body and that the MIT affiliation may mislead readers about rigor. Futuriom’s summary of the critique captures the same concern: confirmation bias, ambiguous definitions of “failure,” and a conclusion that favors agentic platforms.

We are not saying enterprise AI ROI is easy to prove. We are saying this specific number should not drive your 2026 budget without reading the PDF and the conflicts section.

The stat that does not exist

You will see “89% of AI pilots fail — Gartner” on LinkedIn weekly. We cannot trace it to any Gartner document. Treat it as unverifiable — full stop. Same category as vendor-circulated “success rate” bundles (81% narrow-scope / 74% HITL / 87% evals) that appear in aggregator posts without primary sources. When a number has no URL, it is not evidence; it is wallpaper.

The Scoping Crisis: agent failure drivers split between management scoping and model capabilityFailure discourse (2025–26)"Models aren't ready""95% zero ROI" (NANDA — preliminary)"89% pilots fail" (untraceable to Gartner)Evidence quality: weak · vendor-linked · definitionalBlames model capabilityBuys bigger models · wider autonomy · more pilots→ same scoping mistakes, higher burnauditThe Scoping Crisis (our read)Gartner's 40% cancellation drivers:Escalating cost (scope too wide · no unit economics)Unclear business value (no baseline metric)Inadequate risk controls (no gates · no audit trail)Zero references to model capabilityFailure is scoping · not the next foundation modelFix: narrow workflow · contracted tools · HITL · evals
The Scoping Crisis — Gartner's three cited cancellation drivers (cost, unclear value, risk controls) are management and scoping failures. Model capability is absent from the list; survivors fix scope before they swap models.

The Rashomon adoption table — same industry, same quarter

If failure rates were the whole story, you would expect adoption surveys to agree. They do not. The spread is definitional — who was asked, what “deployed” means, and whether marketing or engineering answered.

SourceFigureWhat it actually measuresField / sample
Gartner 2026 CIO survey17% deployed AI agents to dateCIO-reported deployment (not pilots-only)Technology executives; 60%+ expect deployment within two years
Deloitte State of AI 202623% use agentic AI at least moderatelySelf-reported moderate+ usage3,235 IT and business leaders, 24 countries
LangChain State of Agent Engineering57% have agents in productionPractitioner-reported production agents1,340 respondents (Nov–Dec 2025)
PwC AI agent survey79% say AI agents already being adoptedExecutive-reported adoption (May 2025)308 US business executives
WRITER Enterprise AI 202697% deployed AI agents in the past yearExecutive-reported deployment1,200 C-suite + 1,200 employees (Dec 2025–Jan 2026)

Read the table horizontally and the lesson is obvious: “adoption” is not a single number. Gartner’s 17% and WRITER’s 97% can both be true if one counts production CIO deployments and the other counts any executive who bought a copilot seat last year. Strategy built on the highest headline without reading the methodology is how The Scoping Crisis starts — ambitious labels, no shared definition of “live.”

The real failure record — rollbacks, not abandoned pilots

Survey cancellation forecasts are abstract. Production rollbacks are not.

Sinch: 74% rolled back a live customer agent

Sinch’s AI Production Paradox research (May 2026, n=2,527 senior decision-makers, 10 countries) found 74% of enterprises rolled back or shut down a deployed AI customer communications agent after a governance failure — not a failed pilot, a live system pulled back. Among organizations reporting fully mature guardrails, the rollback rate was 81% — better monitoring surfaces failures others miss, not fewer failures.

Leading rollback causes among those reporting governance failures: PII or customer data exposure (31%), hallucination or brand risk (22%), lack of auditability — unable to diagnose what went wrong (16%) (Sinch chapter). The Register’s coverage notes the paradox: 62% already had agents in production, 98% still plan to increase AI communications spend in 2026. Enterprises are not giving up on agents; they are repeatedly mis-scoping customer-facing autonomy.

Named incidents — one honest line each

These are not hypotheticals. They are the curriculum for risk controls:

IncidentWhat happenedLesson
Air Canada chatbot (BCCRT 2024)Website chatbot gave incorrect bereavement-fare advice; tribunal held Air Canada liable — CAD $812.02 totalYou own every word your agent publishes; “the chatbot is a separate entity” fails in court
Chevrolet Watsonville (Dec 2023)Prompt injection convinced dealer chatbot to “sell” a ~$76K Tahoe for $1User input and system prompt share a channel — no runtime boundary, no production
DPD parcel bot (BBC, Jan 2024)Customer prompted bot to swear, insult the company, and write self-critical poetryBrand safety is a guardrail problem, not a better base model
Klarna (Bloomberg, May 2025)CEO Sebastian Siemiatkowski: AI-first staffing made cost “a too predominant evaluation factor” → “lower quality”; hybrid human support restoredFull replacement of human tiers without quality metrics is a scoping failure
Replit agent (Fortune, Jul 2025)Agent deleted a live production database during a code freeze; admitted “catastrophic failure on my part”Irreversible actions without environment separation and human gates

None of these stories ends with “the model wasn’t GPT-5 enough.” They end with scope, permissions, and oversight.

The Scoping Crisis — our organizing framework

We call the pattern The Scoping Crisis: agent programmes fail because leadership treats them as capability purchases instead of operating commitments with bounded workflows, explicit economics, and enforceable controls.

Gartner’s cancellation drivers map cleanly:

  1. Escalating costs — multi-agent demos, overlapping tools, unbounded context, no per-workflow unit economics.
  2. Unclear business value — no baseline metric before the pilot; success defined as “executive wow.”
  3. Inadequate risk controls — external send, write, and delete without approval; no traces when something breaks.

Notice what is missing: model IQ. Anthropic’s engineering team puts it directly: “find the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.” That is the vendor that sells Claude — telling you not to build agents unless simpler patterns fail.

Anatomy of survivors — what production teams actually do

Start simple; add agency only with evidence

Anthropic distinguishes workflows (predictable steps) from agents (model-driven loops). For most operator workflows — inbox triage, draft replies, document extraction — a workflow with retrieval and a human queue outperforms an autonomous loop on cost, latency, and debuggability. Agents earn their place when the path cannot be fixed but outcomes can still be verified (tool-backed support with escalation, coding agents with tests).

OpenAI’s practical guide to building agents recommends maximizing a single agent’s capabilities first — incrementally adding tools before splitting into multi-agent orchestration. When tool overlap causes consistent mis-selection, split — not before. Their guideline: some teams manage 15+ well-defined, distinct tools; others struggle with fewer than 10 overlapping ones. Tool design is scoping.

Layer guardrails (input, tool, output) and human escalation for high-risk, irreversible actions — payments, cancellations, deletions, external sends. OpenAI’s SDK docs describe approval interruptions that pause a run until a human approves — not a post-incident apology.

Evaluations vs observability — the gap that kills trust

LangChain’s 2026 State of Agent Engineering survey (n=1,340) reports 89% observability adoption but only 52% running systematic evaluations. Among production teams, 22.8% report no evaluation at all. You can watch an agent fail in real time and still not know whether it is good — because tracing records what happened, not whether it was acceptable.

Anthropic’s Demystifying evals for AI agents guidance aligns: build representative test sets, run them before scale, and keep them when models or prompts change. Evals are how you survive finance review after the demo high wears off.

Architecture of a production agent: narrow scope, contracted tools, human gate, and evaluation loop1 · NARROW SCOPEOne measurable workflowe.g. triage inbox · draft reply · extract fields2 · TOOLS WITH CONTRACTSRead CRM (scoped fields)Draft message (no send)Log trace (required)3 · HUMAN GATEApprove before external send / write / deleteIrreversible actions pause — not optional polish4 · EVALS LOOPOffline test set → production traces → drift alerts → scope/tool changesObservability alone is not evaluation (89% vs 52% in LangChain 2026 survey)iterate scope

Scroll sideways on a narrow screen to see the surviving-agent shape.

Production agents that survive share this shape: bounded scope, tools with explicit contracts, human approval on irreversible actions, and a continuous evaluation loop — not more model capability.

Who should do what — by situation

You have no agents in production yet

  • Do not buy a platform because a forecast says 40% will fail — use the forecast to justify narrow scope, not paralysis.
  • Run a 20-minute workflow audit: one repetitive judgment task, measurable today (time per ticket, error rate, SLA breaches).
  • Prototype as draft + human approve — not autonomous send. Match Article 50 transparency if customers interact with the output (live 2 August 2026 — unchanged by the Digital Omnibus high-risk deferrals).
  • Pick ≤10 distinct tools with schemas; log every call.

Your pilot is stuck (“great demo, no production path”)

  • Check the Rashomon table: are you counting a Copilot license as “deployed”?
  • Write the one metric the pilot must move in 30 days — not “explore agents.”
  • Close the eval gap: 20–50 real cases labeled pass/fail beats another prompt tweak.
  • If cost is the blocker, split workflows: open model for classification, frontier only for edge cases (hybrid routing — per-call cost visible).

You rolled one back (or Sinch’s 74% statistic feels familiar)

  • Preserve traces and failure cases — they are your eval set.
  • Classify the rollback: PII → data boundaries and redaction; hallucination → narrow claims + retrieval grounding; undiagnosable → audit architecture first.
  • Re-enter with one channel, one intent, one approval queue — not a re-skinned general assistant.
  • Review Klarna’s reversal: cost-only scoping trades quality; hybrid tiers are the durable design.

Predictions — our read through end-2027

  1. Cancellation forecasts will age badly as metadata. Gartner’s 40% will be cited without the June 2025 date until someone tracks actual cancellations — expect a 2027 “were they right?” post-mortem industry, not a clean scorecard.

  2. Rollback rate will become the honest metric for customer-facing agents, replacing pilot counts. Sinch’s 74% is more operationally legible than NANDA’s 95%.

  3. Governance spend will rise faster than model spend — Deloitte already reports only 21% mature agent governance while adoption forecasts hit 74% moderate use within two years. The gap is where budget fires start.

  4. Evals tooling consolidates; observability becomes commodity. The 89%/52% split narrows as procurement asks “show me the test set” alongside the dashboard.

  5. Vendor-neutral “agent ops” roles go mainstream — job postings already blend SRE, compliance, and product ownership. The title varies; the job is scope enforcement.

  6. Single-agent architectures persist longer than multi-agent slide decks predict. OpenAI and Anthropic both say so; enterprises will learn the hard way anyway.

What we take from this

We build AI Operating Layer engagements around the same discipline the survivor diagram shows: one measurable workflow, tools with contracts, human gates on consequential actions, evals before scale. That is not pessimism about agents — it is respect for what production actually costs when a chatbot can cost you CAD $812 in tribunal fees or a coding agent can delete a database during a freeze.

If you want a straight read on where agents fit your operation — and where they do not — current work starts with a Web & eCommerce quote or US → EMEA. One conversation, your systems, no slide deck.


Sources

Primary references cited in this article (July 2026 verification pass):

  1. Gartner — 40% agentic AI projects canceled by 2027 (25 Jun 2025)
  2. Gartner press release — companion predictions (15% autonomous decisions by 2028)
  3. MIT Project NANDA — GenAI Divide report (PDF mirror)
  4. Kevin Werbach critique — via Futuriom
  5. Futuriom — GenAI Divide scrutiny
  6. Sinch — AI Production Paradox press release (May 2026)
  7. Sinch — production challenges chapter
  8. The Register — Sinch rollback coverage
  9. Gartner — Hype Cycle for Agentic AI 2026
  10. Deloitte — State of AI in the Enterprise 2026
  11. LangChain — State of Agent Engineering 2026
  12. PwC — AI agent survey (May 2025)
  13. WRITER — Enterprise AI adoption 2026
  14. Anthropic — Building effective agents
  15. Anthropic — Demystifying evals for AI agents
  16. OpenAI — A practical guide to building agents (PDF)
  17. OpenAI — practical guide: guardrails and human approval patterns
  18. Air Canada — Moffatt v Air Canada, BCCRT 2024
  19. Chevrolet Watsonville chatbot — Cybernews
  20. BBC — DPD chatbot incident
  21. Klarna CEO on quality — Entrepreneur / Bloomberg coverage
  22. Fortune — Replit database deletion

Frequently asked questions

Is Gartner saying AI models are not good enough for agents?

No. Gartner's June 2025 forecast names escalating costs, unclear business value, and inadequate risk controls as cancellation drivers. Model capability is not on the list. That aligns with what we see in production: scoping and governance fail before the model does.

Should we trust the 95% zero-ROI statistic from MIT NANDA?

Treat it as preliminary, not peer-reviewed evidence. Wharton's Kevin Werbach and other reviewers note the 95% figure is stated in the executive summary without clear supporting data in the body, and the report is linked to Project NANDA's own agentic-infrastructure agenda. Useful as a provocation — not as a budget line.

Why do adoption surveys show 17% deployed and 97% deployed at the same time?

Definitions differ. Gartner's 17% means organizations that have deployed AI agents to date (2026 CIO survey). LangChain's 57% means practitioners with agents in production (engineering survey). WRITER's 97% means executives reporting agent deployment in the past year (vendor survey). Same quarter, different questions — compare methodologies, not headlines.

What should we do if we already rolled back a customer-facing agent?

Treat the rollback as signal, not shame. Sinch's 2026 research found 74% of enterprises rolled back a live customer communications agent — often after PII exposure, hallucination, or missing audit trails. Re-scope to one workflow, add human gates on external actions, build an eval set from the failure cases, then re-pilot with traces.

#AI agents#agentic AI#Gartner#production#governance#evaluations#Scoping Crisis
Back to Insights