Data

Gartner: 60% of AI projects will be abandoned over data that was never AI-ready. Here's what 'AI-ready' actually means.

Gartner predicts organizations will abandon 60% of AI projects unsupported by AI-ready data. Cloudera's 2026 numbers show the Readiness Illusion: companies claim strategy while access and governance lag. Operator translation inside.

by Yerbabuena Digital·July 24, 2026·Updated: July 24, 2026·4 min read

Gartner’s 26 February 2025 release predicts that through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data. The same release is often cited for a 63% related figure — treat sample-size wording carefully when you quote it; the directional claim is clear even when secondary write-ups blur the methodology.

That forecast is not “buy a vector database.” It is a diagnosis of unreadiness dressed as adoption.

The Readiness Illusion

We name the pattern the Readiness Illusion, using Cloudera’s own Data Readiness Index 2026 (n=1,270 IT leaders, fielded 22 Jan–3 Mar 2026): 96% report AI integrated into core processes and 85% claim a clear data strategy, while nearly 80% say access still constrains AI initiatives and only 18% call their data fully governed. Companies are not failing to want AI. They are failing to notice they were never ready to feed it.

Deloitte’s State of AI readiness scores often put data and governance near the bottom of capability stacks (illustrative published framing on selected indices: data ~40%, governance ~30%, talent ~20%). The honest counterpoint: many respondents still name skills, not data, as the #1 barrier. Both can be true: strategy decks claim readiness; production agents starve for permissioned, relevant context.

IDC’s FutureScape line on productivity loss from poor data foundations (commonly cited ~15% in secondary summaries) is the cost side of the same story — treat the exact FutureScape ID as secondary unless you open the primary release.

What “AI-ready” means mechanically

Gartner’s public guidance frames a multi-step journey. In operator language:

  1. Find — can a human and an agent locate the source of truth for this workflow?
  2. Access — least privilege that still allows the job; no shared god-mode API keys.
  3. Quality — rules for completeness, freshness, and forbidden fields before retrieval.
  4. Structure for retrieval — chunks, metadata, and eval questions that mirror real work.
  5. Operate — monitors, drift, and an owner when the agent starts lying.

Your agent lies because your data lies

RAGBench and follow-on evaluations show hallucination rates that vary sharply by domain (spreads on the order of low single digits to ~20% depending on task and corpus — quote the paper’s tables, not aggregator slides). Document relevance correlates with answer quality (reported coefficients around 0.66 in RAGBench analyses). Upgrading the model without fixing retrieval is how demos impress and production disappoints.

Market context operators should notice

  • Databricks reported roughly $5.4B run-rate with strong growth in its Feb 2026 communications — do not use unverified “$6.9B June 2026” figures.
  • Snowflake continues to publish high net revenue retention (commonly ~124% in recent periods).
  • Apache Iceberg has become the default open table format conversation across Snowflake, Databricks, and S3 Tables.
  • MCP as access seam: Linux Foundation / AAIF donation (9 Dec 2025), large monthly SDK download figures, and a growing public server count — useful as plumbing adoption, not as proof your estate is governed. Open catalogs (OpenMetadata, DataHub, Unity Catalog OSS) help discovery; they do not invent owners.

Signature stack (operator estate):

LayerIf missing, what breaks
Source systems (ERP, PMS, docs)Agent invents facts
Catalog + quality rulesWrong table / stale rows
MCP / API seamFragile scrapers; vendor lock in connectors
Agent + evalsSilent quality decay

SME and Spain angle

EU Digital Decade tracking still shows SME digital intensity lagging the 90% basic-intensity ambition (Commission stats; ~73% figures appear in recent summaries). Spain’s Kit Digital has issued hundreds of thousands of grants and billions in spend (2022–2025), with the 2026 relaunch under Orden TDF/39/2026 extending support — governments are subsidising plumbing while AI-ready operations remain scarce. Published day-rates for data work often sit €800–1,500/day; readiness assessments from ~$12K appear in vendor marketing — label illustrative / single-source. Compare against our published Discovery / pilot prices rather than invented “market averages.”

Who should do what

No agents yet

Pick one workflow. Map its three source systems. Write ten eval questions from real tickets. Do not fund a lakehouse “for AI.”

Agents in pilot, answers flaky

Measure retrieval hit-rate before prompt tweaks. Fix permissions and chunking. Add human gates on external send.

Rolled back or abandoned project

Preserve failure cases as the eval set. Check whether abandonment was model theatre or missing AI-ready inputs — Gartner’s 60% framing usually points at the latter.

Predictions

  1. “AI-ready” becomes a procurement checkbox with weak definitions — demand workflow-level evidence, not logos.
  2. MCP catalogs grow faster than governed estates — seam without ownership.
  3. Iceberg + open catalogs stay the pragmatic mid-market path vs full platform lock-in.
  4. RAG spend rises (MarketsandMarkets and peers publish multi-billion trajectories) while hallucination incidents stay a board topic.
  5. SMEs that treat Kit Digital as “we did AI” will relearn the Readiness Illusion in 2027 audits.

Quiet close

We treat AI-ready data as part of the same discipline: catalog what the agent can touch, quality rules, MCP seams, and evals — before autonomy. If you want a blunt read on whether your next agent has data to stand on, current work starts with a Web & eCommerce quote or US → EMEA.


Sources

  1. Gartner — AI-ready data abandonment prediction (26 Feb 2025)
  2. Cloudera — Data Readiness Index 2026 (14 Apr 2026)
  3. RAGBench — arXiv 2407.11005
  4. Linux Foundation — AAIF / MCP
  5. Yerbabuena — data governance without the headache
  6. Yerbabuena — Scoping Crisis (agents)

Frequently asked questions

What does Gartner mean by AI-ready data?

In operator language: data that is findable, permissioned, quality-checked for the workflow, documented for retrieval, and monitored after agents touch it — not a lake 'with AI somewhere.' Gartner's public release frames abandonment risk around projects lacking that support.

Is data really the #1 barrier?

It depends which survey you read. Deloitte readiness scores often rank data and governance low while respondents still name skills as the top barrier. Treat 'data is always #1' slogans carefully; the Readiness Illusion is the gap between claimed strategy and governed access.

Why do agents hallucinate on our documents?

Retrieval quality drives answer quality. RAGBench and related work show large hallucination spreads by domain and a meaningful correlation between document relevance and answer quality. If retrieval returns the wrong chunk, a stronger model still lies confidently.

Where should an SME start?

One workflow's source systems: owner, access path, quality rule, and eval set of real questions. Skip company-wide 'AI-ready programmes.' Spain's Kit Digital funds plumbing; it does not invent a catalog for you.

#AI-ready data#data governance#Gartner#RAG#MCP#Readiness Illusion#Cloudera
Back to Insights