Enterprise AI Failure Rate: Why Projects Die Between Pilot and Production


Search for the enterprise AI failure rate and you will find four different numbers, all confidently stated, none measuring the same thing. Eighty percent. Ninety-five percent. Fifty percent. Thirty-seven percent succeeding, which implies sixty-three failing.
They are all real figures from real research. They disagree because they count different populations over different windows, and none of them on its own tells you anything actionable about your own programme.
What the underlying studies do tell you is more useful. AI initiatives fail at five identifiable points, the symptoms differ at each one, the fix differs, and — the part almost nobody discusses — the cost of failing differs by an order of magnitude depending on which one you hit. This article sets out what each headline number measures, where projects die, and how to work out where yours is.
What is the enterprise AI failure rate?
There is no single enterprise AI failure rate, because published figures measure different populations over different horizons. Here is what each one counts.
| Figure | Source | Published | What it actually counts |
|---|---|---|---|
| 80%+ fail | RAND Corporation | 2024 | All AI and ML projects. RAND presents this as an existing estimate rather than its own measurement — their contribution is five root causes drawn from 65 practitioner interviews |
| 50%+ abandoned | Gartner | 2026 | Generative AI projects abandoned after proof of concept. Gartner had forecast at least 30% in July 2024; the observed outcome exceeded the forecast |
| 95% zero return | MIT Project NANDA | July 2025 | Organisations getting no measurable P&L return from generative AI, on a roughly six-month window |
| 40%+ cancelled | Gartner | June 2025 | Agentic AI projects, forecast to be cancelled by end of 2027 |
| 37% live and delivering | S&P Global | June 2026 | AI initiatives launched in the previous twelve months that reached live status and delivered value; a further 46% were on track for return within a year |
Read the labels before the percentages. The 95% counts organisations, not pilots — a company running twenty pilots where one delivers still sits inside it. The 80% covers all AI and machine learning, including work predating generative AI entirely. S&P’s 37% is the only figure that measures a defined cohort over a defined window and reports what happened to it.
One entry deserves attention. Gartner’s 30% was a forecast made in July 2024, attributed to poor data quality, inadequate risk controls, escalating costs, and unclear business value. By 2026 Gartner reported the actual figure at at least 50% — the prediction was not just met but overshot, during a period when model capability improved substantially.
That is the most important pattern in the table. Across three years of research the failure range has stayed broadly stable while models got dramatically better. Whatever is breaking these projects, it is not model quality.
The honest summary: somewhere between a third and a half of enterprise AI initiatives reach production and deliver something measurable. The rest stall — and they stall for reasons that were decidable before any code was written.
Where AI projects actually die
RAND’s 2024 study is the most useful of the group for a practical reason. Rather than counting failures, the authors interviewed 65 data scientists and engineers, each with five or more years building models in industry or academia, and asked what went wrong. Their report is subtitled Avoiding the Anti-Patterns of AI, and the five root causes they identified were:
- Stakeholders misunderstood or miscommunicated the problem to be solved, so models were optimised for the wrong metric or did not fit the business workflow.
- The organisation lacked the data needed to train an effective model.
- Teams focused on the newest technology rather than on solving the actual problem.
- The organisation lacked infrastructure to manage data and deploy models into production.
- The problem was genuinely too difficult for current AI.
Those five sit at five gates in a project’s life. Each has different symptoms, a different fix, and a different price for getting it wrong.
Gate 1 — Problem definition
The anti-pattern: a sponsor asks for “AI” for a function or department without naming a decision it should improve. The team builds something technically impressive that nobody can act on.
RAND found this to be the most common cause, and it is also the least technical. A model optimised for accuracy when the business needed recall, or a tool producing excellent output at a point in the workflow where nobody is looking, both fail this gate regardless of model quality.
The tell: ask three people from different functions what the initiative is meant to change and get three different answers. Or ask what number moves if it works and get a category rather than a metric.
The fix: name the metric, its current value, and the person in finance who already reports it, before any technical work starts.
Gate 2 — Data
The anti-pattern: the use case is clear, but the data turns out to be scattered, inconsistently labelled, insufficient in history, legally constrained, or owned by a team with no capacity to help.
Every study points here. Gartner’s Q3 2024 survey of 248 data management leaders found 63% either lacked the right data management practices for AI or were unsure whether they had them, and Gartner’s February 2025 prediction was that through 2026 organisations would abandon 60% of AI projects unsupported by AI-ready data. S&P Global’s 2026 data, from inside delivery, records data quality as a limiting factor for 38% of respondents, with gaps in data management and governance affecting 57% of initiatives moderately or severely.
The tell: the timeline slips repeatedly on data access rather than on modelling, and nobody can say definitively who owns a given dataset.
The fix: assess readiness against the specific use case on five dimensions — accessible, accurate enough for this decision, sufficient history, legally usable for this purpose, and owned by someone who can support the work. Failing any one blocks the project. An AI readiness assessment exists to establish this before money is committed rather than after.
Find out which gate your initiative is stuck at
Gate 3 — Production infrastructure
The anti-pattern: the model works in a notebook and never leaves it. RAND’s fourth root cause is the absence of infrastructure to manage data and deploy models — monitoring, retraining, versioning, integration with the systems where work actually happens.
This is where pilot purgatory lives: a portfolio of demonstrations that each proved something and none of which shipped. MIT’s 2025 study found only 5% of task-specific enterprise AI tools reached production at all, and that enterprises above $100 million in revenue commonly took nine months or longer to move from pilot to implementation, while mid-market performers did it in about 90 days. The enterprises ran more pilots and assigned more staff, and converted the fewest.
The tell: more pilots this year than last, and the same number of production systems.
The fix: decide before a pilot starts what production would require and who would own it. A pilot with no named path to production is a demonstration, and should be budgeted as one.
Gate 4 — Adoption and learning
The anti-pattern: the system ships, and usage decays. MIT identified this as the central mechanism behind the 95% and called it a learning gap: tools that do not retain feedback, adapt to workflow, or improve with use get abandoned by the people they were built for, regardless of launch-day capability.
Two findings from S&P Global’s 2026 research sharpen the problem. Trust in AI outputs has fallen rather than risen with exposure — 16% of respondents said they completely trust third-party AI models, down from 24% in 2023. And only 22% of AI projects target a fully autonomous end state; the majority sit at minimal autonomy, where AI recommends and a human decides. A system designed on the assumption that users will accept its output unquestioned is designed against both trends.
The tell: strong usage in week one, declining by week six, and the team explaining it as a training problem.
The fix: design for a sceptical user rather than trying to eliminate the scepticism. This is the standard CodeIT applies to every AI system it ships — not a checklist derived from research, but a set of constraints learned from watching adoption succeed and fail across client deployments. In practice that means four things, all decided before the first release rather than retrofitted: :
- Sources attached to claims. Every answer carries a link to the document, ticket, or record it came from, so the user can verify in one click rather than trusting or discarding wholesale.
- Explicit gaps instead of plausible invention. Where the system does not know something, it says so and leaves a marked placeholder. A confident wrong answer costs more trust than an obvious blank.
- A human decision point at the end. The system drafts, summarises, or ranks; a person sends, approves, or acts. This matches where 78% of projects actually sit and removes the largest single objection to adoption.
- Feedback instrumented from release one. Not a satisfaction survey — a record of what was accepted, edited, or discarded, so the learning gap can be closed rather than discovered six months later.
We enforce all four as non-negotiable in every delivery — not as best practice we recommend, but as a gate a build does not pass without. None of this is technically difficult. It is skipped because it is unglamorous and because it is invisible in a demo, which is precisely why systems that demo well so often fail here.
None of this is technically difficult. It is skipped because it is unglamorous and because it is invisible in a demo, which is precisely why systems that demo well so often fail here.
Gate 5 — The feasibility ceiling
The anti-pattern: the problem is genuinely too hard for current AI, and nobody says so.
This is RAND’s fifth root cause and the one most often left out of articles like this, because it is uncomfortable. Some problems are not solvable to the required standard yet: the accuracy threshold is above what current models reach, the reasoning chain is too long, the edge cases are too consequential, or the ground truth needed to evaluate the output does not exist.
The tell: performance plateaus well below the level the business case requires, and each round of work produces smaller gains than the last. Teams respond by trying newer models rather than by questioning the target.
The fix: set the accuracy or reliability threshold the business actually needs before work starts, and treat failure to reach it as information rather than as a problem to be engineered around. Discovering a feasibility ceiling in week six of a proof of concept is a good outcome. Discovering it in month fourteen of a programme is an expensive one.
What failing at each gate costs
The five gates are not equally expensive to fail at, and the ordering is worth understanding before allocating budget. The cost ordering below is a logical consequence of when each failure becomes visible, not a measured figure — but the direction is not in doubt.
| Gate | Typically discovered | What has been spent by then | What is salvageable |
|---|---|---|---|
| 1. Problem definition | Weeks 1–4, if anyone asks | Workshop time and a few sprints | Almost everything. The work redirects |
| 2. Data | Weeks 3–8 | Discovery plus early engineering | Most of it. The data work carries to the next use case |
| 3. Production infrastructure | Months 3–9 | A complete build | The model and the learning; not the timeline or the credibility |
| 4. Adoption | Months 6–12, after launch | Full build plus rollout and change management | Little. The system exists and nobody uses it |
| 5. Feasibility ceiling | Anywhere, and later than it should be | Whatever has been spent to that point | The knowledge that this class of problem is not ready |
Gartner puts the cost of deploying generative AI for business model transformation in the range of $5 million to $20 million. The difference between catching a problem at gate 1 and catching the same problem at gate 4 is most of that range.
This is the practical argument for spending real time on the first two gates. It is not diligence for its own sake — early gates are cheap to fail at, and every gate after them multiplies the cost of the same mistake.
Failing cheaply is a skill
Most writing about AI failure treats every failure as a loss. That framing produces the worst possible behaviour, which is continuing to fund something because stopping would make it a failure.
A pilot that runs for six weeks, hits a feasibility ceiling, and is stopped with a documented reason is not a failed project. It is a cheap answer to a question that would otherwise have been answered in month fourteen. The organisations in the successful minority are not the ones that avoid dead ends — they are the ones that reach them faster and leave sooner.
Three things make that possible, and all three are decided before work starts:
A stated stop condition. The result that would end the initiative, agreed and written down before anyone is invested in the outcome. Without it, every disappointing result gets reinterpreted as a reason for another iteration.
A named person who can call it. Stop conditions with no owner do not get invoked. This should be someone with the authority to reallocate the budget, not the person whose team built the thing.
A defined measurement window. MIT’s 95% rests on a roughly six-month horizon; S&P’s 46% uses twelve months. Neither is correct in the abstract. What matters is that yours is set deliberately at the start rather than inherited from whichever study someone read most recently.
Proofs of concept exist for exactly this reason: to make the expensive questions cheap to answer. A proof of concept with no stop condition is not a proof of concept — it is a build with optimistic scoping.
What the successful minority did differently
Three findings recur across the studies, and none concerns model selection.
They bought or partnered rather than building alone. In MIT’s 2025 sample, external partnerships with customised, learning-capable tools reached deployment about 67% of the time against roughly 33% for internally built tools. The authors caution that this may reflect organisational capability rather than the approach itself — companies choosing partners may differ in procurement sophistication or internal technical capacity — and that correlation does not prove causation. Read carefully, the finding is less about build versus buy than about isolation versus experience.
A worked example of the pattern: a global consumer goods company ran client and vendor support on a custom-built legacy system that required constant maintenance, could not scale with request volume, and had no automation, leaving routine cases to be handled by hand across multiple markets. The organisation already had a custom build; a better one was not the answer. CodeIT assessed the existing infrastructure, mapped case volumes and request types, and replaced the legacy system with Microsoft Dynamics 365 Customer Service and Copilot Studio, configured and integrated by an external team. The result was 50–65% of cases automated, response times 25–35% faster, and support costs down 20–30%, delivered by a team of four. Full write-up: AI support automation for a global FMCG company.
They targeted the back office, not the demo. MIT found roughly half of generative AI budget going to sales and marketing while the sharper returns appeared in back-office automation — elimination of business process outsourcing worth $2–10 million annually in customer service and document processing, a 30% reduction in external creative spend, around $1 million a year saved on outsourced risk checks at a financial services firm. Front-office gains were real but smaller. The misallocation persists because sales metrics are easier to attribute, not because the value is greater.
They decentralised sourcing while keeping accountability. MIT found the successful minority did not route use case selection through a central AI function. Budget holders and domain managers surfaced problems and led their own rollouts, paired with executive accountability rather than executive selection. S&P Global’s 2026 data points the same way from the other end: among organisations with 10,000 or more employees, 44% report a formal documented AI discipline, and S&P found a direct relationship between that discipline and project success rates and returns. Formal structure and central selection are not the same thing, and organisations that conflate them end up with a committee owning a roadmap that no domain owner wants.
How to diagnose your own programme
Work through these in order. The first one you cannot answer cleanly is the gate you are at.
- Problem. Can three people from different functions state the same metric this initiative is meant to move, and its value today?
- Data. For your top use case, can you name every system its data comes from and the person who owns each one?
- Data, again. Has anyone tested whether that data is accurate enough for this specific decision, as opposed to accurate in general?
- Production. Before the current pilot started, was there a named owner and a defined path for what production would require?
- Production, again. Do you have more pilots running than last year and the same number of production systems?
- Adoption. Is usage instrumented, and do you know what it looks like in week six rather than week one?
- Adoption, again. Does the system show its sources and flag what it does not know, or does it always produce a confident answer?
- Feasibility. Is there a stated accuracy or reliability threshold the business needs, and does the current result clear it?
- Stopping. Is there a written stop condition, a named person who can invoke it, and a date by which the outcome is assessed?
Failing questions 1 to 3 means the problem is upstream of engineering, and more delivery capacity will not help. Failing 4 or 5 means the constraint is operational rather than analytical. Failing 6 or 7 means the system shipped but was designed for a user who does not exist. Failing 8 means the target may be outside what current models reach. Failing 9 means that whatever happens, you will find out late.
Organisations that want an outside read on which gate they are at, particularly when internal teams disagree about it, often bring in AI advisory services for the diagnosis rather than the whole programme.
Stop funding pilots that cannot reach production
FAQ
There is no single figure, because published rates measure different things. RAND cites existing estimates of more than 80% for all AI and ML projects. MIT’s 2025 study found 95% of organisations getting no measurable return from generative AI over roughly six months. Gartner reported that at least 50% of generative AI projects were abandoned after proof of concept, exceeding its own 30% forecast. S&P Global’s 2026 research found 37% of initiatives launched in the previous year were live and delivering value, with another 46% on track within twelve months.
Because they count different populations over different windows. The 95% counts organisations rather than individual projects, so a company running twenty pilots with one success still appears in it. The 80% covers all AI and machine learning including pre-generative work. Gartner’s figures track generative and agentic AI specifically. S&P’s 37% is the only one measuring a defined cohort over a defined period.
Misunderstanding or miscommunicating the problem. RAND’s 2024 study, based on interviews with 65 experienced data scientists and engineers, identified this as the leading root cause: models get optimised for the wrong metric or built for a point in the workflow where the output cannot be acted on. It is also the least technical of the five causes they identified, and the cheapest to catch.
A state where an organisation accumulates successful demonstrations that never reach production. MIT’s 2025 study found only 5% of task-specific enterprise AI tools reached production, and that large enterprises typically took nine months or more to move from pilot to implementation while running more pilots than mid-market firms and converting fewer of them.
No. Gartner forecast at least 30% of generative AI projects would be abandoned after proof of concept and subsequently reported the actual figure at 50% or more, during a period of substantial improvement in model capability. Across three years of research the published range has stayed broadly consistent, which is the strongest available evidence that the constraint is not model quality but the surrounding decisions.
The cost compounds by stage. Failing at problem definition costs workshop time and a few sprints and most of the work redirects. Failing at adoption costs a full build plus rollout, and little is salvageable. Gartner puts the cost of deploying generative AI for business model transformation between $5 million and $20 million, and the difference between catching a problem early and catching the same problem after launch is most of that range.
When it hits a stop condition that was written down before work started, called by someone with authority to reallocate the budget. A pilot that runs six weeks, hits a feasibility ceiling, and is stopped with a documented reason is not a failure — it is a cheap answer to a question that would otherwise be answered much later and much more expensively.
