You have seen the number. Ninety-five per cent of enterprise AI pilots fail, or deliver no return, depending on who is repeating it. It turns up in vendor decks, board papers and news coverage, usually with no source attached and always as though it settles an argument.
It comes from one document. That document is worth reading, because what it measured is narrower than the way it gets used, and because sitting a few pages from the famous number is a finding that points the other way entirely.
On 18 August 2025, hours after it went viral, it was pulled from MIT’s servers.
Where the number comes from
The source is The GenAI Divide: State of AI in Business 2025, a July 2025 paper produced in collaboration with Project NANDA out of MIT. Its own cover page describes it as “Preliminary Findings from AI Implementation Research”. The file everyone circulates is named v0.1.
The research design combined three things: a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organisations, and survey responses from 153 senior leaders collected across four major industry conferences.
Page 2 carries a disclaimer worth reading before anyone calls this an MIT study: “The views expressed in this report are solely those of the authors and reviewers and do not reflect the positions of any affiliated employers.” Its only listed reviewer is one of its own authors. It was never peer reviewed.
What it actually says
The headline sentence reads: “Despite $30 to 40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.”
Organisations. Not pilots. The report nowhere contains the sentence “95% of AI pilots fail”, which is the form the statistic almost always takes.
The distinction is not pedantry, because the paper publishes its own funnel a few pages later. Writing about enterprise-grade custom and vendor-sold tools, it reports that “sixty percent of organizations evaluated such tools, but only 20 percent reached pilot stage and just 5 percent reached production”.
Run those numbers. Of the pilots that actually started, roughly a quarter made it to production. That is a poor conversion rate and a real finding. It is not a 95 per cent pilot failure rate, which is the claim the number is used to support.
The sentence nobody quotes
Here is the finding that sits in the same document and almost never travels with the statistic.
“Generic LLM chatbots appear to show high pilot-to-implementation rates (~83%).”
Eighty-three per cent. The 95 per cent attaches specifically to bespoke and vendor-sold enterprise systems. General-purpose AI tools, in the same study, by the same authors, converted from pilot to implementation more than four times out of five.
The report is even clearer about where it locates the problem: “This divide does not seem to be driven by model quality or regulation, but seems to be determined by approach.”
So the study is not evidence that AI does not work. It is evidence that enterprise procurement of AI does not work, which is close to the opposite of how it gets deployed in an argument.
It goes further. While only 40 per cent of companies had bought an official LLM subscription, workers at over 90 per cent of the surveyed companies reported using personal AI tools for work, and the report says this unofficial usage “often delivers better ROI than formal initiatives”. The technology was already working inside these businesses. What failed was the procurement wrapped around it.
What “failure” meant
The report defines success twice, in two places, and the definitions are not the same test.
Section 8.2 gives an objective one: “Success defined as deployment beyond pilot phase with measurable KPIs. ROI impact measured 6 months post-pilot, adjusted for department size.”
Section 3.2 gives a subjective one: “We define successfully implemented for task-specific GenAI tools as ones users or executives have remarked as causing a marked and sustained productivity and/or P&L impact.”
Remarked. Nobody’s accounts were audited. The report says so itself in its limitations: “These figures are directionally accurate based on individual interviews rather than official company reporting.”
The authors also flag the window as possibly too short: “Six-month observation period may be insufficient to fully assess successful deployment for complex enterprise systems, potentially understating success rates for longer-term implementations.”
That is the authors pre-empting their own headline. Anyone citing 95 per cent as a settled failure rate is citing a number its own writers marked as possibly too pessimistic.
Then it disappeared
Fortune covered the paper on 18 August 2025 and it went viral the same day.
The archive record is precise about what happened next. The PDF was serving normally at nanda.media.mit.edu/ai_report_2025.pdf at 14:57 UTC. By 17:28 UTC the same day it returned 404. The whole subdomain now redirects to an MIT Media Lab group page, where the report is reachable only by submitting your details through a Google Form.
It appears on none of Project NANDA’s own publications pages, at MIT or on its own site. MIT’s newsroom never covered it. There has been no correction, no clarification and no republication, and we can be specific about that last point: the version archived from MIT’s own server on 18 August 2025 and the copy in circulation when we checked on 19 August 2026 are byte-for-byte identical. Only one version of this document has ever existed.
We are not suggesting anything improper. Preliminary working papers get taken down for ordinary reasons. But a number this widely quoted now rests on a v0.1 draft that its host stopped serving within three hours of it becoming famous, and that is worth knowing before putting it in a board paper.
The figures that spread were not the ones in the paper
Fortune’s article described the research as “based on 150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments”.
The document says 52 interviews and 153 senior leaders. The article has never been corrected. Checked again on 19 August 2026, the sentence still stands.
The Register published the correct figures the same day, so this was not a case of the paper being unclear. It was one outlet’s error, in the piece that made the number famous, still uncorrected a year later.
The authors have a stake in the answer
The report recommends agentic, memory-capable, interoperable systems as the fix. That is the authors’ own research programme. Project NANDA “builds on Anthropic’s Model Context Protocol (MCP) and the Google/Linux Foundation A2A to create infrastructure for distributed agent intelligence at scale”, and NANDA’s principal investigator is a named author.
This does not make the research dishonest. Plenty of good work comes from people with a view. It does mean the paper deserves the scrutiny you would give a vendor white paper, which is not the scrutiny it usually gets.
So is the direction wrong?
Probably not, and that is the frustrating part.
Anyone who has watched businesses buy AI over the past two years has seen the pattern. Enthusiasm, a subscription, a few weeks of use, no measurable change. The finding rhymes with reality closely enough that it spread, and things spread partly because they are recognisable.
The problem is not that the number was invented. It is that a preliminary, unreviewed, withdrawn paper about enterprise procurement has become a universal fact about whether AI works, quoted by people who have not opened it, to audiences who assume it describes them.
If you run a 20-person business, a statistic about custom enterprise deployments tells you very little about whether drafting your quotes automatically will save you a day a week. The study’s own answer for you is the 83 per cent, not the 95.
What to measure instead
The reason the number is so quotable is that it lets everyone skip the harder question, which is what success would have looked like and whether anybody defined it beforehand.
The failures we see are rarely mysterious. A tool gets bought before a task is chosen. Nobody writes down what the current process costs, so there is no baseline. The pilot runs on a handful of clean test cases rather than the messy real ones. Nobody owns it after launch. Three months later somebody asks whether it worked, the question turns out to be unanswerable, and it gets filed as a failure.
That is a measurement outcome rather than a technology one. It would have looked identical if the software had been perfect.
Write down the baseline before you start. How many hours, how many errors, how long a turnaround. If you cannot state the current number, you cannot state the improvement, and you will end up arguing about impressions.
Pick a window that suits the task. Six months is reasonable for an enterprise transformation programme. It is far too long for “does this draft our customer replies acceptably”, which you will know inside a fortnight.
Run it on the awkward cases. A pilot that only sees tidy inputs tells you nothing, because tidy inputs were never the problem. We wrote about ground truth and evals for the buyer’s version of this.
Decide who owns it before it launches. The most common cause of a stalled pilot is not technical. Nobody’s job description changed, so the tool became everyone’s optional extra.
Count the near misses. A use case saving four hours a week is not a failure because it did not move the annual accounts. Under the 95 per cent measure it would be one.
The point
Scepticism about AI returns is healthy and largely earned. Borrowed scepticism, resting on a number nobody has checked, is a different thing.
If you are going to cite the paper, cite what it found: that in a self-selected sample gathered largely at industry conferences in the first half of 2025, most organisations buying custom or vendor-sold enterprise AI saw no measurable return within six months, while general-purpose tools in the same study converted at around 83 per cent. That is an interesting finding, and a much better guide to what to do next. It is also considerably less useful as a conversation-ender, which is presumably why the shorter version travels further.
Frequently asked questions
Where does the 95% AI failure statistic come from?
From The GenAI Divide: State of AI in Business 2025, a July 2025 preliminary paper produced in collaboration with Project NANDA out of MIT. It drew on a review of over 300 publicly disclosed AI initiatives, structured interviews with 52 organisations and survey responses from 153 senior leaders collected at four industry conferences.
Does the study say 95% of AI pilots fail?
No. Its wording is that “95% of organizations are getting zero return”. The denominator is organisations, not pilots. The report’s own funnel shows 60 per cent evaluating enterprise tools, 20 per cent piloting and 5 per cent reaching production, which implies roughly a quarter of started pilots succeeded rather than 5 per cent.
Is the 95% figure reliable?
It is weaker evidence than its ubiquity suggests. The paper describes itself as preliminary findings, was never peer reviewed, lists one of its own authors as its only reviewer, concedes selection bias in its sample, and states that its figures are “directionally accurate based on individual interviews rather than official company reporting”. Its host stopped serving it within hours of it going viral.
Does the 95% apply to small businesses?
There is no basis to assume so. The failure rate attaches to custom and vendor-sold enterprise systems. The same study reports that generic LLM tools converted from pilot to implementation around 83 per cent of the time, which is the figure closer to a small firm automating one repetitive task.
What did the study blame for the failures?
Not the technology. In its own words, “This divide does not seem to be driven by model quality or regulation, but seems to be determined by approach.” It also found that workers at over 90 per cent of surveyed companies used personal AI tools for work, and that this unofficial usage often returned more than the formal programmes.
What is a better measure of whether an AI project worked?
A baseline recorded before you start, a review window matched to the task rather than to a reporting cycle, testing on realistic inputs rather than clean ones, and a named owner. Most projects called failures were never set up to produce an answer either way.
Flux Dynamics is a fractional CTO who builds. We define what success looks like before we build anything, because a system nobody measured is indistinguishable from a system that did not work. Tell us what you are considering.