Skip to main content
Menu
SuperalignmentEssayAligned

Evidence audit · Evidence

6 minute read

No one has measured the AI pilot failure rate

The widely cited 95 percent figure comes from a narrower study. Here is what it measured, what it did not, and what the evidence supports.

Watercolour of a harbourmaster's office overlooking an estuary, where a clerk works through a tall stack of ledgers by the light of a green-shaded lamp.

What circulated

A preliminary result became the claim that 95 percent of AI pilots fail.

What was measured

One category of enterprise GenAI tools clearing a six-month profit-and-loss threshold.

What to keep

Most enterprise AI pilots have not reached production. That is a narrower, defensible claim.

A slide appears in many AI strategy decks: 95 percent of AI pilots fail. It sounds like a measured rate. It is not.

The figure comes from a preliminary MIT Project NANDA report. The study asked a narrower question about a specific kind of enterprise tool, over a short period, using a small convenience sample. Coverage then removed the qualifiers until the result became a claim about all AI projects.

We have an obvious conflict here. Our research is about why AI deployments stall, so a large failure rate supports our case. Our first blog post also said that "nearly nine in ten agent pilots" never reach production. We checked that number for the same reason we ask teams to inspect their own metrics: the conclusion should follow from what the instrument measured.

01 / Evidence

What the MIT report measured

The source is Project NANDA's The GenAI Divide: State of AI in Business 2025. The 26-page report labels itself preliminary and was not peer reviewed. Its evidence base included 52 structured interviews, 153 survey responses collected at four conferences, and a review of roughly 300 publicly disclosed AI initiatives. The authors describe the evidence as directional.

The 95 percent figure is the end of a particular funnel. For embedded, task-specific enterprise GenAI tools, about 60 percent of organizations evaluated a product, about 20 percent piloted one, and about 5 percent reached the report's definition of successful implementation. That definition was sustained productivity or profit-and-loss impact reported by users or executives, roughly six months after the pilot.

So the finding was not that 95 percent of AI pilots failed. It was that 95 percent of organizations in this sample did not report sustained productivity or P&L impact from this tool category on that timetable.

The same report found that general-purpose LLM tools succeeded at roughly 40 percent of organizations in the sample, while more than 90 percent of employees reported using personal AI tools for work. Its own result is closer to "embedded enterprise tools stall while individual use spreads" than "AI fails."

Project NANDA also works on agentic web infrastructure, and the report recommends systems that learn from use. That does not invalidate the data. It is relevant context when reading the framing and prescription.

02 / Evidence

How the claim widened

The report said that 5 percent of one tool category reached a stated business impact threshold within about six months. Early coverage shortened this to "95 percent of generative AI pilots are failing." Later summaries called it "MIT: 95 percent of AI projects fail."

Four things disappeared along the way:

  • the category, embedded enterprise GenAI tools;
  • the outcome, reported sustained productivity or P&L impact;
  • the time window, roughly six months;
  • the evidence limits, preliminary findings from a convenience sample.

Each summary was easier to repeat than the source. None preserved the original measurand.

03 / Evidence

The other failure rates measure different things

S&P Global Market Intelligence surveyed more than a thousand companies in early 2025. It found that 42 percent had abandoned most of their AI initiatives and that the average organization scrapped 46 percent of proofs of concept. Abandonment is a real measurement, but it is not automatically failure. Stopping cheap experiments can be a healthy portfolio decision.

Gartner's widely quoted 40 percent is a forecast that agentic AI projects will be canceled by the end of 2027. It is often quoted as an observed rate. It cannot be evaluated as an observed rate until the forecast period ends.

RAND's report is frequently summarized as finding that 80 percent of AI projects fail. The 80 percent appears as "by some estimates" and traces through earlier trade press rather than a documented measurement. RAND's original contribution was 65 interviews about why projects fail, not a new estimate of how many fail.

Vendor studies move in the other direction too. An IDC study commissioned by Lenovo put proof-of-concept-to-production at about 12 percent in early 2025. A later report in the same research line reported 46 percent. The result changed by 34 points in a year, which is a reason to inspect changes in the sample and definitions before treating the two figures as a trend.

There is also counter-evidence. A Wharton survey found 74 percent of enterprise leaders reporting positive GenAI ROI. McKinsey's 2025 State of AI found 88 percent of organizations using AI somewhere, about a third scaling it, and 10 percent or fewer scaling agents in any single function.

Those findings can all be true. An organization can get value from AI, stop many experiments, and still have few systems in production. They describe different units and outcomes.

04 / Evidence

How to read any failure statistic

Before repeating a failure rate, check four things:

  1. Study type. Separate observed measurements from forecasts and expert estimates.
  2. Unit. Organizations, initiatives, pilots, proofs of concept, and projects are different denominators.
  3. Outcome. No P&L impact, lack of scale, cancellation, and technical failure describe different events.
  4. Sponsor. A commercial interest does not make a result false, but it belongs in the interpretation.

05 / Evidence

The strongest case for the failure-rate shorthand

A skeptic can reasonably argue that the exact rate is secondary. Multiple surveys agree that many enterprise pilots stall before production, and leaders need a compact way to communicate that deployment is hard.

We agree with the underlying point. Most enterprise AI work has not crossed from pilot to scaled operation. But the shorthand changes the diagnosis. "Failure" can mean no P&L impact in six months, cancellation, lack of scale, or technical malfunction. Those conditions call for different decisions. A single rate hides the difference.

06 / Evidence

What the evidence supports

As of August 2026, we think four claims survive the source audit:

  1. Most enterprise GenAI pilots and proofs of concept have not reached production or measurable financial impact.
  2. Agentic AI remains mostly pre-production. In McKinsey's survey, 10 percent or fewer of organizations were scaling agents in any single function.
  3. Many organizations report positive AI ROI while scaling few agent systems. Value from individual tools does not imply production deployment at scale.
  4. No source we found supports the general statement that 95 percent of AI pilots fail.

Our earlier "nearly nine in ten" line rested on a vendor-commissioned branch of this evidence. We replaced it with the narrower statement that most agent pilots end before production. The argument did not depend on the larger number, and the source did not justify it.

07 / Evidence

FAQ

Is it true that 95 percent of AI pilots fail?

No source supports that claim. The figure describes organizations in a preliminary convenience sample that did not report sustained productivity or P&L impact from one category of embedded enterprise GenAI tool within about six months.

Where does the 95 percent number come from?

It comes from Project NANDA's 2025 report, The GenAI Divide. Coverage removed the tool category, outcome definition, time window, and preliminary-study qualifier.

How many AI pilots reach production?

There is no reliable general rate. Surveys use different samples, units, time windows, and definitions of production. The common finding is that most enterprise AI pilots have not reached scaled operation.

Do most AI projects fail?

That depends on the outcome being counted. Many experiments are canceled, many deployments remain unscaled, and many organizations still report useful returns from AI. A single failure label cannot distinguish those cases.

08 / Evidence

What would change our mind

A useful failure-rate study would define failure before collecting data, fix the denominator, follow a representative cohort for a stated period, publish the disposition of every pilot, and separate technical failure from cancellation, lack of scale, and lack of financial impact. Its sponsor and commercial interest should be explicit.

We have not found that study. If one exists, the useful question is not whether its headline number is higher or lower. It is whether two independent readers can reconstruct exactly what entered the denominator and exactly what counted as failure.