Why AI projects fail
Everyone quotes the 95% number. Almost nobody quotes what it measured, or what its authors said about their own confidence in it. Here it is with its receipts — and then the argument the evidence supports but rarely states: AI projects don’t fail on the model. They fail because nobody in the room is paid to say “kill it.”
DRAFT COPY — AWAITING APPROVAL
01THE NUMBER EVERYONE QUOTES
It comes from one document: The GenAI Divide: State of AI in Business 2025, published July 2025 by MIT’s NANDA initiative (Challapally, Pease, Raskar and Chari). Its executive summary: “Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.” The split, it says, “does not seem to be driven by model quality or regulation, but seems to be determined by approach.”
Now the part that gets dropped. It is preliminary findings, version 0.1, not peer-reviewed, covering January to June 2025: over 300 publicly disclosed AI initiatives reviewed, structured interviews at 52 organisations, 153 survey responses.
And this is the sentence that should travel with the statistic. The report counts an implementation successful when “users or executives have remarked as causing a marked and sustained productivity and/or P&L impact.” So “95% fail” does not mean the software broke; it means that in 95% of cases nobody could point to a sustained effect on the business. Its limitations page adds that the figures are “directionally accurate based on individual company reporting” and that “success definitions may differ across organizations.” That is a weaker claim than the headline and a more useful one: not that AI doesn’t work, but that almost nobody can prove theirs did.
02THE OTHER TWO NUMBERS
Two other sources get quoted at you. In July 2024, Gartner predicted at least 30% of generative-AI projects would be abandoned after proof of concept by the end of 2025, blaming poor data quality, inadequate risk controls, escalating costs or unclear business value. In January 2026 the same firm reported that more than 50% had been abandoned after proof of concept, for those same four reasons — its own forecast was optimistic by twenty points. Note where the deaths cluster: after the proof of concept, not during it.
The most-quoted figure of all — “more than 80 percent of AI projects fail, twice the rate of IT projects that don’t involve AI,” from RAND’s 2024 report on the root causes of AI project failure (Ryseff, De Bruhl and Newberry) — needs two qualifications. RAND is repeating an outside estimate there, not a measurement of its own, and it studied machine-learning projects, explicitly excluding work that just uses pre-trained models: it is not a study of generative-AI pilots. What RAND did measure matters more — 65 interviews, five root causes, and only the fifth, applying AI to problems too hard for AI, is technical. The most common is that nobody agreed which problem was being solved.
03NOBODY IS PAID TO SAY KILL IT
Put the three sources side by side and the technology barely appears. What appears is a room in which no one has the job of stopping things.
The vendor cannot be the judge
The MIT report’s data on how executives find AI tools is the tell: roughly a fifth arrive through an existing vendor partnership, 15% through partner referrals, 13% through peer recommendations, 10% through a board member or advisor — against 6% from industry publications. Most AI projects are introduced by someone with a relationship to protect, and that someone usually writes the success criteria, runs the pilot and reports the result. Nobody in that chain is paid for the answer “this isn’t working, stop.”
No baseline means no verdict
You cannot fail a test you never set. The MIT report found roughly half of GenAI budgets going to sales and marketing while the larger savings sat in back-office automation, and named the cause plainly: the bias “reflects easier metric attribution, not actual value.” Money follows the number that is easy to claim. Ask a pilot owner what the process cost beforehand and you usually get an estimate — and an estimate cannot convict anybody.
No kill date means no kill
Gartner’s four causes of abandonment are all things you could have found in week two, and nobody looked until month nine. RAND offers the cleanest test here: be prepared to commit a team to the problem for at least a year, and if it isn’t worth that, it probably isn’t worth starting. A project with no date on which it can be declared dead does not get killed. It gets renewed, quietly, because cancelling it would embarrass someone.
04PROOF OF CONCEPT, PILOT, THEATRE
What an AI proof of concept is actually for
Three different things get called a pilot, and confusing them is how budgets disappear. An AI proof of concept answers one question: can this technology do this task at all, on our data, well enough to be interesting? Cheap, short, allowed to fail. A pilot answers a harder one: does it work inside our real workflow, with our real people, against a number we recorded before we started? Theatre is neither — a demo built to be shown, with no baseline, no decision attached and no date on which anyone will judge it.
The MIT report tracks the drop-off for task-specific tools: about 60% of organisations evaluated them, 20% reached pilot stage, 5% reached production. Very little dies at the technical stage. Almost everything dies in the gap between “it works” and “it works here, and we can prove it.”
One question separates the three: what decision does this settle, and when? If the honest answer is “it shows the board we’re doing AI,” you have theatre — and admitting that now is the cheapest thing you will ever do.
05WHAT THE MINORITY DOES DIFFERENTLY
The MIT report describes the buyers who cross what it calls the GenAI divide as acting less like software purchasers and more like clients of a business-services firm. Four habits recur: they demand customisation aligned to their own processes and data; they benchmark tools on operational outcomes rather than model benchmarks; they stay through early failures instead of restarting with a new vendor; and they source initiatives from frontline managers rather than a central lab.
The same report found externally partnered deployments reaching production around 67% of the time against roughly 33% for internal builds — while warning, to its credit, that the correlation “does not necessarily prove causation.” Its fast movers were mid-market firms averaging 90 days from pilot to full implementation, against nine months or more at large enterprises. That is not recklessness; it is what happens when fewer people can defer a decision. None of it needs a data science team. It needs someone whose job is to judge, and who loses nothing by saying no.
06HOW TO TELL IN WEEK ONE
You do not need nine months to know. Seven signals, all available in the first week, none needing technical knowledge to check.
- 1
The baseline. Ask what the number is today — hours, error rate, cost per case, days to close. If nobody measured it before the pilot started, the pilot cannot be judged, only defended.
- 2
The kill date. Ask for the date this gets stopped if the number hasn’t moved. A pilot without an end date is a subscription.
- 3
The owner of no. Name who can cancel this without asking permission, and check whether anything bad happens to them if they do. If that is the person who proposed it, you have no brake.
- 4
Who wrote the criteria. If the vendor drafted the success metrics, the vendor has already passed. Rewrite them yourself, in your own units, before the work starts.
- 5
The workflow, not the demo. Watch one real employee do one real task with it, unassisted, on a normal day. Demos run on clean data, driven by people who built the thing.
- 6
The second invoice. Ask what it costs to go from working pilot to everyday use: integration, licences, retraining, the person who maintains it. Gartner names escalating cost among the reasons projects die after proof of concept, and the escalation lives in that second number.
- 7
The year test. RAND’s question, and the fastest: would you commit a team to this problem for a year? If not, you have an experiment rather than a project — fine, as long as it is funded and named as one.
07WHERE THIS LEAVES YOU
If the checklist made you uncomfortable, the missing piece is a person, not a platform. Seven real buy-or-kill decisions are laid out on the homepage: what was proposed, what I said, and what it turned out to cost.
One live decision — a proposal, a quote, an architecture — is a Decision Review, 2 WEEKS. A whole stack read end to end is a Technical Assessment, 3–4 WEEKS; if that is the shape of it, start with the AI readiness assessment. Decisions arriving weekly are a retainer, 3-MONTH MINIMUM. Smaller companies usually want AI advice sized for a small business instead of any of those.
Fees are fixed before the work starts and invoiced 50/50; anything paid for a Decision Review is credited in full against an Assessment within 90 days. Capacity is ONE NEW ENGAGEMENT PER MONTH, next start SEPTEMBER. I take no commission, referral fee or revenue share from any vendor — which is the only reason any of the above is worth reading.
08QUESTIONS
What percentage of AI projects fail?
Nobody knows, and the three most-quoted figures measure different things. MIT’s NANDA initiative reported in The GenAI Divide (July 2025) that 95% of organisations were getting zero return — where “return” meant a marked, sustained productivity or P&L impact somebody actually remarked on, in a preliminary study of 52 organisations. Gartner reported in January 2026 that more than 50% of GenAI projects were abandoned after proof of concept. RAND’s 2024 report cites an outside estimate of more than 80% for AI projects generally. Abandoned, unmeasurable and failed are three different words; treat all three numbers as directional.
Why do generative AI pilots fail?
Rarely because the model was bad. The MIT report puts the divide down to approach rather than model quality or regulation, naming the barrier as systems that don’t learn or integrate with real workflows. Gartner attributes abandonment to poor data quality, inadequate risk controls, escalating costs and unclear business value. RAND found only one of its five root causes is technical. The pattern repeats: no baseline recorded before the pilot, no date on which it could be stopped, and nobody in the room whose job it was to stop it.
Is the 95% figure reliable?
It is a real finding from a real MIT-affiliated research group, and also a self-described preliminary version 0.1, not peer-reviewed, built on a six-month window and 52 interviewed organisations. The report warns its figures are only “directionally accurate” and that success definitions vary between organisations. I quote it because it matches what I see in the room, not because 95 is precise. A consultant who cites it without the caveats is telling you something about the consultant.
What is the difference between an AI proof of concept and a pilot?
A proof of concept asks whether the technology can do the task at all on your data — a technical question, and it should be cheap and short. A pilot asks whether it works inside your real workflow, with your real staff, against a number you recorded beforehand — a business question, and it needs a baseline and a kill date. Most of what gets called a pilot is a proof of concept given neither, which is why Gartner finds projects dying after the proof-of-concept stage rather than during it.