August 25, 20269 min readAI & Automation

Why 95% of AI projects fail, and how a small business avoids it

The short answer

MIT's NANDA research found roughly 95% of enterprise generative AI pilots deliver no measurable profit-and-loss impact, and RAND puts the broader AI project failure rate above 80%, about twice that of conventional IT projects. The causes are consistent and none of them are the model: no agreed definition of success, weak data foundations, and no integration into real workflows. Projects with quantified success metrics defined upfront succeed at 54%. Those without succeed at 12%.

Why AI projects fail and how to avoid it

What the failure numbers actually say

Five independent studies from Gartner, MIT, RAND, BCG and McKinsey converged on the same range across 2025 and 2026.

95%

of GenAI pilots deliver no measurable P&L impact (MIT NANDA)

80%+

AI project failure rate, twice that of conventional IT (RAND)

73%

of failed projects had no agreed definition of success

61%

were approved on projected ROI that was never measured after launch

Why 95% failing does not mean AI does not work

MIT defined successfully implemented strictly: sustained productivity gains and documented P&L impact, verified by both end users and executives. Most deployments are simply not measured to that standard, which means the figure captures value realisation rather than model capability.

That distinction matters, because the common reading of the statistic is that the technology does not deliver. The research says something different and more useful: most organisations underestimate the data governance and engineering rigour required to move from an impressive demo to a reliable production system, and then never check whether the result was worth the money.

Read the list of causes again. Not one item is about the model. When five separate analyses using different methodologies converge on the same conclusion, you are looking at a systemic problem in how projects are run, not an implementation detail.

The four failure patterns

No definition of done

We want to use AI is not a project. Reduce the time to answer a customer order-status question from four hours to four minutes is a project. The second can be measured; the first can only be declared successful.

Data that is not ready

Models inherit whatever mess exists underneath. If customer data sits in three systems with different spellings of the same company name, an AI layer produces confident nonsense. The audit belongs before the pilot.

Bolting AI onto a broken process

Automating a bad workflow makes it fast and bad. If a process needs six approvals because nobody trusts the data, AI does not fix that, it just produces the untrusted output faster.

Chasing visibility over value

Over half of AI budgets in 2025 went into sales and marketing pilots, which are high visibility and low ROI. MIT found the real returns came from back-office automation, which nobody demonstrates at a board meeting.

What the successful minority do differently

  • They scope tightly. One process, one measurable outcome, one team. Not a transformation programme.
  • They partner rather than build alone. MIT found pilots blending internal people with external expertise succeeded around 67% of the time, against 22% for IT-only internal builds.
  • They fix the data first. Unglamorous, undemoable, and the single largest predictor of whether anything works.
  • They capture a baseline before starting. Without it you cannot prove anything, and unproven projects get cancelled at the first budget review.
  • They accept a no-go outcome. A pilot that concludes the process should not be automated has produced a useful answer and saved the build cost.

A realistic first project for a small business

The shape matters more than the technology. Pick something high-volume, rule-heavy, and currently done by a person who dislikes doing it.

Pick the process

Order status queries. Invoice data entry. First-line support triage. Document classification. All four are repetitive, rule-bound and measurable, which is exactly what makes them suitable.

Measure it as it is today

How long does it take, how often does it go wrong, how many times a week does it happen. Write the numbers down before anything changes, because reconstructing them afterwards is guesswork.

Check the data behind it

Can a system reach the data it needs through an API or an export, and does that data agree with itself? If not, that is the project, and it is worth doing regardless of whether AI follows.

Automate one narrow slice

Not the whole process. The most repetitive step in it. A narrow automation that works beats a broad one that half works.

Measure again after six weeks

Same metrics, same method. If the numbers moved, scale. If they did not, stop and say so. Both are acceptable outcomes; carrying on without measuring is not.

From our own projects

Our clearest measurable result came from an e-commerce platform where AI generates the question bank. The number that mattered was time: work that took a full day now does not. That is the shape of a project worth doing, one process with a before and an after you can state plainly.

We have also processed 7.6 million documents and now scrape 200,000 documents a month automatically for a client CRM. That is automation rather than AI, and it is worth being precise about the difference, because a lot of what gets sold as AI is deterministic automation that would work more reliably without a model in the loop.

The scoping question we ask first is not which model. It is which process, measured how, and what happens when the system gets it wrong. Projects that cannot answer those three do not get built.

Frequently asked questions

Thinking about a first AI project?

We scope against a measurable process, capture the baseline before building, and tell you when the honest answer is that automation is the better fit.