Skip to content

Delivery

Why 95% of enterprise AI pilots never reach production, and what the 5% do differently

Most enterprise AI pilots fail for reasons that have nothing to do with the models. They are built outside production constraints, they have no evaluation, nobody owns the outcome, and the organisation around them never changes. The ones that succeed pick narrow processes, prove on real data, work with specialists, and run the system as an operation for as long as it exists.

Author
Luka Kokot, Founder and Chief Executive, Applicat AI
Published
Updated
Reading time
8 min read

The numbers

In August 2025, MIT's NANDA initiative published The GenAI Divide: State of AI in Business 2025. The headline finding: around 95% of enterprise generative AI pilots delivered no measurable impact on profit and loss, despite tens of billions of dollars of enterprise investment. A year earlier, Gartner had predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value.

Both numbers describe one phenomenon from different angles. Generative AI is easy to start and hard to finish. The MIT study also carried the most useful detail for anyone trying to be in the 5%: deployments built with specialised external firms succeeded about twice as often as internal builds, and the successful group concentrated on narrow, high-value processes and integrated deeply into workflows.

Why pilots fail: four structural causes

1. The pilot was built outside production conditions

Sample data, a sandbox, no identity integration, no audit trail, no cost profile. The demonstration works, and then every one of those omissions becomes a project of its own. By the time integration, permissions and security review are complete, the sponsor has moved on. This is the pilot trap. It is the single most common failure mode we see.

2. There was no evaluation

Without an evaluation suite built from the organisation's own cases, nobody can say whether the system is good enough, whether it got worse after a change, or whether a different model would do better. Decisions revert to opinion, and opinion does not survive a risk committee.

3. Nobody owned the outcome

A strategy firm framed it, a vendor supplied a licence, an integrator built to specification and an internal team inherited it. When performance fell short, each could point at another. Systems whose behaviour is probabilistic need a single accountable owner with the authority to change the model, the prompts, the tools and the process.

4. The organisation was not changed

MIT's study noted that the tools that stalled could not retain feedback, adapt to context or improve over time, and that adoption of generic tools was high while transformation was rare. A system that is dropped into an unchanged process, with no checkpoint design, no training and no adoption measurement, is optional. Optional systems are not used.

What the 5% do differently

  • They choose narrow processes with a number attached. A claims queue, a reconciliation, a category of customer request, with a baseline someone already reports on.
  • They prove on real data inside the perimeter. The proof is the first release, so there is no second project to reach production.
  • They evaluate before, during and after. Every decision about go-live, model choice and scope is made on measured results.
  • They design people in. Checkpoints are defined, fast and measured. Judgment stays with people; throughput moves to the system.
  • They work with specialists. Twice the success rate in MIT's sample. The frontier moves too quickly for most organisations to keep a bench current.
  • They run it as an operation. Monitoring, regression evaluation, model upgrades and monthly value reporting, for as long as the system exists.

A method that avoids the trap

The Applicat Method encodes those six behaviours into five stages: Frame, Prove, Build, Deploy and Run. Each stage has a defined output and ends with a decision made on evidence. The Prove stage in particular is built to be the first production release, which is why proofs convert. You can read the full method on the method page.

Sources

  1. MIT NANDA, The GenAI Divide: State of AI in Business 2025
  2. Gartner, "Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025", July 2024

Questions on this topic

Bring the frontier into production.

Tell us about the process you want to change. We reply within one business day.