Skip to content

[ Field guide ]

Why many AI pilots never become operating processes

[ In short ]

An AI pilot stays a pilot when one of three things is missing: an owner with the authority to change the process, a success metric written before launch, and integration into daily work rather than a separate demo session. The most common cause is not model quality, it is the absence of one of these three things.

Published
Reading
8 min

What actually happens after a successful AI pilot?

In most cases, nothing. The pilot works in the demo, the team that built it shows it off with justified enthusiasm, and then the project enters a quiet phase where nobody officially cancels it but nobody moves it forward either. Six months later, whoever funded the pilot struggles to say precisely what happened in between.

It is not a visible technical failure. It is subtler than that: the pilot does not break, it simply never gets folded into a process someone uses daily without having to think about it. It stays a tool one person knows how to use well, instead of becoming a piece of how the company works.

How many pilots actually reach measurable impact?

Few, and it is documented. The GenAI Divide: State of AI in Business 2025, a MIT NANDA report built on interviews with dozens of companies, a survey of hundreds of leaders, and an analysis of over three hundred publicly known AI initiatives, finds that roughly 95% of enterprise generative AI initiatives produce no measurable impact on P&L. Only a minority, around 5%, converts the pilot into a verifiable operational or financial gain.

The number alone does not say much. The useful part of the report is what separates that 5%, because it is not the technology they use: they use the same models everyone else has access to.

A second independent source converges on the same point with a different methodology. Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs, or unclear business value as the causes. None of those four is model capability, and the overlap with the MIT report makes it hard to dismiss the pattern as an artifact of a single study.

If it is not model quality, what is it?

The report describes the cause as a learning gap, not a technology gap. A generic tool, used by a single person on a personal subscription, works well precisely because it is flexible: it adapts to whoever is using it. The same tool, dropped into a company process with multiple people, specific data, and operating constraints, stops adapting because it does not learn from the organization's specific context. It stays generally good and specifically useless.

This explains why a more powerful model does not fix the problem most companies are actually facing. The problem is not "the model gets it wrong too often," it is "nobody built the bridge between what the model can do and the specific way this company works."

What actually changes between a pilot and a process?

The difference is not the quality of what the pilot produces the first time. It is what happens the tenth time, the hundredth time, when the person who built it is on leave or has changed roles.

A pilot can score well on the first row and still stay a pilot on every other one.
AspectPilotProcess
OwnerWhoever built it, informallyA defined role, independent of the person
MetricDefined after the fact, if asked forWritten before launch, verifiable
UseSeparate session, on individual initiativeInside the daily workflow
ContinuityStops if the person changes rolesSurvives turnover of whoever runs it
Quality controlInformal check, when someone noticesDeclared checklist, applied every time

Why does building in-house fail more often than partnering with someone who already has?

The same report finds that projects built with a specialized external partner reach measurable impact at a significantly higher rate, roughly double, compared to projects built in-house without outside support. Not because the external partner has a better model: nobody has a better model, the models are the same for everyone.

The advantage sits in three things a team on its first attempt has not built up yet: accumulated knowledge of what works and what does not in similar contexts, the ability to integrate the solution into specific processes instead of handing over a generic tool, and a system that keeps learning from use instead of staying frozen at launch-day quality.

This does not mean building in-house is always the wrong call. It means doing it for the first time, on the first case, without ever having seen where the friction hides, almost always costs a learning cycle a partner with more iterations behind them has already worked through.

What concretely separates the 5% that succeeds?

Three things, always the same three: a defined output before construction starts, an owner with the authority to enforce the change in the process, and a budget allocated not just to build the tool but to integrate, fix, and maintain it after launch. None of the three is about the model. All three are about the organizational structure around the model.

That is why a list of tools is never the hard part of an AI project. The hard part is deciding who owns it, how you verify it works, and what happens the day after launch, once the initial enthusiasm has passed and all that is left is the work of making it run every day.

What are the early signs a pilot is about to stall?

Well before an obvious failure, there are recognizable signs a few weeks in advance.

  • Only the person who built it uses it regularly. If nobody else has adopted it, it is not a process yet.
  • Nobody can answer "how do we measure whether it is working" in one sentence. If the metric is not immediate, it probably did not exist before launch.
  • The project calendar has a launch date but no follow-up review date. Without a fixed appointment to decide whether to scale or stop, the pilot stays suspended by default.
  • The internal conversation has shifted from "what are we building" to "who maintains this," and nobody has answered that question yet.

What do you do if you recognize these signs in your own pilot?

Do not start over and do not immediately add a new tool: in most cases the pilot already technically works, and that is not where the problem sits. The first step is naming an owner with the authority to decide, not just to follow the project, even if that means temporarily taking it away from whoever originally built it. The second is writing down now, late but before going any further, the metric that should have been written at the start: what changes, by how much, measured how, by when.

The third step is the one most often skipped: setting a date, within a few weeks, when someone with decision-making power looks at that metric and explicitly decides whether the project scales, stops, or continues for another defined cycle. A pilot without that date never formally dies, and precisely because of that it keeps occupying time and attention without ever reaching a verdict. Stopping a project explicitly, with a written decision, frees up more resources than letting it quietly drift.

[ What to take away ]

  • Roughly 95% of enterprise AI initiatives produce no measurable impact: the most common cause is organizational, not technological.
  • Two independent studies, MIT NANDA and Gartner, converge on the same causes: data, controls, costs, and unclear value, never model capability.
  • Projects with a specialized external partner reach measurable impact at roughly double the rate of projects built in-house on the first attempt.
  • A pilot becomes a process when it has an owner, a metric written before launch, and integration into daily work, not when the model improves.
  • Set a review date before launch. Without it, a pilot that works still stays suspended.

How visibility inside a generative engine gets measured, written out in full. Read the article

[ Author ]

Nicola Dussin

Founder of Creaitivo. Every analysis is run directly by me.

Full profile

Which task would you like AI to handle?

Describe the work, who performs it, and how quality is checked. Within 48 hours, I will indicate whether the case suits a workflow, a prototype, or a hands-on working session with the team.

All field notes