The pilot worked.

Everyone agreed it worked. There was a demo, the room was impressed, somebody said this changes things. Then a quarter passed, then another, and it is still a pilot. Nobody killed it. It just never became the way the work gets done.

This is the most common outcome in business AI, and it is worth being precise about why, because the reasons are decided long before the pilot starts.

The size of the problem

The July 2025 research from Project NANDA at MIT publishes a funnel for enterprise-grade AI tools. Sixty per cent of organisations evaluated such tools, 20% reached pilot stage, and 5% reached production.

Read that as a conversion rate rather than a headline. Of the organisations that got as far as running a pilot, roughly a quarter made it to production. Three in four stalled.

The same report is specific that this is not a technology problem:

This divide does not seem to be driven by model quality or regulation, but seems to be determined by approach.

And the detail that settles it. While only 40% of the companies surveyed had bought an official AI subscription, workers at over 90% of them were already using personal AI tools for their job, with the report noting that this unofficial usage “often delivers better ROI than formal initiatives”.

The technology was working in those businesses the whole time. It was the formal programmes that stalled. We looked at what that study actually measured separately, because the popular version of it says something quite different.

The question most pilots are built to answer

Here is the underlying mistake, and it is easy to make.

Most pilots are designed to answer “does the technology work”. That question was settled before you started. The models are good, the demo will succeed, and you will finish the pilot knowing something you already knew.

The question that decides whether it reaches production is different. Will this survive contact with the business? That covers the messy inputs, the person whose job changes, the security review nobody scheduled, and the system it has to talk to that belongs to a supplier who is not returning calls.

A pilot that only tests the first question produces a confident yes and no route forward.

The five things that decide it

It ran on the good cases. Somebody assembled the pilot data, and being helpful, they assembled the clean examples. Production is the customer who replied to the wrong email thread, the invoice with three PDFs stapled together, and the enquiry written in a hurry from a phone. If the pilot never saw those, it has not been tested.

No route to production was agreed before it began. Who signs it off. What security review it has to pass. Which system it needs to write to, and who owns that system. These take weeks and they are sequential, so discovering them at the end adds a quarter. Decided at the start they cost nothing, because they run alongside the build.

Nobody’s objectives changed. The person doing the task manually is the person who has to stop, and unless somebody has changed what they are measured on, the old process keeps running in parallel. Parallel running is comfortable and it never ends on its own.

There was no baseline. Nobody recorded what the process cost before the pilot started, so when the review meeting asks whether it worked, the honest answer is that the question cannot be answered. Projects rarely die from a bad number. They die from an unanswerable one, because in the absence of evidence the safe decision is to keep piloting. We set out why the acceptance criterion should be a number agreed in advance.

The pilot budget was the whole budget. The money ran out at exactly the point where the work becomes valuable. Production means integration, permissions, error handling, someone to call when it misbehaves, and a plan for what happens when the model provider changes something. If none of that was funded, the pilot was always going to be the end of it.

What to decide before you start

None of this requires a bigger pilot. It requires four decisions taken at the front rather than the back.

What has to be true for this to go live. Write the list on day one. Security sign-off, data agreements, who owns the system it integrates with, what the acceptance number is. It is usually shorter than people fear, and every item on it has a lead time.

Which real cases it has to handle. Not a sample of good ones. Deliberately include the awkward examples, because those are the ones that decide whether it survives.

Whose week changes, and what they get out of it. If you cannot name the person, adoption has not been planned. If naming them makes you uncomfortable, that is the conversation to have before the build rather than after.

What production costs, roughly. You do not need a precise figure. You need to know whether it is a similar order to the pilot or ten times it, because that determines whether there is any point running the pilot at all.

The obvious objection

Is a pilot not supposed to be cheap and disposable? Doesn’t all this front-loading defeat the purpose?

It would, if the thing being made disposable were the decisions. It is not. Throw away the code by all means, that is what a pilot is for, and current tooling makes rebuilding the cheap half straightforward. What should never be disposable is the understanding of what the process actually is, what it costs today, and what has to be true for it to change.

There is a second objection worth taking seriously. Sometimes a pilot stalls because it should. The process turned out not to matter, the volume was too low, the value was smaller than expected. That is a good outcome, and a pilot that produces a clear no has done its job. The failures worth worrying about are the ones where the answer was yes and nothing happened anyway.

What good looks like

A pilot that reaches production usually looks slightly disappointing at the demo.

It handles a narrower problem than the impressive version. It has been run against the difficult cases, so its accuracy number is lower than a curated demo would produce and considerably more believable. Somebody in the business has already agreed to change how they work. The security conversation happened in week two rather than week twelve.

That version rarely gets the room excited. It is also the one still running a year later, which is the only test that has ever mattered.

Frequently asked questions

What percentage of AI pilots reach production?

The July 2025 Project NANDA research reported a funnel for enterprise-grade AI tools of 60% of organisations evaluating, 20% piloting and 5% reaching production. Expressed as a conversion rate, roughly a quarter of pilots that started reached production.

Why do AI pilots fail even when the technology works?

Because the pilot usually tests whether the model performs, which was rarely in doubt, rather than whether the result survives contact with the business. Integration, sign-off, messy real inputs and someone changing how they work are what decide it, and none of those are tested by a demo.

What should be agreed before an AI pilot starts?

What has to be true for it to go live, which real cases it must handle, whose job changes, and roughly what production will cost. Each of those has a lead time, so discovering them at the end adds months.

Is it a problem if a pilot does not go into production?

Not always. A pilot that produces a clear no has done its job, and stopping is the right call when the process turns out not to matter enough. The expensive failures are the ones where everyone agreed it worked and nothing changed.

How big should an AI pilot be?

Small enough to finish, and narrow enough that one process is properly covered rather than three partially. The constraint that matters is not size but whether the awkward real cases are included, since those decide whether it holds up in production.


Flux Dynamics is a fractional CTO who builds. We agree what has to be true for something to go live before the build starts, because a pilot that cannot become production was never worth running. Tell us what you are considering.

Flux Dynamics
Software & AI Consultancy

Flux Dynamics is a UK software and AI consultancy: a fractional CTO who also builds, shipping custom web applications and software for businesses.