[insight]

Why most AI pilots never reach production

Six reasons pilots stall between demo and deployment, and what to decide before the first line of code.

6 min read · 2026-09-14

The pattern

Across financial services, telecommunications, government, and utilities the story repeats. A pilot is funded, a model is built, a demo impresses a steering committee, and then nothing happens. A year or two later the organisation has a portfolio of prototypes and not one line in the P&L it can point to.

The problem is rarely the model. It's that the pilot was designed to prove AI could work, not that a specific decision or workflow would change and move a number.

Reason one: no owner for the outcome

A pilot sponsored by a data team or an innovation function has no one whose budget improves when it succeeds. Without a business owner who reports the target metric, the pilot has no route to funding for production.

Reason two: the workflow was never designed

Models produce predictions. Value is produced when a person or system acts differently because of them. Pilots that stop at a dashboard leave the hardest part — changing how a front-line team works — for later, and later never comes.

Reason three: pilot data is not production data

Curated extracts hide the quality, latency, and access problems that appear the moment a system has to run every day. The rebuild needed to move from extract to pipeline is often larger than the pilot itself.

Reason four: economics were not modelled

Inference cost, licence cost, and the people needed to run the system are discovered at the point of scaling. A use case that looked attractive at pilot scale becomes marginal once unit economics are known.

Reason five: governance arrives late

Risk, security, and compliance teams are consulted when the pilot is finished. Controls that could have been designed in for a fraction of the cost are now retrofitted, and delivery stalls while the argument plays out.

Reason six: success was measured in pilot metrics

Accuracy, precision, and user satisfaction scores are inputs. They do not tell a CFO whether volume, margin, cost to serve, or risk events moved. A pilot that reports only model metrics has not proved anything the business needs.

What to decide first

Before a model is trained: who owns the metric, what workflow changes, which production data feeds it, what it costs per outcome at scale, which controls apply, and what result would cause the team to stop. Those six answers are the difference between a pilot and the first step of a production programme.

This is what our Find phase produces, and why our engagements are measured by P&L impact rather than pilot metrics.

Ready to move from pilots to P&L?

Tell us about the decision or workflow you want to change. We'll come back with an honest view on whether it's worth proving, and what it would take.

Start a conversation