Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsLeadership

Why Most Analytics Pilots Never Reach a Second Release

RealAIJun 20, 20238 min read
LeadershipData StrategyFMCGAnalyticsMLOps

Eight years ago I co-wrote a proposal for a European food and consumer goods manufacturer that wanted its first serious analytics capability. The scope ran across marketing, sales and operations, out along the supply chain. It was a short, cheap, deliberately small piece of work: one proving case, delivered in weeks, with the platform decision left for later.

I went back to that document this month because I keep seeing the same programme die the same way. A company runs a first analytics or machine learning pilot. The pilot works. Everyone claps. Nothing ships a second time. The model sits in a notebook, the dashboard gets a bookmark, and later a new sponsor commissions a new pilot on the same data.

What makes the old proposal worth rereading is that it named its own failure modes on the slide, in a sales document, before anyone had signed anything. Six of them. Not one was technical.

The failure list was printed in the sales deck

The method slide laid out five stages: form hypotheses and a business case, prepare data, analyse, validate, then run benefits realisation. Underneath three of them sat short lists headed as challenges to overcome.

At the hypothesis stage: lack of senior sponsorship, and no real case identified. At the data stage: data sources not available across silos, and in some cases actively protected by the stakeholders who hold them, plus a company lacking the combination of data-driven business professionals and data scientists. At the benefits stage: no discipline to run continuous benefit reporting, and no data-driven mentality to keep improving the fact base.

Two political, two organisational, two cultural. Nothing about model accuracy. Nothing about tooling. A commercial document has every incentive to be optimistic about what the buyer will manage; this one printed the buyer's likely deficits instead. Eight years on, I would change the wording and keep the list.

Score your last stalled pilot against those six. Most fail on at least three, and the missing pairing of business people with data scientists is the quietest killer: a data scientist alone produces a correct answer nobody acts on, a commercial manager alone an action nobody can defend.

Five times return on effort

The most useful number in the document is a selection bar. A case only counted as worth running if it returned roughly five times the effort spent on it.

That is a crude ratio, and its crudeness is the feature. Weighted scoring matrices, the kind with impact and feasibility on two axes and a colour scale between them, exist so that nothing ever gets rejected. Every sponsor finds their idea somewhere on the grid. A single ratio forces two honest sentences instead: what is this worth, and what will it cost us to do. Most ideas do not survive the second one, because the effort side has to include getting the data, which is the part everyone leaves out.

Applied to the machine learning backlogs I see in 2023, this bar removes most of the list: the churn model feeding a retention team with no budget to act, the demand forecast that beats the current one but arrives after the planner has committed the order. Both are good models. Neither returns five times its effort, because the decision downstream does not move.

Four to six of the six to eight weeks were the data block

The schedule in that proposal is an honest confession. One week to form hypotheses and a business case. Four to six weeks for data preparation, analysis and validation together. One week for benefits realisation. Six to eight weeks in total.

Two thirds to three quarters of the elapsed time sat in that middle block, and inside it the heaviest activities were making the data accessible and integrating and transforming it to test the hypotheses. Both were put at 90 percent on the delivery team.

We have had several technology generations since. We have managed pipelines, feature stores, orchestration tools, a whole discipline called MLOps that did not have a name when that deck was written. The share of a project that goes into getting data into a usable, described, trustworthy state has barely moved. That is the strongest argument I know for treating data readiness and lineage as the real programme and the model as the small part at the end. It is why a RealAI Platform engagement opens on the estate and its lineage, and why we quote the data weeks separately: a schedule that hides them will slip in public.

The percentages nobody argues about until it is too late

The effort planning slide is the piece I would hand to any leader commissioning a pilot today. Fourteen delivery objectives, each carrying an explicit percentage split between the delivery team and the client, written down before any work started. The slide called itself an example, which makes it more useful rather than less: this is what the split looked like when nobody was yet defending a position.

The shape of those splits encodes a philosophy. Hypothesis formation was 50/50, data crunching and transformation 90/10 towards the delivery team, analysis and validation back to 50/50. Then three rows point firmly the other way: choosing the analytics platform, 90 percent the client; providing security and technical safeguards before any data was gathered, 90 percent the client; running the internal conversation about the business case and the benefits, 80 percent the client.

That last one is the reason pilots do not reach a second release. External credibility does not transfer. An outside team can produce the number; it cannot make a commercial function change how it decides. If nobody inside the company has contracted, in writing and in advance, to own the argument, the argument does not happen, and a good result becomes a slide in an archive.

The table is not perfect, incidentally. One row, on engaging stakeholders to make key issues visible, is printed as 20 percent and 0 percent and does not sum. I quote it as it stands, because the row that got fumbled under deadline is exactly the row everyone under-specifies.

The immaturity premium was in a footnote

At the bottom of the commercial slide sat a sentence worth more than the rest of the page. A case of this kind typically runs four to six weeks for an established analytics environment. This one was quoted at six to eight, and the footnote said plainly that the duration depended on the data available and the support the client could give.

That is a data readiness premium, priced. A third to a half longer for the same work, because the estate was not ready, printed in small type under a price table, which is roughly where most organisations still keep it. Every machine learning business case I read in 2023 assumes the mature timeline and meets the immature one in month two.

The money was small on purpose: 1.3 full-time equivalents at 7,300 euros a week for the entry option, or 2.1 and 15,000 euros a week for a second case plus a sandbox platform with licensing included. Those are quoted prices in a proposal, not delivered outcomes. What matters is the design intent: a first proof of value small enough to approve without a board paper, short enough to finish, concrete enough to argue about afterwards.

Benefit tracking was a stage, not a slide

The fifth stage carried its own effort split: the delivery team took 40 percent of the ongoing monitoring, the client 60. Someone had written it down, before any work began, that measurement would continue after the report.

Almost nothing I see today does this. The benefit is booked in the business case, the pilot delivers, and nobody returns to the number. Eventually someone asks what the analytics team has produced, and there is no answer, because no one ran the measurement that would have settled it. That, more than any technical shortfall, is why the second release never gets commissioned. The same trap is open in front of every organisation running its first careful language model pilots: real enthusiasm, and almost no measurement. Enthusiasm does not survive a budget cycle.

5x
Return on effort, the bar for a case worth running
4-6 of 6-8 wks
Spent making data usable and testing on it
90%
Client share of platform choice and security safeguards
80%
Client share of selling the result internally

A pilot with no continuing measurement has no evidence to fund its successor. That is why the second release never gets commissioned.

What I would carry forward

Pick one case and hold it to the five times bar honestly, with the cost of getting the data on the effort side. Assume two thirds to three quarters of your calendar goes into preparing data, analysing it and validating the result, and plan the sponsorship conversations for the weeks when nothing visible is happening. Write down who owns platform choice, who owns the security safeguards, and above all who owns the internal argument, in percentages, before work starts. Then keep a standing benefit measurement with a named owner on both sides, because that measurement is the only thing that buys a second release.

None of this is new, which is the point. The document I have been quoting is eight years old and describes a world of data lakes and dashboards, before deployment pipelines had a discipline attached to them and well before anyone in an FMCG boardroom was asking about language models. The constraint it named never moved from the organisation chart to the model.

Six ways the work could fail were printed on the slide, and every one of them lived in the organisation rather than in the code.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?