Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

Test the Bet and Build the Exit

RealAIMar 13, 20248 min read
Delivery AssuranceProgramme RiskEnergyData ReadinessMLOps

The reflex, when you walk into a programme and find a technical decision you would not have made, is to argue about it. Reopen the selection. Ask who signed it off, and on what evidence. When the decision is already placed and half the estate is being built on top of it, that is close to the least useful move available. It spends credibility you will need later, and at the end the platform is still the platform.

I keep returning to one document for the cleanest statement of the better move.

It is an independent readiness check on a European energy retailer's business and SME division, run while that division rebuilt the sales, contracting and billing platform behind its business-customer operation on a bespoke system borrowed from an affiliated business. Four weeks of fieldwork, jointly staffed by the parent group's internal audit function and an outside assurance firm, with nineteen interviews across programme staff and steering committee members and a pass over the programme's own documentation. It produced findings and recommendations about a programme still in front of everybody: four headline findings, twelve recommendations.

Read the twelve for what is not in them. Not one asks the programme to reconsider the platform choice, revisit how it was made, or compare it against what it beat. The bet was placed. The review treated it as placed.

Two exits, for two different failures

The second and third of those recommendations look like one recommendation written twice. They are not, and the review kept them apart on purpose.

The first covers a component being wrong. The base system may turn out to be an unsound foundation for what is built on it, in which case you need measures, or a different route to the same outcome, designed before the evidence arrives rather than after. The second covers the programme being wrong: the platform is fine and it is simply not there, or not there for everyone, on the day the business needs it. One is answered by an assessment, the other by a date passing. Different triggers, different owners, different costs.

Most organisations that do any of this do the second and skip the first, because a contingency slide about being late is a familiar object and a designed alternative to a technology already being built on is not. The review also asked for the existing plan to be stressed against its unhappy flows, the piece almost nobody writes, because it means naming in advance which parts of your own plan are load bearing.

The test had a date on it

What separates this from a wish is that the assessment was given a date in the plan, and paired.

The proposed operating model was continuous quality reporting into the steering committee on the same rhythm the programme itself reported, plus a limited number of thorough assessments at moments chosen for what would be knowable then. The first was set to run in conjunction with the independent assessment of the base system, so the platform question and the delivery question would land in the same room, in the same pack, on the same day. Checkpoints later in the calendar were left explicitly marked to be determined, on the honest ground that what to examine then depended on a detailed plan that did not yet exist.

That coupling is the part I would steal. An independent evaluation arriving alone reads as a technical opinion and gets filed. The same evaluation arriving beside the programme's own status, in front of the body that controls the money, is a decision point.

Why the exit is cheap and the discovery is not

The calendar the assurance plan was drawn against had a business readiness gate on it, and the opening of the selling season two months after that. That season belonged to the market, not the programme.

Run the two branches. If the base system is found wanting at the gate, the exit design is already written and the organisation has two months of very bad weeks. If nobody looked, and it is found wanting when the season opens, there is no time at all, and the only short-term fallback came at high cost and could never deliver the business model the programme existed to create. The source gives the direction and puts no figure on either side, so take it as direction: designing the exit is work measured in weeks, and discovering you need one at the season opening is measured in a commercial year.

That asymmetry is wide enough that the exit design should not require a business case, and asking for one is usually a way of not doing it.

What the programme actually did

Three of the twelve carried an addressed mark by the time the review closed. Nine did not.

The three that moved were the approval of the plan of approach, the architecture and business design work with its roadmap, and the phased planning artefact. All three are documents, producible from material the programme already held, by people already on it, without money and without a decision from anyone senior. The nine outstanding included the independent assessment, which needed an outside party commissioned, the alternative route, which needed a budget and an admission, and the plan B, which needed someone to write down that the thing might not arrive.

That is not negligence. It is the default behaviour of any programme under pressure, predictable enough to plan around. Anything requiring a decision, an outside party or a line of budget waits, and the items that wait are disproportionately the ones that would have told you something you did not already believe. So if you want the independent test to happen, commission it in the meeting that receives the recommendation, name the party, and put its report on the steering agenda before anyone leaves the room.

3 of 12
Recommendations addressed when the review closed
9 of 12
Still outstanding, including every one attached to the platform bet
2
Separate exits recommended, one for the component and one for the programme
Two months
Between the readiness gate and the opening of the selling season

The same shape, in this year's decisions

Every organisation I speak to is placing a bet of this kind right now, faster than it placed the last one.

A model provider gets chosen for the retrieval-augmented assistant, and within a quarter the prompts, the evaluation sets, the retrieval layer and every copilot in every function are standing on it. A vector store gets chosen for one corpus and quietly becomes the store for all of them. An embedding model gets chosen once and every document in the organisation is encoded with it. None gets the scrutiny of a core system replacement, because each arrives as a pilot and the commitment accumulates without a moment where anybody signs.

I am not going to tell you to reopen those choices, for the same reason the review did not. Do the other three things.

Test the bet independently. An evaluation set built on your own data, for your own tasks, held by you, run on a schedule, with someone other than the team that chose the provider reading the result. A vendor benchmark answers the vendor's question. Yours is the only instrument that answers yours, and the only thing that will tell you whether a version change underneath you has moved behaviour on the cases you care about.

Build the component exit. Name what breaks if you have to move: whether calls go through an abstraction you control or provider-specific code scattered across services, where the prompts and evaluation sets live and whether they are portable, whether retrieval is coupled to one embedding model, what re-encoding the corpus costs in money and days. On a RealAI Platform engagement that inventory gets written down before the first provider call is wired into a service, because it is cheap to list while the code is still small and expensive to reconstruct once it is not. Then keep one alternative warm on a slice of traffic, so the answer is a measurement rather than a hope.

Build the programme exit too, and keep it separate. The system may be fine and simply not ready when the business needs it. That fallback is a human path kept staffed and skilled on real volume. It is expensive, which is why it gets quietly removed from plans shortly before it is needed. Write down the trigger and the owner for each, because an exit nobody is empowered to pull is a document, not an exit.

Reopening a placed bet costs you the argument and buys you nothing. Testing it costs weeks. Designing the exit costs weeks. Finding out you needed one on the morning the selling season opens costs the year.

One more thing is worth taking from that review. Its programme management wrote down that the plan was about fifty percent reliable, because the teams were new and no measured delivery rate existed yet. Almost nobody building assistants and early, carefully scoped autonomous experiments holds an equivalent figure, and the answer to an estimate that weak has never been a firmer date. It is a designed exit.

The regulatory version of the same structure is already visible. With the European AI Act having reached political agreement, organisations are drawing timetables against a date nobody in the room controls, on top of a stack nobody in the room has independently tested. That is the energy programme's position exactly, with a different owner of the calendar.

Test the bet. Build the exit. Do not reopen the argument, because the argument is the one part of this that has never paid.

Drawn from an independent programme readiness check on a European energy retailer's business and SME division: its headline findings, its twelve recommendations with their addressed and not-yet-addressed status, and the assurance model it proposed. That review produced findings and recommendations about a programme still in flight, not delivered results. Reading its platform recommendations as a general discipline for irreversible technical bets is ours.

Reopening a placed bet costs you the argument and buys you nothing. Testing it costs weeks. Designing the exit costs weeks. Finding out you needed one on the morning the selling season opens costs the year.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?