Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesInsurance change portfolio

Case study
Insurance change portfolioA European composite insurance group

Credit every funded programme in full, at the improvement its own sponsor claims, and ten of twenty-two capabilities are still short

A European composite insurance group assessed twenty-two capabilities across its retail insurance business, then did the one piece of arithmetic most capability reviews skip: it credited every funded programme in full, at the improvement its own owners expected, and measured what remained. Ten capabilities came through still short of what the strategy needed. The residual set did not cluster under one function or one budget, and part of it was not technology at all. That scatter is the finding, because it explains why programmes organised by function can each succeed on their own terms and leave the same list standing. The engagement produced a scored baseline, a shortlist of ten and a set of hypotheses to validate, not results. The method transfers directly to how an AI business case should be argued: against the baseline that exists after the funded work lands, not the one visible today.

10 of 22Capabilities still short once every funded programme is credited in full
Client
A European composite insurance group
Duration
Assessment pilot, four weeks, findings and recommendations
AI · RIDGE E31.8 N37.5ρmax 1.00
9Separate change roadmaps credited against the same capability set
~29Stakeholder interviews behind the baseline reading
2 to 12Respondents behind a single capability reading

A change portfolio is usually defended with a list of what is already funded. The list is long, the programmes are real, and the argument it carries is that the remaining distance will be closed by work under way. That argument is almost never tested, because testing it means crediting every programme in full, at the improvement its own sponsor claims for it, and then measuring what is left over. The test is unpleasant for the same reason it is useful: everything it turns up is work nobody has budgeted.

A European composite insurance group ran that test on itself. It assessed twenty-two capabilities across its retail insurance business and read each one three ways: where it stood, where the committed programmes were expected to carry it, and where the strategy needed it to be. Then it subtracted the second reading from the third. Ten capabilities came through the subtraction still short. That count, not the baseline behind it, is what the engagement actually delivered.

The challenge

The generosity of the arithmetic is the point. Every initiative in flight was credited at the level its own owners expected it to reach, not at what it had shipped. Nine separate change roadmaps were read this way, across several consumer brands and group-wide programmes, and each one was laid against the same capability set so that the credits could be added up rather than argued about brand by brand. A programme that had spent heavily and delivered little would still have carried its full expected improvement. The residual list survived that treatment.

This matters because it removes the standard objection. When a review reports a gap, the room answers that the gap is covered by a programme already running. Here the programme has been counted, at the number its sponsor put on it, before the gap was reported. Ten items came through anyway.

The second half of the finding is the shape of those ten, and it is the half that gets lost when the number is quoted on its own. The residual capabilities did not pile up under one function. They were scattered across the assessed set, which means no single sponsor and no single budget line reaches more than a fraction of them. Part of the residual set was not technology at all, which is why no amount of platform spend closes it.

That scatter explains a pattern anyone who has watched a digital or AI programme stall will recognise. Each function funds its own capability. Marketing funds customer insight, technology funds a migration, operations funds process automation. Every one of those programmes can hit its own target, be reported green, and close successfully. The residual list is unaffected, because the items on it sit between the sponsors rather than inside any one of them. Nobody failed. The list simply belonged to nobody.

The approach

The recommendation that follows from a scattered residual is not another programme. It is a single coordinated roadmap across the brands rather than nine roadmaps that each stop at their own boundary. That sounds like an organisational preference and it is really an arithmetic consequence: if the ten items sit between sponsors, then the only unit of planning that can hold them is one above every sponsor. One of the brands had already defined its own digital strategy and phased it into hundred-day increments, so the machinery for delivering in short cycles existed somewhere inside the group. What did not exist was a single view of what the increments were collectively supposed to close.

The reading itself was built from an online survey plus roughly twenty-nine stakeholder interviews across technology and the business, with two deep dives into individual consumer brands. That is a reasonable evidence base for a pilot and a thin one for a mandate, and the assessment says so. Some capability readings rest on as few as two respondents; others carry twelve. A shortlist assembled this way is a set of hypotheses about where the money should go next, and the follow-up phase was explicitly scoped to validate them and build the integrated plan, not to start building.

The same care applies to the credit itself. An expected improvement is a claim by a sponsor, not a measurement, and a portfolio credited generously will still overstate itself where the sponsors are optimistic. That cuts one way only, which is what makes the exercise safe: if the credits are too generous, the residual list is too short, and every item on it is real.

The outcome

What this phase produced was a close-out: a scored baseline, a shortlist of ten capabilities needing action beyond everything funded, external benchmarks against each of them, and a case for one coordinated plan across the brands. No capability was built in this phase and no roadmap was funded on the strength of it. The numbers are readings and expectations, not outcomes.

The method is what carries forward, and it carries directly into how AI work gets argued today. Almost every retrieval-augmented pilot and copilot business case we are asked to review is written against the baseline visible now. That is the wrong denominator. The right one is the baseline that exists after the data platform work already funded has landed, after the migration under way has reached the systems it will actually reach, after the analytics work already running finishes. Credit all of it, generously, then ask what the copilot is for. In more than one case the answer turns out to be a capability the funded work was going to deliver anyway, more slowly and with less noise.

When the residual list is drawn honestly, the items on it are rarely models. They are data readiness, lineage that an auditor can follow, and the evaluation sets that decide whether a retrieval system is answering from the right document or an adjacent one. None of those belongs to the team that wants the copilot. They sit between functions, exactly like the ten items on this insurer's list, which is why they are still open after each function has spent its budget successfully. The EU AI Act has reached political agreement, and the documentation it points towards is made of that same between-the-functions material.

Our Platform work starts on the residual set rather than the demonstration, because a pilot built on top of unclaimed data readiness produces a convincing meeting and nothing that survives contact with an audit. Early autonomous experiments are worth running, carefully scoped, but they belong after the arithmetic and not instead of it.

Ten of twenty-two is a small enough number to fund and a scattered enough set to be nobody's job. Both halves of that sentence are the finding, and organisations act on the first half and get defeated by the second.

NEXT STEP

Ready to make AI real?