Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesInsurance distribution

Case study
Insurance distributionAn international health insurance group

The same process takes about a day in one unit and about twenty-six in another

An international health insurance group was deciding whether to fund a standing customer insight function, and asked to see the method produce something before paying for it internally. The proof case came from a completed engagement in a retail lending business, where a mined event log ranked 141 comparable units of that organisation on three measures: steps per product, throughput time per product, and conversion. Average throughput ran from roughly one day at the fast end to roughly twenty-six at the slow one, and 52 of the 141 sat above the ten-day mark the chart splits at. The process is the constant in that picture. Everything that produced the range sits outside the process document, in local sequencing, handover and waiting, which is precisely the layer an automation programme buys without looking at. What was delivered to the group was a proposition and a proof case, not a result inside the group.

1 to 26 daysAverage throughput per unit in the proof-case benchmark
Client
An international health insurance group
Duration
Proposition and proof case, pre-engagement
AI · RIDGE E54.5 N62.5ρmax 1.00
141Comparable units ranked on one documented process
52 of 141Units above the ten-day mark on that chart
3Measures each unit was ranked on

A process document is a promise that the same thing will happen every time. One was measured against that promise across 141 comparable units of a single organisation, in one period, off one event log, with the measure taken per product so that no unit was flattered or punished by its mix. The answer came back as a range rather than a number. Average throughput at the fast end was around one day. At the slow end it was around twenty-six.

That is not variation around a process. It is what you see when the document everyone signed off is not the thing determining how long a customer waits.

The challenge

An international health insurance group was working out whether to fund a standing customer insight function, and wanted the method to have produced something somewhere before it paid for the method internally. The evidence we brought came from a completed engagement in a retail lending business, where the same question had already run from hypothesis to benchmark.

The ranking was built from a mined event log rather than from a survey of the units. Every unit was normalised per product, so that a branch handling a harder mix was not penalised for the mix, and each was then scored on three measures: how many steps a case took, how long a case took end to end, and what share of cases converted. Throughput time is the measure that produced the picture people remember. Sorted longest to shortest, one bar per unit, it descends from roughly twenty-six days to roughly one.

The chart splits at ten days. Fifty-two of the 141 units sit above the split. That is the number worth holding, because the two ends describe two units and nothing else, while 52 above the ten-day mark describes more than a third of the network. The spread is not decoration on a healthy middle. The middle is where the recoverable time is.

Now put the constant back into the sentence. One organisation. One documented process, written once, distributed to every unit, audited against in every unit. One event log, which the units can only share because they share the systems that write it. Whatever stretched a day into twenty-six, it was not the process, because the process did not vary. It sat in the part of the work that no document specifies: the order in which a unit does things it is free to order, how much of a case it completes before handing on, how appointment calendars are kept, how long a case rests between two people who are each individually doing their job correctly.

The source ranks the units. It does not say why any given unit is fast, and it would be dishonest to claim otherwise. What it establishes is where the answer cannot be, and that is a finding with real consequences, because it rules out the two responses an organisation reaches for first. Rewriting the process will not close a gap the process did not open. Buying a system will not close it either, since the ranking was mined from a log all 141 units were already writing to.

The approach

The engagement opened on a hypothesis blunt enough to come back false, about whether the distribution network actually delivered the experience the organisation had promised, and it was written with guard rails attached: throughput time down and customer effort down, while conversion and satisfaction hold. That clause is what separates an efficiency exercise from one that trades a week of waiting for a percentage point of lost business, and it also decided which of the three measures could be read on its own. None of them could.

Reading the three together is where the benchmark stops being a league table. A unit can be fast because it waits less, or fast because it does less, and the throughput ranking alone cannot tell those apart. Set steps per case beside throughput and the difference resolves: a unit low on both is running a leaner path, while a unit low on time and ordinary on steps has taken the waiting out from between the steps. Bring conversion in as the third reading and the units that got fast by cutting work they should have done identify themselves, because their conversion moves the wrong way while their throughput improves.

The improvement target that comes out of this is internal, and that matters more than it sounds. An external average invites the meeting to argue about comparability, and that argument is unwinnable because it is partly true in every organisation. A fast unit inside the same business, running the same documented process and writing to the same log, removes the argument. Somebody already does this in a day. The distance between the current distribution and that unit is a benefit estimate with a source attached rather than an assertion, and it is available before anyone trains anything.

The outcome

What was delivered to the health insurance group was a proposition and a proof case. No benchmark was run inside the group, no unit of the group was ranked, and no number inside the group moved. The distribution above belongs to the lending engagement, where it was evidence of what the method surfaces. The proposition itself was a standing function running short-cycle projects, each opened with a business case and closed with an explicit benefit claim, which is a design and not a result.

Three things would be done differently if the same benchmark opened today, and none of them changes the arithmetic.

The first is what the range means for automation. When a copilot or an agent loop is scoped, it is almost always scoped against the documented process, because the documented process is the artefact that exists and is easy to read. This chart says the documented process is running in 141 materially different ways. A harness built for the middle of the distribution lands in a unit that already runs in a day and slows it down, and lands in a unit at the slow end without touching the waiting that made it slow. Scope from the distribution, not from the document, and scope per unit.

The second is evaluation. An evaluation set assembled from the process as written will pass in the lab and predict nothing in the field, for the same reason. The set has to be sampled across the spread, with cases drawn from the fast end, the mass around the threshold and the slow tail, because a system that behaves well on the first and badly on the third has a defect the average score will hide. Building that sampling into the evaluation set is ordinary work, and it is close to the first thing our Platform team does with a client's own event history.

The third is autonomy. Graded autonomy is usually set once, per process, at the level the organisation is comfortable with on average. A spread this wide says the comfortable level is different in different places, and that a unit with a long queue between two steps may be exactly where more autonomy pays and exactly where it is riskiest. Setting the grade per unit means recording, per unit, what the system was allowed to do and what it did, which is the same per-unit record that model risk management wants and the same one that the EU AI Act, now in force, expects an operator to be able to produce. Consult engagements now open by asking whether that record can be assembled at all, because the answer sets the ceiling on how much autonomy anyone can responsibly grant.

The process was never the thing under test. It was the control variable, held constant across 141 units so that everything else could be seen, and everything else turned out to be most of the answer.

NEXT STEP

Ready to make AI real?