A capability review of a European composite insurance group put a number on the group's own flagship programme, and the number was 2.9 out of 5.
That score is not where the insurer stood on straight-through processing at the time of the review. It is where the assessors expected it to stand afterwards, once the roll-out of the core policy and claims platform had landed in full. Everything the programme was going to deliver was already priced into the number. The strategy that commissioned the review had asked this same capability to reach 5.
The challenge
The group's direction was stated plainly. Compete on cost, standardise, consolidate lines of business, prepare parts of the operation to be outsourced. Read against that direction, most of the graded capabilities were assigned a target of 4. A short list was assigned a target of 5, and straight-through processing was on it. The strategy said in as many words that the focus belonged on straight-through processing and standardisation.
The programme meant to deliver it was the core platform roll-out. Nobody in the review disputed that it worked. It was producing high maturity across a number of lines of business. Selected products already ran fully automated commercial and technical acceptance with no person in the path at all. Consumer and non-complex products could be bought and fulfilled online with no paper trail and no waiting step. By any ordinary definition of automation, parts of the estate had arrived.
And the planned score still came out at 2.9, on a scale where 3 sits in the middle. A flagship programme, credited in full, moving a capability the strategy had put in its top band to just under the midpoint.
The same review ran the arithmetic across everything else in flight, and expressed each programme's expected contribution in maturity points rather than in adjectives, banded from under half a point to more than two. The largest expected step change in the whole portfolio landed on breaking products into configurable components. The smallest landed on straight-through processing. Both of those capabilities had been assigned the same target of 5. The portfolio was set to deliver its biggest step on one of them and its smallest on the other.
The reason sits in one respondent comment, and it is the most useful sentence in the pack. Fully automated acceptance already existed on selected products. What was slow was changing the pricing and the acceptance criteria behind that acceptance: complex, and demanding too much lead time. The constraint was never the decision. It was the rule change behind the decision.
The assessors added a second constraint in their own voice, and it is the one most programmes never write down: the capability still required a great deal of craftsmanship from employees. Underwriting judgement that lived in people rather than in anything a platform could read.
One more thing about the method, because it is rarer than it should be. Three to four people answered on each sub-capability in this part of the review, and the review printed that count beside every bar it drew. It told the reader exactly how much weight the number would carry.
The approach
This was an assessment. What it produced was a graded position, a set of gaps and a list of recommended actions. No system was built and no process was rewired inside its scope, and the piece is worth reading for how it reached its finding rather than for a result it did not claim.
Three moves did the work.
The first was to grade the planned position separately from the current one. Programme reporting almost always compares the planned end state to today, which flatters every programme ever funded, because the baseline was chosen by the people who wrote the business case. This review compared the planned end state to the level the strategy demanded, and the distance between them stayed on the page after the programme had been given full credit.
The second was to express what each live initiative would actually add, in points on the same scale, banded. That turns a portfolio review from a list of things being described into a list of things being measured. It is also what exposed the mismatch: the two capabilities at either end of that range carried the same top target, and the portfolio's largest expected gain landed on one of them while its smallest landed on the other. Equal ambition on paper, and nothing like equal movement behind it.
The third was to name the constraint rather than the symptom. The batch-oriented legacy underneath parts of the estate was going to outlive the migration, and no programme in the portfolio was going to reach it. The recommendation followed from that: extend automation into the legacy environments with a process orchestration layer above them, instead of waiting for every line of business to arrive on the new platform. That accepts the legacy as a standing condition rather than as a phase that ends.
The outcome
What was handed over was a finding and a set of recommended actions. The 2.9 is a forecast about a plan, not a measurement of a result, and it should be read that way. Nothing on those pages had run yet.
The durable part is where the constraint was found, because that answer has aged better than anything else in the deck. The transaction had been automated. The path by which a pricing table or an acceptance rule changes had not been: the request, the impact analysis, the test, the sign-off, the release. That path was manual, slow, and absent from the business case of the programme everyone pointed to. A group can automate every quote it issues and still be capped at 2.9, because the thing it cannot do quickly is change what a quote means.
That path is where an operating platform earns its keep, and it is the first place we point one. Acceptance criteria and pricing rules held as data rather than compiled into a system. A change proposed against a versioned rule set, evaluated on held-out history, replayed against the quotes it would have altered, with lineage back to every one of them, and released behind a switch that can be closed again. Machine learning belongs on the same path rather than beside it, because a rating factor and an acceptance threshold are the same object under different names, and both need the same deployment discipline, the same drift monitoring, and the same record of what changed and who approved it. The Platform is built around that record.
The craftsmanship finding points the same way. Judgement sitting in underwriters' heads is not a permanent ceiling on automation. It is an undocumented dataset that nobody has been paid to capture. Start capturing the decisions with their reasons attached and the ceiling becomes a data readiness problem, which is tractable, rather than a talent problem, which is not. Early language-model pilots have a narrow and honest job here: draft the change memo, assemble candidate test cases from prior decisions, summarise the rationale a human underwriter then signs. Useful, checkable, and nowhere near the release switch.
There is a last caution in the same review, and it is the one that should be fixed before anything else. The group recorded that time to market, process lead times, claims leakage and pay-out times were not measured across the organisation, and were produced only on request. Even if the roll-out had lifted straight-through processing to the level the strategy asked for, nobody in the group could have proved it. A score assigned to a plan is only ever as good as the instrumentation waiting to check it.
