Most capability assessments ask two questions. Where are you, and where do you need to be. The distance between the answers becomes a gap, each gap becomes an action, the actions become a roadmap, and the roadmap lands on the desk of someone who is already running a programme that covers half of it. Nobody set out to fund the same work twice. The scoring method did it for them.
A digital capability assessment run with a European composite insurance group asked a third question, and that third question is the entire method. Where does this capability stand today. Where will it stand once every programme already funded and staffed has landed. Where does the strategy need it. Twenty-two capabilities, scored three times each on a five-point scale. The gap that generated an action was never the distance from today to the target. It was the distance from the plan to the target.
Twenty-two capabilities went in. Ten came out.
The challenge
The group was not starting from nothing, which was precisely the problem. Roughly twenty digital initiatives were already running across the operating brands and the technology function: a customer relationship platform rolling out, value-chain automation programmes, a business process management platform, a data warehouse for campaign work, a real-time query layer, and a set of culture and collaboration programmes. The assessment's own framing of the failure mode was organisational rather than technical. The initiatives were scattered, poorly shared between brands, and many of them stalled short of value because the back office behind them had not been integrated. Maturity therefore varied widely between parts of the same group.
Score that estate with two levels and you produce a document that is both accurate and useless. Accurate, because the present really is uneven. Useless, because every gap it names is partly being closed already by someone with a budget code, and the reader cannot tell which parts. The programme director whose work is already underway reads the new action list, recognises the overlap, and quietly discounts the whole assessment. That is the ordinary death of a capability review, and it is a scoring failure before it is a political one.
There was a second problem underneath the first. The three scores cannot honestly come from the same place. Where a capability stands today and what the funded portfolio will add to it are facts the organisation holds, recoverable from the people running the work. What the capability needs to reach is not a fact at all. It is a consequence of strategy, and asking the delivery teams to supply it produces a target set to whatever they think they can hit.
The approach
The assessment separated those two evidence types and never mixed them. Current and planned levels were gathered bottom-up, through an online survey instrument and interviews across the technology function and two operating brands. Target levels were derived top-down, from the group's stated strategy and from discussion with the management team. Same scale, two sources, and the difference between them is where the argument sits.
Target calibration was where the method got unusually disciplined. Rather than benchmarking, the targets were read off three strategy statements: hold parity with the competitive field, grow volume in the core propositions through new segments and new distribution, and cut cost through process work, consolidation and readiness to outsource. Those statements were then translated into a single rule. Level four, the second-highest band, became the default target for all twenty-two capabilities. Level five was reserved for exactly five of them, all sitting in customer data, product construction, pricing or processing, and none of them cultural or infrastructural. Seventeen capabilities were deliberately capped below best in class. An assessment that lets every capability aspire to the top has not made a decision, it has made a wish list.
Then the portfolio was mapped onto the same grid. Each in-flight initiative was plotted against the capabilities it touches, and each capability was given an expected improvement in maturity points, banded: under half a point, half to one, one to one and a half, one and a half to two, and more than two. That banding is the credit. It is the mechanism that lets a review say, in a number, how much of the distance the organisation is already paying to cover.
Only then were gaps computed, and only from planned to target, under a published threshold rule. Less than one point apart counted as no gap or a small one. One to two points, medium. More than two, large. A rule printed on the page means a leadership team can dispute a score but cannot argue about what the picture means.
The outcome
Ten of the twenty-two capabilities still fell short of target once the in-flight portfolio was credited. They spread three, two, one, two and two across the five assessed dimensions, which is a more useful shape than a ranked list: no dimension was clean, and none was hopeless.
The sharpest result came from holding the three levels side by side. One of the five capabilities set to the top target level, processing a request end to end without a person re-keying it, was the one the existing portfolio was expected to move least. The largest expected step change landed instead on a capability whose migration programmes were already funded and already running. A two-level score would have shown both as gaps of similar size and said nothing about which one the organisation was actually working on. Ambition and momentum were pointing at different capabilities, and only the middle level made that visible.
The absences carried information too. One operating brand declined to score fifteen of the twenty-two capabilities; another left two unscored. Those blanks are not missing data to be chased. A brand scores what it believes it owns, so the unscored cells map ownership more honestly than any responsibility matrix in the pack.
What was delivered here is an assessment. A graded baseline, a set of target levels traced to strategy statements, a portfolio credit, ten residual priorities, and an action longlist sequenced onto a transformation roadmap. No capability moved because of this document. The scores describe the weeks they were collected in, the targets are ambitions, and the improvement bands are expectations about programmes that had not yet landed. Anyone reading it as a result is reading it wrong.
The method transfers directly, and the failure it prevents is more expensive now than it was then. Machine learning portfolios accumulate the same overlap: a feature store funded twice because the first sits inside a data programme nobody scored, a lineage effort duplicated because governance and engineering each wrote their own gap, deployment tooling bought while an MLOps stream was already standing it up. Score the present alone and every one of those becomes a fresh line item. Score the plan and most of them disappear before they reach a budget.
One part of the original method is worth replacing rather than repeating. The current and planned levels came from a survey, which means they came from opinion, they were stale within weeks, and they cost a round of interviews to refresh. A delivery estate already emits the evidence those scores were guessing at: pipeline run histories, deployment records, lineage coverage, incident and rollback logs, how long a model sits between build and production. Read the grade off the artefacts through the Platform and the assessment stops being an engagement and becomes an instrument that refreshes itself. The three-level structure stays exactly as it was. Only the survey goes.
