Every conversation about automating a claims chain arrives at the same place, and it is never the model. It is how anyone would know afterwards whether the thing worked. Somebody produces the operational pack, and the pack is organised the way the company is organised: intake has its numbers, assessment has its numbers, payments has its numbers, and each of those numbers belongs to a manager who can be held to it.
That arrangement looks like rigour. It is the single most common reason a process automation programme cannot prove its own value.
I keep coming back to one review for the sharpest statement of the problem I have seen written down. It was a digital readiness review at a European composite insurance group, run through surveys and stakeholder interviews across IT and parts of the retail business, producing a gap analysis and a set of proposed measures rather than a delivered result. On the page about how a flexible organisation steers itself, the review proposed two measures. The first was the cadence of performance review and feedback. The second was second-order: what proportion of the measure set attaches to chains of work and projects rather than to departments.
A measure about measures. It sounds like the kind of thing that gets added when a committee runs out of ideas. I think it is the most useful line in the document.
What a department-shaped scoreboard cannot see
Take a retail claim from first notification of loss through intake, adjudication, settlement and billing. That path crosses several organisational boundaries, and at each one the work changes hands. Now score it with the pack. Intake reports contacts handled and average handling time. Assessment reports files closed per assessor. Payments reports payment runs completed within the window. Every one of those boxes can be green in a month during which a customer waited weeks for a decision that should have taken days.
Nobody is lying. The measurement system has no place to record what happened, because what happened happened between the boxes. A file that bounces back to intake for missing information reads, from intake's side, as additional contacts handled, and from assessment's, as additional files touched. The rework is not invisible because someone hid it. It is invisible because nobody owns the interval, so no number is attached to it.
The review put this in the plainest terms in a different section, quoting a respondent on the prevailing culture: my own tile is clean, yours is dirty. That is not a character flaw in the people involved. It is an accurate description of what their scoreboard rewards. If every measure you are held to lives inside your own boundary, keeping your own boundary clean is exactly rational, and pushing a difficult file across the boundary improves your position.
This is also why throughput targets get flagged. Claims handled per day is a real measure of a real thing. It is also the measure most easily improved by handing the hard cases to someone else, and by closing the easy ones first while the complex file ages. The alternative the page itself offers, the share of claims handled correctly and on time, has the opposite property: you cannot improve it by moving work sideways, because the clock and the quality check both survive the handover.
The measures that were missing were the chain measures
The diagnostic half of the review is where this stops being theory. On monitoring, the group scored comparatively well. Operational management tooling, an internal IT platform programme, a rationalisation effort and an enterprise resource platform migration were all credited on that page with pushing monitoring towards something end to end and uniform. By the standards of the ladder in use, that section came out as the strongest one in its group.
And in the same section, one respondent wrote that operational performance in terms of time to market, process lead times, claims leakage and pay-out times was not in place enterprise-wide and was only performed on request. Read those four items. Time to market is a chain measure. Process lead time is a chain measure. Claims leakage is a chain measure, and arguably the chain measure in general insurance. Pay-out time is a chain measure. Every single thing on the missing list was a number that spans departments, and the reason they were produced only on request is that no standing report was shaped to hold them.
The mechanism is boring and worth stating anyway. Standing reports get built by whoever owns a system, and systems are procured by departments. The chain measure has no natural home, so it becomes a project: someone pulls extracts, reconciles keys by hand, produces a figure for a steering meeting. Because it costs real effort every time, it gets produced when someone senior asks, which is to say when there is already a problem. The measurement arrives as an investigation rather than an instrument.
- On request only
- Time to market, process lead times, claims leakage and pay-out times, as recorded by a respondent
- Claims per day
- The throughput goal the review places near the bottom of its ladder
- Correctly and on time
- The single claims measure it offers instead, quality and speed together
- Two to four
- Respondents behind each sub-dimension of the steering scorecard, a count the assessment printed beside the score
That last figure matters for how much weight any of this can carry. The steering section rested on two to four respondents per sub-dimension, and the assessment printed that count beside the score instead of letting the reader assume a denominator. The finding I am drawing out is the proposed measure, not a measured state of the organisation.
Why a data team should care about a performance management page
The reason I mine an organisational design page for machine learning purposes is that the scoreboard decides which pipelines can ever be justified.
Start with the label. A model that predicts which claims will leak, or which will breach the service window, needs an outcome column, and the outcome is a property of the whole file rather than of any one step. If leakage is computed by hand on request, there is no history of it, and there is no history because nobody was paid to keep one. The absence of the chain measure is the absence of the training label. No amount of feature engineering recovers it.
Then lineage. A metric assembled from department extracts each time somebody asks drifts between askings, so two figures produced at different times are not comparable. A feature store built on them promises what it cannot deliver: that a value means the same thing in the training window as in production.
Then adoption, which is where most of these programmes actually die. Suppose the automation works. Suppose straight-through handling of the simple claims lifts, lead time through the chain falls, and the customer gets a decision in days rather than weeks. Some department is now absorbing more exception work than before, because that is what happens when you automate the easy path. On a department-shaped scoreboard, that department's numbers deteriorate. The person accountable for them will resist, and will be right to, because you have measured them worse for a change that made the company better. The pilot gets described as promising and quietly not scaled.
A department-shaped scoreboard does not tolerate cross-department failure. It cannot detect it. Every box reports green while the claim sits in the gap between two of them, because no number in the building moves when a handover stalls.
What to do before the modelling starts
None of the repairs needs a data scientist, and all of them have to land before hiring one pays. It is the four-step pass RealAI's Consult team runs at the front of a claims engagement, and it is deliberately unglamorous.
Name the chains. Not the departments and not the systems, the chains: notification to decision, decision to payment, first contact to closure. A chain that has no name cannot be measured, and most organisations have never written the list down.
Give each chain one interval measure and one quality measure, and let the quality measure carry a definition that survives a handover. The example on the page is a good one precisely because it is joint: correctly and on time, not correctly, and separately, on time.
Put one named person against each chain number. Not a board, not a forum. A person who can be asked why it moved and who is allowed to look inside every department the chain crosses.
Then count. Take your current pack and sort every number into chain or department. The proportion you get is the measure the review was reaching for, and you can compute it this afternoon without buying anything.
On one capability the review never filled the proposed-measure field in at all, leaving a placeholder where the measures should have been: the page about experimenting and sharing what is learned. That blank is the exception. Everything a claims chain does can be counted along the chain, and the reason it usually is not has nothing to do with difficulty.
Drawn from a digital readiness review at a European composite insurance group: its capability scoring, the measures it proposed, and the stakeholder interviews recorded alongside them. That engagement produced findings and recommendations ahead of a coordinated roadmap, not delivered results. The measure discussed here was proposed, not implemented. Reading it as a precondition for machine learning work in claims is ours.
“A department-shaped scoreboard does not tolerate cross-department failure. It cannot detect it. Every box reports green while the claim sits in the gap between two of them, because no number in the building moves when a handover stalls.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
