Ask a claims director what the fraud rate is and you will get a number. Ask where it comes from and you will get a process: claims that tripped a referral rule, claims an investigator picked up, claims closed under a particular code. The figure is real. It is also a measurement of the investigation function rather than of the customers, and it routinely becomes the training label for the first fraud model anybody builds.
I keep returning to an analytics mapping proposal I worked on for a European composite insurance group that was standing up a group-level analytics capability across its operating companies. The task was to lay out on one sheet where analytics could move the profit and loss: revenue on one arm, cost and claim cost on another, risk on a third. A set of candidate cases was placed against that structure. One sits in claims, and its wording is unusual enough that I have been quoting it ever since.
It does not ask for a fraud model. It asks for a model of claims behaviour, built from behaviour patterns, attitude patterns and the strength of those patterns, informing pricing, underwriting, product design and the claims handling approach. Fraud detection is named as a consequence of that model rather than as its target, and the stated result is that fraud falls and the claims process gets faster. That is a different sentence from the one insurers usually write.
What a flag actually records
A fraud flag is created by an investigation, and investigations are allocated. Somebody wrote the referral rules, somebody staffed the unit, somebody decided this month that motor total losses deserved more attention than household escape of water. Every one of those decisions sits upstream of the label. What ends up in the column is a joint product of the customer's conduct and the organisation's attention, and the second term is usually the larger.
Train a classifier on that column and you get something that scores well and does the wrong job: it has learned where the searchlight points. Claims that were never referred sit in the negative class beside claims that were referred and cleared, and the model cannot separate those populations because the data does not. That is ordinary selection bias, and no amount of extra history, a better algorithm or a larger feature store repairs it. It is a property of how the label was manufactured.
There is a second problem stacked on top. A flag is binary and terminal, a yes or a no about the whole claim, recorded after the file closes. Almost nothing that matters in claims handling behaves like that. Positions harden, an itemisation gets revised, a claimant goes quiet and then arrives with a lawyer, and all of it happens while the money has not gone out yet, which is the only period in which a model can change anything. A label defined at closure describes the one moment when intervention is no longer possible.
The signal that made the case worth reading
The proposal names something specific enough to build against: sensitivity to the deviation between the amount requested and the amount paid.
Both figures already exist. Every claim has an amount asked for and an amount settled, recorded because the handling system has to record them in order to pay anybody, and the gap between them is arithmetic. What varies across customers is the conduct in response to that gap: how hard the difference is contested, how quickly a position moves once a first offer lands, whether the supporting documentation changes shape afterwards, how often the file bounces back for another round. None of that needs an investigator to have formed an opinion, and none of it is missing for the claims nobody investigated.
That is what makes it a behavioural feature rather than a verdict. It is present for the whole book, continuous rather than binary, accruing during the life of the claim, produced as a by-product of settling claims at all. It also does something a flag cannot: it separates the customer who disputes a settlement because the settlement is wrong from the customer whose disputing follows a pattern across every claim they have ever made. Those two look identical in a flag column and completely different in a conduct series.
The proposal then treats the same conduct model as an input to pricing, underwriting and product design, not only to claims. That works only because the quantity is behavioural. You cannot underwrite on a variable that exists only for the people your own investigators happened to examine.
A label that was already built properly
The same file carries an older piece of retention work that is worth reading beside the claims case, because it shows what a target looks like when the behaviour is observable.
The model predicted whether a client would cancel one policy, cancel two, or cancel neither, inside ninety days. It ran on a deliberately small subset of around fifteen variables, trained on 80 percent of one year of data and validated on the remaining 20 percent, and reported 83.3 percent exact predictions across the three categories, the other 16.7 percent being predicted cancellations that did not occur.
The accuracy is not the interesting part, and it should be read with its limits: one year of data, a single hold-out, no comparison against the client's existing model quoted anywhere. The target is the interesting part. Cancellation is an event the administration system records because it has to process it, and nobody adjudicates a cancellation. The label carries a horizon, ninety days, so the question has a deadline attached. It carries a granularity, one policy or two, so the size of the event is part of the prediction. And its negative class means something, because a customer who did not cancel is genuinely a customer who did not cancel, not a customer nobody got around to checking.
Set the usual fraud label beside that. No horizon, no granularity, and a negative class that mixes clean claims with unexamined ones, created by a person at the end for a small and non-random slice of the book. It fails every property the retention label satisfies, and the difference has nothing to do with modelling technique.
- 83.3%
- Exact predictions across three cancellation outcomes, hold-out validation
- 16.7%
- Predicted cancellations that did not occur
- ~15
- Variables in the feature subset
- 90 days
- Prediction horizon written into the label
The loop that makes a bad label permanent
A fraud model trained on referral outcomes goes into production and starts recommending referrals. Those recommendations become investigations, and those investigations become next year's labels. The searchlight it learned from is now partly its own output, and each retraining cycle tightens the agreement. Performance metrics improve. The customers nobody has ever looked at stay as unexamined as they were, and nothing in the pipeline can report it, because internally the arrangement is consistent.
It is the same failure I wrote about in service operations earlier this year: a number that looks like the ones beside it, means something different, and passes straight through a pipeline with no way to object.
A flag records who was looked at. Behaviour records what the customer did. Only one of those is a property of the claim.
What we would build instead
None of these repairs are exotic, and all are cheaper than the model.
Define the target as an observable event with a horizon and a size, the way the retention work did, and accept that the event is a proxy for fraud rather than fraud itself. Score conduct across the life of the claim instead of once at closure, so the score exists while the file is open and the decision is live. Keep the investigation outcome as a separate audited variable used to calibrate the model, never as the thing it is trained to reproduce. Record the referral decision itself as data, including who was examined and why, so the searchlight can be modelled on its own rather than silently absorbed. And keep a person on the repudiation decision, with the conduct signals that drove the score written in language a customer could be shown, because a score about behaviour carries an obligation to explain that a score about a transaction does not.
Do those five and the operational prize is the one the proposal names. Claims that carry no adverse conduct signal are the files that can move to straight-through settlement, while handling effort concentrates where the signal is. Faster claims and lower fraud leakage become the same programme rather than two competing ones. That is the sequencing a RealAI Platform engagement starts from, and it is why the first question we ask in claims is not which model you want but which column you were planning to train it on.
Figures are as recorded in an analytics mapping proposal prepared for a European composite insurance group building a group-level analytics capability, together with the earlier engagement records appended to it. The claims case described here is a candidate placed against the profit and loss for evaluation, not a delivered programme, and the retention figures are the authoring firm's own account of prior work. Reading the pair as an argument about label design is ours.
“A flag records who was looked at. Behaviour records what the customer did. Only one of those is a property of the claim.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
