Insurance analytics is a large-numbers business and almost every method in it says so out loud. You price against a population, you reserve against a population, you accept that any single policy will be wrong and you rely on the wrongness cancelling across the book. Cross-validation works because there is more where that came from. A holdout set works because the rows you set aside look like the rows you kept.
That assumption is so far inside the tooling that nobody states it. It gets stated only when it breaks.
I saw it break inside one bid appendix. A competitive proposal I worked on for a European composite insurance group, standing up a group-level analytics function, carried a library of eight candidate cases, each one a page, each written to the same shape: what the data is, what the model does, who benefits. Seven are population machines. The eighth is not, and reading them side by side is the clearest lesson in method I have had out of any deck.
Seven cases that need a crowd
Read the seven for what they assume rather than what they promise and they collapse into one method with seven faces.
The loyalty case builds a real-time recommendation score for customer groups out of social data, sorting people by how likely they are to recommend, repurchase and forgive. It needs groups to score. The price sensitivity case estimates willingness to pay per individual so that customers who want value for money and customers who want extra service can be offered different propositions. It needs a distribution to sit that individual in. The lead generation case predicts a customer's life journey from archetype life cycles, on the stated basis that life-cycle stage is a key driver in most behavioural predictive models, and archetypes are averages with a name on them. The claims behaviour case learns who will claim and how the gap between amount requested and amount paid tends to move, a pattern read across many claims. The connected-home case reads devices in the house to price behaviour and mitigate risk, and only earns the integration effort across many houses. Even the trend case, watching social media for topics that emerge before the real world does, is a crowd instrument by construction.
None of that is wrong. It is the right method for a book of many small, similar, replaceable risks, which is what most of a composite insurer looks like. The error on one policy is diluted by every other policy. You can afford to be wrong per customer as long as you are right per cohort, and cohorts are the unit the whole toolchain speaks in.
The one that needs the opposite
Then there is the commercial asset case, and it flips the sign on all of it.
The book here is a small number of very large single risks: a power installation, a construction programme, a vessel. The proposal's own line is that there are often only a few customers of this size. The title of this piece says five; the slide does not, and the exact count is not the point. The point is that it is small enough to count on your fingers, which is a different regime, not a smaller version of the same one.
The case proposes gathering real-time data from the operator's own safety measurement devices and augmenting it with external sources, the ones it names being weather data and age and wear data from comparable installations. That instinct is worth pausing on. When you cannot borrow strength from other customers, because there are almost none, you borrow it from physics instead: the same class of asset elsewhere, the environment the asset sits in, the wear curve of the equipment. The population has not disappeared. It has moved from the customer axis to the asset axis, and the model has to be rebuilt around the new one.
What inverts, precisely
Three things change, and each one breaks a habit rather than a tool.
Per-risk accuracy stops being an input and becomes the answer. In a mass book, mispricing one policy is a rounding error on the year, absorbed by every other policy priced beside it. With a handful of very large risks, the mispricing and the book are the same object. There is no averaging layer between the model's error on one customer and the result the group reports. The case's own line is that improving the risk models for these customers is extremely important, and on a book this shape that is a statement about arithmetic rather than about emphasis.
Sampling stops helping. A twenty percent holdout on a book you can count on your fingers sets one contract aside. Cross-validation on a handful of contracts is not validation, it is arithmetic performed for the comfort of the person doing it. Backtesting against claims history is not much better, because a book like this generates very few loss events by design, and each is unlike the others. The machinery that gives you confidence in a mass book gives you nothing here, and it gives you nothing while still producing an output that looks like a score.
Direct observation beats history. If you cannot learn the risk from siblings and cannot learn it from your own loss record, the remaining source is the asset itself, measured now. That is exactly what the case reaches for: the operator's safety instrumentation, streamed rather than sampled at renewal. It is the substitute for the missing population, and the reason a case about underwriting turns out to be an engineering and data-access problem long before it is a modelling one.
- 1 in 8
- Candidate cases in the library that assume few customers rather than many
- ~25%
- Cost reduction the volume-side process case reports from another organisation, about 3.3 million euros
- Direction only
- The evidence the large-risk case offers, because there is no population to measure against
Readiness means something different here
Data readiness for the seven population cases is the familiar work: get the policy and claims history into one place, agree the definitions, build the features, keep the lineage honest. Scale is the challenge and it is a solved kind of challenge.
Readiness for the eighth case is a different job with the same name. The data you need is generated by equipment your customer owns, on a network your customer controls, under a commercial agreement that has to be negotiated before an engineer can touch anything. The pipeline is thin and continuous rather than wide and batched, the feature store is organised per asset rather than per cohort, and a single drifting sensor feed does not degrade an average, it moves the price of a material share of the book. Access agreements, clock discipline and a monitor on every feed are where RealAI's Platform work starts on a book like this, and lineage matters more here, not less.
The case's benefits line is the usual pair, lower cost and lower exposure to large losses for the insurer, a more competitive premium for the customer. The route to it runs through a conversation with the customer's operations engineers, not through a segmentation.
With a mass book, mispricing one policy is a rounding error on the year. With a handful of very large risks, the mispricing and the book are the same object.
How you would know it worked
This is the question that catches teams out, and the library answers it by accident.
The process case elsewhere in the same appendix could put a number on itself because a support process runs constantly and leaves a trail, so mining it end to end yields a before and an after: fewer steps, about a quarter of the cost gone. Repetition is what makes measurement cheap.
A book of a few very large risks has no such trail. There is no control group, and neither a good year nor a bad one proves anything. So validation moves from loss experience to the physical claim the model is making. Does the predicted stress on the asset match the measured stress? Does the wear model track what the instruments report? You test the mechanism rather than the outcome, because the outcome arrives too rarely to test.
Teams trained on marketing analytics find that unfamiliar and engineers do not, which tells you who belongs in the room.
Most books are closer to five than to five million
The regime is not confined to nuclear plants and hulls. Commercial property programmes, marine, aviation, energy and large specialty lines sit in it, and so does much business-to-business work outside insurance, wherever a handful of accounts carry most of the exposure.
The failure mode is quiet. A team copies a method built for retail volumes, applies it to a book it could list on one page, and gets an output. Nothing errors. The confidence interval is meaningless and no part of the stack says so.
The first question in scoping is therefore not which use case, and not which model. It is how many independent risks are actually in this book, and it is the first thing a Consult engagement settles, because the answer decides whether your instinct to sample is a method or a superstition.
Source: the appendix case library of a competitive proposal to a European composite insurance group standing up a group-level analytics function. Those cases are indicative examples offered in a bid, and the figures quoted in them describe work at other organisations, not results delivered to this group. The proposal produced a recommended shape of work. Reading the library as a statement about portfolio shape is ours.
“With a mass book, mispricing one policy is a rounding error on the year. With a handful of very large risks, the mispricing and the book are the same object.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
