Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsFinance

Stop Benchmarking Against the Market

RealAIApr 23, 20238 min read
FinanceProcess MiningRetail BankingAnalyticsData Strategy

The benchmark slide has a standard shape. One column for the metric, one for what the agreement or the plan requires, one headed industry average. It is the third column that gets the discussion, and the discussion almost always ends the same way. Our mix is different. Our customers are older, our regions are rural, our regulation is heavier, our book is not their book. Everyone in the room knows the sentence is partly true, nobody can say which part, and the meeting closes with the number noted and nothing scheduled.

I have watched that happen enough times to think the problem is not the sincerity of the people in the room. It is the object on the screen. An industry average is a fact about other people. It has no instruction inside it.

The most useful measurement chart I have worked with had no industry column at all. It came out of a journey mining engagement at a European cooperative banking group, on the residential mortgage process, and it plotted a single quantity for every local unit in the group: the time from a customer becoming an identifiable lead to that customer sitting in a first appointment. Lead time in days on one axis, local units on the other, sorted longest to shortest. That is the entire design.

What the chart does that a market benchmark cannot

Put the two objects side by side and the difference is not precision, it is what each one licenses you to say next.

An industry average is a single point summarising organisations you cannot inspect, measuring a quantity they defined for themselves, over a period they chose, in a mix you do not share. Every one of those four gaps is a legitimate reason to discount it, and because they are legitimate they get used. The number survives the meeting. The behaviour does not change.

A sorted internal distribution has none of those gaps. Every bar is a unit of the same group selling the same product in the same window, and every number was computed by the same query over the same event stream. When the fast end of that distribution reaches about one day, the target is not an aspiration and not an import. It is a description of what some of your own colleagues did last month. There is no version of our mix is different that survives it, because the comparison group is you.

That is why I now ask a different opening question in operations work. Not how do we compare to the market. Where is our own best decile, and what is between us and it.

The line, and why its colour is the technique

The chart carries one more design decision worth stealing. The bars are not simply drawn to their height. Each one is split at a constant reference level of roughly nine and a half days, with everything below the line in one colour and everything above it in another. The eye does not read lead time. The eye reads excess.

That reframes the whole discussion. A bar of twenty days is not a bar of twenty days, it is roughly nine and a half of process and roughly ten and a half of waste, and the waste is coloured. Nothing about that split is technically clever: any competent query over an event store will produce it, and the reference level is a choice you make and defend. The persuasion is in refusing to plot the total, because the total is a fact and the excess is a decision.

The shape is not the spread

There is a trap in this chart, and it is the trap in every sorted internal distribution I have built since.

Thirty times is a headline, and the headline belongs to a single unit. The slowest sat around thirty-six days while the second slowest sat near twenty-two, a gap of about fourteen days between rank one and rank two, after which the curve settles into an ordinary slope. The mean is dragged above the median by that one bar. So the sentence the chart supports is not there is a thirty-fold performance range in this group. It is one unit is behaving unlike the others, and separately, a large group of units is sitting a little over the line.

Those are two different programmes. The outlier is a case for a visit: something is broken there, probably one queue, one vacancy, one system, and it will not be fixed by a policy. The crowd just above the line is a case for a standard, and it holds almost all of the recoverable time, because there are so many of them.

Any measurement you intend to run continuously has to make that separation on its own, or every alert it raises will be about the same one unit and people will stop reading them within a month. That belongs in the specification of the reporting, not in the analyst's head.

~36 days
Lead time to first appointment, slowest local unit
~1 day
Lead time at the best-performing end of the same group
~9.6 days
Reference line at which each bar was split into process and excess
~2 in 5
Local units carrying visible excess above that line

The instruction lives in the drill-down

A distribution on its own still stops short of an instruction. It tells a slow unit that it is slow. It does not tell it what the fast ones do differently, and that is the sentence the manager of a slow unit actually needs.

The engagement closed that gap by drilling into the discovered process of a single unit. What came back was a six-activity model built from customer relationship events rather than drawn in a workshop: sales opportunity, inbound call, outbound call, appointment, quote, purchase, wired with the transitions the event log actually contained. Two of those details do the work. Inbound and outbound calls appear as separate activities, which a designed funnel would have collapsed into contact. And the appointment is reachable by two routes, straight from the opportunity or by way of an outbound call, so what the model names is a routing choice rather than an effort level. The drill-down was run on one unit, so nobody compared those two routes across the fast and the slow ends of the distribution. That comparison is the obvious next query, not a result already in hand.

That is the shape of a usable instruction. Not work harder, and not match the industry. Take the sequence of steps your fastest units run, and run it.

An industry average is a fact about other people. A sorted distribution of your own units is a fact about you, and the fast end of it is a target nobody can argue is unrealistic, because your own people already hit it.

What has to be true before any of it compares

The uncomfortable line in this work was not on the chart. It sat in eight-point type in the corner of a diagram, marked as a precondition: a uniform way of working in the customer relationship system, so that steering and comparing become stronger.

Everything above depends on it. A lead time per unit is only meaningful if every unit creates the opportunity record at the same moment in the real process, and books the appointment activity at the same moment too. Where recording discipline varies, the distribution measures data entry habits and presents them as performance, with the same clean bars and the same confident colours. There is no statistical repair for that, and a model trained on it inherits it silently. This is ordinary data readiness work, the same lineage and definition discipline any deployment needs, and it has to land before the league table is published rather than after somebody disputes it. It is the opening block of a RealAI Platform engagement for that reason, and it is why we ask which number you want to move before we ask which model you want to build.

Where the measurement is supposed to live

The last thing worth carrying across is organisational rather than analytical. The proposal did not ask for an analytics team. It asked to embed the process indicators in three places the group already had: the reporting that runs across the straight-through processing chain, the existing programme by which local units are compared, and the development and marketing team that changes the product. The forecast and the output were to be steered jointly with the local unit, not issued to it, which in a cooperative structure is the difference between a target and an insult.

That is the whole argument compressed. Move from steering on the output at the end of a quarter to steering inside the process while it runs, put the number where the decision already gets made, and set the target from your own fast end.

Figures are as read off the charts of a journey mining engagement on the cross-channel mortgage process at a European cooperative banking group: a per-unit lead-time distribution computed from the group's own transaction records, and a discovered process model for one unit. The lead-time values are pixel reads from a plotted chart, good to roughly a third of a day, not figures stated in the source. That engagement produced findings and recommendations, including the proposal to embed the measurement permanently. Reading it as an argument against market benchmarking is ours.

An industry average is a fact about other people. A sorted distribution of your own units is a fact about you, and the fast end of it is a target nobody can argue is unrealistic, because your own people already hit it.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?