A sorted bar chart is the most persuasive object anyone brings to an operations meeting, and the easiest one to read wrong. Rank every unit of the business on the same measure, longest bar on the left, and the room agrees on the shape of the problem in about four seconds. That speed is the appeal, and it is where the reading goes wrong, because the eye lands on the left edge and the mouth reports a ratio.
I have been carrying one of these charts around for a while now. It came out of a process mining engagement on the cross-channel mortgage journey at a large European cooperative banking group, where the point of the work was to stop steering on output after the fact and start steering inside the process while it ran. Event data from the mobile, online and customer-relationship systems was stitched into a single case history, and lead time to the customer's first appointment was measured per local unit, on the group's own transaction records rather than on survey returns or self-reported management information.
The result is a descending bar chart, one bar per local unit, with a best-practice group annotated at the short end. It is also the chart that gets quoted incorrectly every time, including by me the first time I saw it.
The number the room repeats
Somebody in every readout says the spread. Thirty-one times between the slowest local unit and the fastest, and the number does its job, which is to make people sit up. It is also true. The arithmetic behind it is roughly 36 days over roughly 1.1, both figures read off the plotted bars against a calibrated axis rather than printed anywhere on the page, so treat them as accurate to about a third of a day.
What the number is not is a statement about the organisation. It is a statement about two units out of the whole population, one at each end, joined by a division sign, and neither of them is where the recoverable time sits.
That sounds pedantic until you notice what gets commissioned off the back of it. A thirty-one-times headline commissions a hunt for the worst offender. A distribution with two in five units a few days over a reference line commissions something else: a standard, applied everywhere, with the best-practice group's method written down and moved.
The cliff, and why it changes the diagnosis
Go back to the left edge and look at the first two bars rather than the first and the last. The slowest local unit is near 36 days. The next one is near 22. Then the steps between neighbours fall under two days, and quickly to tenths, all the way down to the best-practice group.
Fourteen days between rank one and rank two, in a population of well over a hundred units running the same product on the same systems in the same period, is not a tail. Tails are gradual by construction. A step that size means the mechanism changed. Something at that unit is doing a different thing: a queue that only forms there, an appointment calendar maintained differently, a handover waiting on a person who is not always present, or, the one worth checking first, a recording convention that starts the clock at a different event.
The practical consequence is that the outlier and the mass need two separate programmes. Fixing the outlier is casework. Somebody goes there, watches the process, and finds the one thing. Fixing the mass is standard-setting, and it has to survive being applied in every other unit at once. Neither does the other's job. Send a standard to the outlier and it will comply while staying at 36 days, because compliance was never what was broken. Send caseworkers to the mass and you are staffing two units in every five with a caseworker apiece.
What the mean quietly does
The median across the population is about 8.9 days. The mean is about 9.4. Most of that half-day gap is the ordinary right skew of a distribution with a long thin top, and attributing all of it to the outlier would be wrong. The part that really is one bar is smaller and more interesting. Take the 36-day unit out and the mean falls to about 9.2 while the median settles at about 8.8: two tenths off one statistic, a rounding error off the other.
Two tenths of a day sounds like nothing until you remember what the mean is used for. It goes into the monthly pack. Targets get set against it. Forecasts are anchored on it, and once anyone builds a model it becomes a feature: average lead time per unit, per month, fed into whatever comes next. Every one of those uses inherits a contribution from a single unit whose process is not the process being described.
The reference line on the chart, at about 9.7 days, is drawn above both. Each bar is split at that level so that the portion above it reads as excess, which is a good piece of persuasion because it converts a ranking into a quantity of recoverable time. But a threshold sitting a little above the mean, when the mean is itself being pulled up by an outlier, is a threshold that has been softened by the very unit it should be catching. Take the outlier out and the same logic would draw the line lower, and more units would find themselves on the wrong side of it.
- ~36 days
- Lead time to first appointment, slowest local unit
- ~22 days
- Second slowest, a gap of about fourteen days
- ~8.9 days
- Median across the population, against a mean near 9.4
- 2 in 5
- Units above the reference line at about 9.7 days
Where this reaches a machine learning pipeline
Nobody builds a model to find the worst branch. You can see it from the doorway. Models get built for the thing the chart cannot do by itself, which is to keep producing this reading every month, per unit, without a consultant in the room.
Take alerting first. Point a monitor at per-unit lead time with a threshold anywhere near the group mean, and the same unit trips it every single period. Alerts about a known condition are not information, and a queue of them trains the people receiving them to stop reading. The fix is not a cleverer detector. It is deciding, before the monitor is built, that this unit is out of scope for the monitor and inside the scope of a person.
Take features. Per-unit aggregates computed as means across a population with a structural outlier carry that unit's signature into every downstream model, and nothing in the pipeline flags it, because there is nothing internally inconsistent about a mean. A feature store is supposed to settle exactly this, by pinning down what a value means and how it was computed rather than letting each consumer re-derive it. Store the median and the count above threshold alongside the mean, and the choice becomes visible at the point of use instead of buried in a query somebody wrote once. Pinning that down is the opening block of a RealAI Platform engagement, before a single feature is served to anything.
Take the definition underneath all of it, which is the load-bearing question and the one the source never answers. Lead time to a first appointment is a difference between two timestamps, and both of them are events somebody's local convention decides. If one unit logs the case at the enquiry and another at the completed application, the chart is not ranking speed. It is ranking bookkeeping, and the fourteen-day cliff is precisely the shape you would expect if one unit's clock starts earlier than everyone else's.
That is a testable question, and the first thing I would test. It cannot be tested from the chart, only by going back to the event log and comparing the activities that open each case, which is ordinary lineage work of the sort that has to land before anything is modelled.
A gap of fourteen days between rank one and rank two is not a tail. It is a seam. On one side of it sits a process that works badly, and on the other sits a process that is not the same process.
What we would do with this chart
Four things, in order, none needing a data scientist, all of them ahead of hiring one.
Check whether the outlier is a process or a definition, by reading the opening activity of its cases against everybody else's. Remove it from every group summary, keeping it named and tracked on its own, so the mean, the target and any threshold describe the population they are meant to describe. Publish the median and the count above the line alongside the mean, so the number in the pack cannot be moved by one unit again. Then take the best-practice end seriously, because the improvement target here is internal and already demonstrated inside the same organisation, which removes the usual objection that an external benchmark measures a different kind of market. Those four steps are what RealAI's Consult team works through on an operations distribution before anybody asks what could be modelled.
And treat the steering as joint rather than issued. Where local units run themselves, a central ranking that arrives without local co-ownership of the forecast reads as an accusation, and gets handled as one. The measurement can be central. The commitment cannot be.
None of that is modelling. All of it decides whether the modelling would have meant anything.
Figures are read off the lead-time chart produced in a process mining engagement on the cross-channel mortgage journey of a large European cooperative banking group, measured on the group's own transaction data. The bar count behind the proportions, the per-unit values and the position of the reference line were measured from the plotted chart rather than stated in the source, and are approximate. That engagement produced findings and a proposal for embedding process indicators permanently, not delivered results. Reading the distribution's shape as a diagnosis is ours.
“A gap of fourteen days between rank one and rank two is not a tail. It is a seam. On one side of it sits a process that works badly, and on the other sits a process that is not the same process.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
