Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsProcess Intelligence

Stop Steering on Output

RealAINov 22, 20237 min read
Process MiningFinancial ServicesLendingOperationsData Readiness

Every operating review I sit in opens with the same page: results. Volumes booked, conversion percentages, a funnel drawn as four or five tidy rectangles with a number under each. Nobody in the room disputes the figures. The dispute worth having is about their timing.

Output is a record of decisions already taken. By the time a conversion rate lands in a monthly pack, every case behind it has finished. The one that failed at document upload three weeks ago sits in the denominator, unreachable. You can learn from it the way you learn from a post-mortem. Nobody can act on it, least of all the person who was handling it, who is now working the next batch under identical conditions.

The sharpest version of this argument I have worked from sits in a cross-channel process review carried out for a large European cooperative banking group. The review put it in one line: move from steering on output to steering inside the process. See concrete customer behaviour rather than inferred behaviour, close the feedback loop, and do it during the rebuild, because while you are replacing a chain your output measures describe a chain that no longer exists.

The instrument the review chose

Given a mandate to improve conversion on a rebuilt cross-channel process, the obvious thing to measure is conversion. The review measured lead time to first appointment instead, one figure per local unit, computed from real transaction and event data rather than from surveys or self-reported management information.

That choice is the argument in miniature. A conversion rate is terminal. Lead time to first appointment is not: a case that has been waiting six days for an appointment is still a case, and somebody can still ring that customer today. The measure and the intervention window overlap, which is what makes a number steerable. Almost nothing on a standard results page has that property.

The distribution is the part people remember. One bar per local unit, sorted descending, each coloured up to a fixed reference level and in a second colour above it, so the upper section of every bar was literally the waste. Same brand, same product, same period. Read off that chart, the slowest unit took around thirty-six days to reach a first appointment and the fastest a little over one, with a median just under nine, a mean slightly above, and the reference line at roughly nine and a half. About four in ten units carried an excess segment.

I should be plain about those figures. The slide states the axis, the units and the label on the best-performing tail, and nothing else. Every number I have quoted was measured off the chart against its own gridlines, accurate to perhaps a third of a day. They are reads, not quotations.

Read the shape, not the ratio

The gap between best and worst is about thirty to one, and that is the number a slide wants to carry. It is also the least useful number on the chart.

Rank one measured about thirty-six days. Rank two measured about twenty-two. A fourteen-day cliff between the two worst units, in a distribution whose median is under nine, means the eye-catching ratio is a statement about one local unit. Fixing that unit is a local programme with specific people and a specific queue. Moving the four in ten units sitting a day or two over the line is a different programme, and it is where the recoverable time lives.

This matters practically, because whoever builds the monitoring will build alerting on top of it, and a rule tuned to the headline spread fires on the same outlier forever. Sorted internal distributions have to be read for shape before extremes. Their compensation is that they remove an excuse: an external benchmark invites the answer that our market is different, and a target set by your own units does not.

Why per-channel instrumentation cannot see it

The finding that reframed the engagement was not a rate at all. It was an ordering: customers made contact with the bank first, and went online afterwards. In the digital-first climate of the time this inverted the assumed sequence, in which people orient online and then book an appointment.

No single channel could have produced it. Web analytics sees a session beginning, telephony sees a call beginning, the CRM sees an opportunity created. Each is correct about itself and structurally incapable of noticing it is the second step. The finding required stitching mobile events, online clickstream and CRM click data into one case-level model, which is unglamorous plumbing no channel owner is incentivised to fund.

One caveat belongs with it, and it is ours rather than the review's. The online surface the review puts beside that finding is the group's own product page, and it carried an acquisition funnel down the middle and a servicing rail for existing borrowers down the side, so some of that early contact will have been current customers changing a payment. The direction holds. The magnitude was never stated and I would not defend one.

~36 days
Slowest local unit to first appointment (chart read)
~1 day
Best-performing unit, same measure (chart read)
~9.5 days
Fixed reference level splitting each bar at its excess
~4 in 10
Units sitting above that line

The designed process and the discovered one

The programme had a journey map: eight stages, fifteen touchpoints, three owning functions splitting those stages one, four and three, which puts two internal handovers directly inside the customer's path. It is a good map. Discovered from the event log, the same journey came back as roughly thirty-five to forty activities with dense crossing edges and long loop-backs from late in the process to early in it.

That illegibility is not a failure of the drawing. It is the finding. The tidy version is what the organisation believes it does, and any control system built on it steers a process that exists only in a document.

Scoped to a single local unit, the discovered model resolved into six activities: sales opportunity, inbound call, outbound call, appointment, quote, purchase. Two details in that small graph are worth more than the whole map. The appointment is reachable two ways, directly from the opportunity and via an outbound call, so the phone step is a route rather than an overhead. And a visible path leaves the funnel without reaching purchase, which is the population every output measure quietly drops.

A conversion rate is an obituary. It tells you, accurately and too late, what happened to cases whose fate was settled weeks ago by people who have moved on to the next ones.

In-process measurement is harder, and that is the objection

Output measures are cheap because a terminal event is unambiguous. Sold or not sold needs no committee. In-flight measures need three things most estates do not have: a case identity that survives a customer moving between channels, a timestamp on every activity rather than only on stage transitions, and a vocabulary of activity names that means the same thing in every unit.

The review named the third as its own precondition, and I have not read a more honest sentence in a consulting document: the comparing and the steering only get stronger once a uniform way of working is embedded in the CRM. Treat that as a data readiness statement and it is the same lineage problem that stalls machine learning work everywhere else. A benchmark built on inconsistently recorded events produces confident, wrong league tables, and nothing downstream detects the inconsistency, because there is nothing internally contradictory about it.

The organisational half of the recommendation is worth as much. Rather than a central analytics function, the proposal asked for process indicators embedded in the reporting already running over the straight-through processing chain, in the existing internal comparison programme, and inside the team that ships the product. That is measurement placed where the change gets made rather than where the reports get written. To be clear: the distribution was measured, the embedding was only proposed.

What I would carry forward

One line in the approach proposed alongside the findings is worth more than the rest of it. The pilot was scoped at six to eight weeks, and a monitoring baseline sat in the deliverables list beside the findings rather than after them.

That is the cheapest upgrade available to almost any analytics programme, and it is the opening block of a RealAI Platform engagement for exactly that reason. Instrument the before-state as part of the work rather than as a follow-on nobody funds, and the next cycle can prove whether the change worked. Skip it and the counterfactual is gone, which is how organisations end up with a portfolio of pilots and no evidence.

The modelling end has never been easier. Deployment is close to routine in most stacks, feature stores are ordinary infrastructure, and the first cautious language-model pilots crossing my desk stand up in days. None of that touches the timing problem. A model trained on output learns what already finished and reports on what already finished. Point the same effort at what a process is still holding, and you get something a person can act on before the case closes.

Figures are as recorded in a cross-channel process review carried out for a large European cooperative banking group: its lead-time distribution, its sequence findings and the pilot scoping proposed alongside them. Lead-time values were measured off the published chart against its own gridlines and are approximate. The review produced findings and recommendations, not delivered results; the permanent process indicators were proposed, not built. Reading it as a measurement-timing problem is ours.

A conversion rate is an obituary. It tells you, accurately and too late, what happened to cases whose fate was settled weeks ago by people who have moved on to the next ones.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?