Every customer journey proposal I read ends in the same place, and the place is not a system. It is a readout.
That is not a complaint about consultants; I have written enough of these to have no standing to make it. The analysis is genuinely hard and it gets done, and then it is compressed into something a room can absorb in forty minutes. Whatever survives the compression is what the organisation gets.
I have been re-reading a session pack put together for a European travel and tourism group moving from a mostly physical selling estate into online selling, and running its web work the way everybody did: page by page, instrumented with a mainstream analytics package and steered by user experience judgement. Its diagnosis was that page-level work has a ceiling, because the page is the wrong unit of analysis once a customer's path crosses channels. Its method was three steps. Capture click and transaction events out of the channel systems and the CRM, which are already writing them. Reconstruct the end-to-end journey from that event log and cut it by customer segment. Then hand the result to online marketing and to the website developers.
Then it named the six measures the work was supposed to move: conversion rate, online marketing cost per lead, throughput time, net promoter score, margin, market share.
Naming them is the easy half.
Six measures, six different clocks
Put the six side by side and they do not belong to one instrument. Conversion rate falls out of the same event stream that produced the analysis and can be read daily. Online marketing cost per lead needs spend data from another system and a rule about what counts as a lead. Throughput time is readable from the log if, and only if, somebody has fixed the start and end events. Net promoter score arrives on a survey cadence and is a sample. Margin comes out of a finance close. Market share comes from outside the business entirely.
Two of the six live in the log that the finding came from. The other four live somewhere else, on slower clocks, owned by people who were not in the room. That is my reading rather than the pack's, but the consequence is not subtle: a finding drawn from the event log can be checked against two of its six targets by the team that drew it, and against the other four only by somebody else, later, on an instrument that cannot see the change.
This is why so much journey work produces an agreed diagnosis and no measured effect. Nobody lied. The measurement simply lives on a different clock from the change, and the two never get joined up.
A drawn map is an assertion, a mined graph is a measurement
The pack does one thing I still think is the right instinct: it puts a drawn journey map and a discovered one in the same document.
The drawn one is the familiar object: stages laid left to right, channel lanes beneath, a customer's path traced through as a clean sequence of moments. It comes from a prior engagement in another industry, shown to make the point that the method travels. Read as a picture of how a business believes its customers move it is useful. Read as data it is an assertion with a designer's confidence.
The discovered one is what an event log says. Twenty-nine activity nodes, each carrying an absolute count. Directed edges between them carrying transition frequencies from three figures to the high hundreds of thousands. Loops. Several competing routes between the same pair of screens. Two screens carrying the traffic: one reads over a million events, the heaviest single transition between the two carries more than eight hundred thousand on its own, and most of the remaining nodes read in the tens of thousands. The node labels are application screen names, so the log is screen-level navigation, which is why journey reconstruction is possible at all.
Those two exhibits are the argument, and it survives without a single number from either being precise. A drawn map holds a handful of states because a human drew it and a human can hold a handful. A mined graph holds whatever the customers did. The gap between them is the long tail of real behaviour, and it is the part any change you ship has to survive.
What gets lost between the graph and the bullet
Now do the handoff.
The graph goes on a slide. It will not read at that size, so it gets cropped and the finding is written underneath as a sentence, in the shape of: this segment loops between these two screens before dropping out, so reduce the friction there.
A developer or a marketer picks that sentence up two weeks later and has to turn it back into something actionable. Which two screens, at what volume, for which customers, entering from where. None of that is in the sentence, so they reconstruct it, from the picture, from memory of the session, from their own knowledge of the site. Their reconstruction is competent and it is not the same one. The filter differs slightly, the segment boundary differs slightly, the start event differs slightly.
Then they ship a change against their version of it, and the original analysis is never re-run to check that what they changed is what was measured. The finding did not get ignored. It got quietly rewritten by the person acting on it, which is much harder to notice, because everyone involved will tell you truthfully that they acted on the analysis.
That is the failure I would put ahead of every modelling concern in this domain. The analysis was right. The handoff was lossy, and nothing in the process was watching the loss.
The segment cut goes first
The method insists the journey be cut by segment before any conclusion is drawn, and the worked example shows three clusters. Its axes are unlabelled, so what separates them is not recoverable and I will not guess. The principle stands without that: an aggregate graph over mixed segments describes a customer who does not exist, and optimising for that composite optimises for nobody.
Segmentation is also the first casualty of the readout. One slide holds one picture, so three segment-specific graphs become one blended graph, or one representative graph with a note that the others differ. By the time it reaches the person changing the checkout, the note has gone. The change lands on everybody, gets measured on a pooled conversion rate, moves it ambiguously, and the programme concludes that journey mining is interesting but hard to prove out.
Hand over the query, not the conclusion
The repair is unglamorous and mostly not analytical.
Ship the event definition alongside the finding: which system emits the start event, which the end, what a case is, what closes it. Ship the segment assignment as a stored rule rather than a description, so the cohort a change targets is the cohort the effect is read on. Ship the query itself, in a form that runs unchanged next month, so before and after come from one instrument rather than two. Name one owner per measure, a person rather than a function, and say plainly that four of the six report on a slower clock than the change.
None of it needs a data scientist, which is the uncomfortable part, because all of it has to be true before hiring one pays. This is the block a RealAI Consult engagement runs before anyone scopes a model, and it is why we ask which measure you intend to move, and who reads it, before we ask what analysis you want.
The straight-through version is the goal: discovered path defect, shipped change, same query re-run on the next window, movement attributed or not. Every step of that loop was manual in the pack I have been reading, and manual is why it would have run once.
A finding that arrives as a sentence has to be turned back into a query before anyone can act on it, and the person doing that reconstruction is never the person who ran it the first time.
Why this matters more now, not less
Discovery has got cheap. Event logs are larger and easier to reach, mining tools are close to commodity, and the first language-model pilots crossing my desk can already turn a process graph into a readable paragraph in seconds.
That capability is pointed at the wrong end of the problem. The summary was never the bottleneck, and a machine-written one is as lossy as a human-written one, delivered sooner. The part worth automating is the re-derivation: keep the query, the cohort rule and the event definitions attached to the finding as it travels, so whoever acts on it acts on the object that was measured, and the measurement can be repeated the day after the change ships.
Get that right and the six measures stop being an aspiration slide: two become a control loop and the other four a slower audit of it. Get it wrong and you will do excellent analysis for years and never be able to say whether any of it worked.
Source: a pre-sales session pack on cross-channel journey mining prepared for a European travel and tourism group: its situation read, its three-step method, its six named target measures and three rendered exhibits. It is a proposal, not a delivered engagement, and reports no baseline and no result for any of the six. The process graph and the segmentation exhibit carry no stated data source, so they are treated here as illustrative and attributed to nobody. Reading the pack as a handoff problem rather than an analysis problem is ours.
“A finding that arrives as a sentence has to be turned back into a query before anyone can act on it, and the person doing that reconstruction is never the person who ran it the first time.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
