Almost every customer journey programme I am asked to scope opens with a collection plan. A tagging sprint, a tracking schema, a panel to fill the gaps, and weeks of instrumentation before anybody looks at anything, on the reasoning that you cannot analyse a journey you have not yet recorded.
That reasoning is usually wrong, and it is wrong in a way that costs a quarter.
I keep returning to a demo pack prepared for a European travel and tourism group that was moving from a branch network and one main website towards several further web shops. The pack made one claim that has held up better than anything else in it. The records needed to reconstruct the customer's real path were already being written. Millions of transactional data records captured by each process, in the pack's own phrasing. The gloss is mine and it is the part that matters: those records are a byproduct of the business simply running. Nothing new to instrument, nothing new to ask customers. The path was already on disk, spread across the transaction systems of each channel and the CRM.
That is what journey mining is. Not a research method. A reading method, applied to exhaust data.
The page was the wrong unit of analysis
The group was not short of analytics. It was optimising individual web pages using a web analytics tool and user experience expertise, which was the correct practice of the day and is still a large part of the practice. The pack's diagnosis was not that the tooling was bad. It was that the method had a ceiling built into it: optimising single web pages would have limited impact on the conversion ratio of the end-to-end web experience.
That single sentence carries the whole methodological argument. A page test tells you which of two pages performs better for the traffic that reached that page. It tells you nothing about the path that produced the traffic, and nothing at all about the customer who begins online, phones to ask a question, and finishes in a branch. Optimise every page independently and you get a set of local optima, each one true, none of them adding up to a better journey.
Journey mining moves the unit of analysis from the page to the path, and it can do that only because the path was recorded somewhere, in fragments, by systems built for another purpose. Order records, contact records, session records, case records. Each one is a poor analytics artefact and a perfect event: a thing that happened, to a case, at a time.
A drawn map is an assertion, a mined graph is a measurement
The pack's proof exhibit was a journey map its authors carried across from retail banking work of their own, drawn as a grid. It is their track record, not mine and not RealAI's, and I am reading it here as an artefact rather than claiming it. Lifecycle stages ran across the top, three channel lanes ran down the side, and a mortgage buyer's path was plotted across the grid, starting from a life event, passing through a calculator on a tablet, a call-back request from the web, a signed contract and a set of keys, and ending with the customer recommending the bank to friends. Below the path sat a value-driver band, three measures named down its side and only one of them actually drawn: a wedge that begins wide at the top of the funnel, narrows to a single point at purchase, then widens again through the post-purchase phase and into advocacy, with the crowd of figures thinning to one and then filling back out.
It is a good exhibit and it makes a real argument, which is that the same customer crosses channels inside a single purchase. It is also a drawing. The stages were chosen. The lanes were chosen. The touchpoints were chosen. The wedge has no axis, no scale and no data label, so it carries a direction and not a magnitude. Having drawn that kind of map more than once myself, I would put it plainly: a workshop produces a hypothesis about behaviour with a very high production value.
What the exhaust looks like once you read it
The next exhibit in the same pack is the interesting one, because it was not drawn. It is a discovered process graph: twenty-nine activity nodes, directed edges between them, every node labelled with an absolute event count, every edge labelled with a transition frequency, edge thickness scaled to volume.
The counts have a shape worth describing. The quietest nodes sit near ten thousand events, the busy ones run into the hundreds of thousands, and one screen carries roughly 1.42 million, which is around three times the second busiest node and well over a hundred times the quietest. The transition labels spread just as widely, from three figures on the thinnest edges to roughly half a million on the heaviest. Two caveats, both mine and both load-bearing. The dataset behind that exhibit is not attributed anywhere in the pack, and its activity names come from a financial services CRM, so it is not the travel group's data and I will not present it as such. And I read those figures off a low-resolution embedded image, so treat the digits as approximate and only the ratios between them as solid.
The ratios are the finding. A real journey graph does not look like a funnel. It looks like a dense network in which the same pair of screens is joined by several competing routes, in which the busiest screen absorbs more than a hundred times the traffic of the quietest, and in which the long tail of rare paths is where the abandoned cases live. No workshop draws that, because no one in the room remembers it.
- millions
- transactional records per process, already captured
- 29
- activity nodes in the discovered process graph
- ~1.42M
- events on the busiest activity, over 100x the quietest
- 6
- KPIs named as targets of the approach
Averages hide the segments that matter
The pack pairs the process graph with a second exhibit, a scatter plot separating three customer clusters, labelled only as A, B and C. The axes carry no titles and no tick values, so what those two dimensions measure is not recoverable from the document, and I am not going to guess. What is stated in words on the slide is the design intent: analyse the end-to-end journey by client segment, not only in aggregate.
That intent matters more than the picture. A journey graph built over all traffic is an average of several different behaviours, and the path taken by a customer who arrives with a fixed destination and a date is not the path taken by one who is browsing an idea. A single graph blends the two into a route that nobody walked. Mine per segment and the competing paths separate. Mine in aggregate and you have replaced one fictional journey, the drawn one, with another.
A drawn journey map is a set of claims about how customers behave. A mined journey graph is a count of how they behaved. The two look similar on a slide and they are not the same kind of object.
What it is actually for
The method in the pack runs in three steps and the third is the one that decides whether any of it was worth doing. Capture click behaviour from the transaction systems of the channels and the CRM. Analyse the end-to-end journey, by segment. Then hand the result to online marketing and to the people who build the site, as input to their work. The output is operational rather than advisory. Six measures were named as the targets: conversion rate, online marketing cost per lead, throughput time, net promoter score, margin and market share.
Two of those deserve a note. Throughput time is a process mining measure that page analytics cannot produce at all, because it needs a case identity that survives across channels. Cost per lead is the measure that punishes local optimisation hardest, since spend can buy a better page while the path behind it stays broken.
I should be exact about what this pack was. It was a demo, prepared to show what the method could do. It did not price a programme, it did not report a result, and it contains no delivered outcome for the travel group. The argument is what I am carrying forward, not a claim of impact.
Where journey mining actually fails is never volume. It is case identity and clock discipline. Does a branch record carry an identifier that ties it to the same person's web session, and who owns that mapping. Are the timestamps written in one clock or in four. Is the event that marks a purchase written once, or once per system under three different names. Those are lineage questions, answerable in weeks by people who already work there, and they are the first thing a RealAI Consult engagement puts on the table when a journey brief arrives, ahead of any modelling. Straight-through processing carries the same dependency: you cannot automate a handover on a case identity that two systems disagree about.
The modelling end of this has never been cheaper. Pipelines are assembly work, deployment is close to routine in most stacks, and the first language model pilots crossing my desk stand up in days. The scarce thing is a description of the customer's real path that somebody counted rather than drew. That description is being written into your transaction logs while you read this, and it is the cheapest asset in your estate because you are already paying to produce it.
Drawn from a customer journey mining demo pack prepared for a European travel and tourism group. That document is a demo: it stated an aim and showed a method, and it reports no delivered outcome for that group. The process graph and segmentation exhibits inside it carry no data source, and their activity names indicate a financial services system rather than the travel group's own, so nothing quantitative here is presented as that group's result. Figures read from embedded images are approximate and are used for their ratios only. The reading of exhaust data as the cheapest available grounding for journey work is ours.
“A drawn journey map is a set of claims about how customers behave. A mined journey graph is a count of how they behaved. The two look similar on a slide and they are not the same kind of object.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
