Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsRetail

The Page Is the Wrong Unit of Analysis: Why Conversion Work Stalls at Local Optima

RealAIJan 6, 20268 min read
RetailCustomer ExperienceAgentic AI

In February 2014 a European travel and tourism group running a large branch network alongside its main website was preparing to open additional web shops. Its conversion work looked orderly. Pages were being optimised with a commercial web analytics suite and user experience expertise, one page and one test at a time. The demo material put in front of the company said something none of those page reports could say on their own: optimising single web pages would have limited impact on the conversion ratio of the end-to-end web experience.

That single sentence is the whole argument, and it has aged better than most things written about digital commerce in that decade. The company was not short of analytics. It was short of a unit of analysis that matched the way its customers actually behaved, which was to cross a website, a branch counter and soon a set of new web shops inside what they experienced as one decision.

The ceiling you cannot see from a page

Page optimisation is a well-behaved discipline. You pick a surface, form a hypothesis, split the traffic, and keep the winner. Every step is measurable and every win is real. The problem is what the method holds constant. A page test can only find the best version of a page. It cannot tell you that the page should not have been there at all, or that the abandonment happening on it is the tail end of something that went wrong two channels earlier.

That is the definition of a local optimum. Each surface improves, the aggregate barely moves, and the programme slowly runs out of hypotheses that matter. Teams in that position usually conclude they need better testing tooling. The 2014 diagnosis pointed somewhere else: the constraint was not the quality of the tests but the size of the thing being tested.

Loading exhibit
Exhibit 1Where a page win actually landsEvery tile is one cell of the reference journey grid: eight lifecycle stages by three channel lanes. Tile height is conversion at that cell; the plane across the middle four stages is end-to-end conversion, the product of them.

Exhibit 1 puts that arithmetic on a surface you can walk around. Apply the same page-sized win to any single tile and it rises alone while everything next to it stays where it was, which is what a local optimum looks like when you give it a floor plan. The plane barely registers the change, because it is the product of four stages and a win in one of them is diluted by the three it has to pass through. The same win is worth 2.15 points of end-to-end conversion in the branch lane at proposal, a cell no page test can reach; 1.56 points on the best of the four cells a page test can see; and nothing at all in the twelve cells that fall outside the four stages the ratio counts, where a 158 percent gain on the call-centre referral cell leaves the conversion number exactly where it was. Nothing in a page report distinguishes those three cases, because a page report has no way to say where the page sits.

For a retailer with a large physical estate the frame is wrong by a wide margin. A website, a branch network of that size and new web shops on the way add up to a surface count no page-by-page programme can reason across. The customer sees one attempt to buy a holiday.

What the alternative was designed to do

The proposed replacement was process mining applied across channels, a technique the 2014 material credits to a joint lineage between a technical university and a consulting firm, and describes as a way to reconstruct how processes work in reality rather than how they were designed to work. Applied to commerce, that means reconstructing the customer journey from the transaction records the business is already generating instead of drawing it in a workshop.

The economics of that idea were attractive then and are more attractive now. It pointed at the millions of transactional data records that digital processes capture as a by-product, and argued that they were enough to visualise and optimise each customer journey end to end. No new instrumentation project, no survey panel. The evidence of what customers did was already sitting in channel transaction systems and CRM, unread.

The design had three steps. Capture click behaviour from the transaction systems of each channel and from CRM. Analyse the end-to-end journey with process mining, cut by client segment rather than in aggregate. Then feed the result to online marketing and to website developers as concrete input for change. Six KPIs were named as the scoreboard: conversion rate, online marketing cost per lead, throughput time, NPS, margin, and market share.

Two design choices in that pipeline deserve more credit than they usually get. The first is the insistence on segmentation before conclusion. An aggregate process graph over mixed customer types describes an average customer who does not exist, and optimising for that average is another way to stall. The second is the named recipient. Step three does not hand findings to a committee. It hands them to the two roles that can change the thing being measured.

A journey that starts before you and ends after you

The reference journey map used to explain the approach came from a prior engagement in another sector, and it is worth reading for its shape rather than its content. It ran eight stages across three lifecycle bands, from awareness and orientation through inform, advice, proposal and buy, then on into service and maintenance, renewal and referral. Cutting the other way were three channel lanes: online, call centre, branch. Eleven labelled touchpoints were plotted across that grid.

The interesting part is where the map begins and ends. It opens at a life event and a search that happens entirely outside the institution, before any contact at all. It closes past the transaction, at renewal and at a customer recommending the business to friends. Underneath the journey sat three named value drivers, volume, margin and satisfaction, plus a customer experience band, and three outcome measures the map tied back to them: market share, conversion rate and loyalty, all resolving into a single line about increasing profitable revenue.

Most conversion analytics, then and now, starts at first-party session start and stops at checkout, which excludes both ends of that map. It is a measurement boundary drawn around the systems a company happens to own, and it quietly becomes the boundary of what the company believes the journey to be.

The same failure mode, on newer surfaces

The reason to revisit a diagnosis written in 2014 is that the failure mode it names has multiplied rather than faded. In retail the surface count has grown well past a website and a branch: marketplace listings, retail media placements, a loyalty app, a store associate holding a tablet, a subscription flow, and now a conversational assistant that answers questions the product page used to answer. Each surface has an owner, a dashboard and a test backlog. Almost none of them share a unit of analysis.

AI has made the trap easier to fall into. Prompt-level optimisation is page-level optimisation with a newer name. A team can spend a quarter improving the phrasing of a shopping assistant, measure a genuine lift in assistant satisfaction, and never discover that the customers using it are the ones who walked out of a branch two days earlier holding a quote they had no way to compare against anything. The assistant improves. The end-to-end conversion ratio does not.

What has genuinely improved is the raw material. Event logs are cheaper to retain and easier to join, process discovery is commodity capability rather than research, and the 2014 argument that the data already exists is now close to unarguable for any retailer with a loyalty programme and a CRM.

What agents change, and what they do not

Agentic AI changes three specific things about this work, and it is worth being precise about which.

The first is cadence. The 2014 pipeline was a batch process with a human in every loop: pull the logs, mine once, cut by segment, write a document, hand it over, wait for a release. An agent can hold that pipeline open. It can re-run discovery over the event log on a schedule, compare the current journey graph against the previous one, and raise the specific change worth looking at, such as a transition between two states that has thinned out since last week or a new loop that has appeared between the web shop checkout and the call centre. The analysis stops being a study and becomes a standing instrument.

The second is the handoff. Step three of the original design was the honest weak point of every analytics programme: insight arrives, and then organisational latency eats it. An agent can carry a discovered path defect through to a drafted change, a test design and a measured outcome, with a human approving the change rather than assembling it. That is where most of the recoverable value sits, because the diagnosis was rarely the bottleneck.

The third is breadth. Agents can hold context across the stretches of the journey that funnel instrumentation structurally cannot see, the pre-contact browsing and the post-purchase service life, provided the events exist and the customer has agreed to their use. That last condition is doing real work. Cross-channel joins that were a technical question in 2014 are a consent and data-minimisation question now, and branch and call-centre events tend to have the thinnest lawful basis of all. Design the join and its legal grounding before anyone builds the model.

Breadth has a physical side too, and for a retailer with this shape it is the largest blind spot of all. The branch lane on that journey map is where the longest conversations happen and where almost nothing machine-readable is produced: a printed quote marked up at the counter, a signed booking form, a brochure page a customer pointed at, a confirmation scanned and filed after they left. In 2014 that lane could only enter an event log if a member of staff retyped it, which is why branch behaviour tends to arrive as a monthly total rather than as a path. Computer vision changes what is available there. Document models read counter paperwork and scanned confirmations into structured events with timestamps and identifiers, which is the difference between a branch that appears in the journey graph as a step and one that appears as a gap. The consent question travels with the image, and it is stricter for a document carrying a traveller's details than for a click.

Put those pieces together and the original design stops being a study and becomes a running pipeline. Capture reads from the channel transaction systems, from CRM, and now from the branch documents. Discovery re-mines the log and cuts it by segment. A drafting agent turns a thinned transition, or a new loop between the web shop checkout and the call centre, into a proposed change with a test design attached. Measurement reports back against the same scoreboard and reopens the loop. Each stage hands structured output to the next rather than a slide, humans approve rather than assemble, and at that point agents are carrying a real share of how the commercial operation runs instead of producing a document about how it might.

What does not change is more important. You still need a real event log with stable identity across channels, because an agent reasoning over a broken join will produce confident nonsense faster than a human would. You still need segments, because a model trained on pooled behaviour learns the mean path and optimises for nobody. You still need one person accountable for the number. And you still have to accept that a mined journey from a real business is not a tidy eight-box diagram. It is dense, looping and full of competing routes between the same two states. An agent handed the idealised journey as its operating policy will perform well on the path that was drawn and fail on the long tail that was not, and the event log is the only place that tail is written down.

How much of that pipeline should run unattended is an engineering decision rather than a slogan. The stages safe to make autonomous are the ones where a wrong answer is cheap and the correction is fast: re-running discovery over last night's log, flagging a transition that has thinned, preparing the segment cut. The stages that change what a customer sees stay behind an approval gate until the loop has a record worth trusting. What separates the two is the harness the loop runs inside: the identity join across web, call centre and branch, the consent rules attached to each source, the segment definitions, an evaluation set that can say whether a proposed change was actually good, and the gate a named human signs. Building that harness is most of the work, and it is why journey work that arrived in 2014 as a one-off document can now run at the cadence of the event log rather than the cadence of a study.

Where to start

Three moves, in order. Name the unit of analysis before funding any more tests, and make it the journey rather than the surface. Inventory the event data you already hold across channels and CRM, and treat the gaps in identity and consent as the first work package rather than a later problem. Then close the loop by giving the discovered defects an owner and a KPI, so the analysis is judged on whether the number moved rather than on whether the document was good.

The diagnosis from February 2014 needed no new technology to be true. It needed a change in what the organisation chose to measure. That part is still the hard part, and it is still the one that decides whether the AI sitting on top of it is worth anything.

What has changed is how fast the answer arrives once the choice is made. Build the loop and the harness around it, and the gap between noticing a broken path and shipping a tested change is measured the way a release is measured, not the way a study is.

A page test can only find the best version of a page. It cannot tell you that the page should not have been there at all.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?