Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsData Strategy

One Page of Methods, Four Steps of Something Else

RealAIFeb 17, 20238 min read
Data StrategyBankingAnalyticsProcess MiningData Quality

I went through a stack of old proposals this month looking for something unrelated, and lost an afternoon to one of them. It was written for an Austrian retail and corporate bank, the better part of a decade ago, at a point when standing up an analytics function still involved a serious conversation about where the data would physically sit.

The back half of the document is a delivery method. Near the middle of it sits one page listing every analytical technique on offer, filed under five headings. That page stopped me, and not because anything on it is wrong. It stopped me because of how short it is, and because of what the pages either side of it spend their length on.

What was actually on the page

Under one heading: pivot tables, cubes, interactive dashboards. Under another: outliers, trend, correlations. Under a third: repetition of patterns, frequent item sets, association rules, process mining. Under a fourth: k-means, hierarchical clustering, density-based clustering. Under a fifth: decision trees, regression, neural networks. A row of vendor marks sits on the page too, unreadable in the copy I have. The whole page carries a stamp marking it illustrative, so nobody would mistake it for a committed stack.

Neural networks gets one line. Not a heading, not a section, not a strategy. One entry of sixteen.

A few pages earlier the same document puts a question to the room: which analytical method suits a mortgage case. Six options are offered. Neural networks is one of them. Big data is another, listed as though it were a method rather than a statement about volume. No answer is printed. The exercise was to make the room argue about method selection before anyone sold them one.

What survives best from that page is not the inventory but the filing system, and the filing is only half consistent. Two of the headings are plain method families, the sort of label a textbook contents page would carry. The other three describe what you are doing to a question. That is the half worth keeping, because sorting by the shape of the question rather than by algorithm lineage makes the method a dependent variable of the question rather than the reverse.

Four steps that are not about algorithms

The catalogue is one step of five. The four around it are not about algorithms at all.

The method opens with a business question, deliberately, before any dataset is located. Two are printed as examples of the right kind: how different customer segments would want to be served, and what the next product to sell is. Neither is answerable from a schema, and both have money attached.

Then data preparation, defined in one sentence with five verbs in it: locate, extract, understand, clean and combine the relevant data, in line with data governance. Understand sits between extract and clean, which is the part worth noticing. You are told to work out what a field means before you decide which of its values are wrong, the opposite of how most data quality work gets sequenced. Governance sits inside the step rather than beside it as a compliance workstream.

Then the analysis. Then a step whose written definition amounts to a single instruction: validate the conclusions with experts, meaning the people who run the business argue with the finding before it counts. Then a final step about embedding the outcome into how the business operates and monitoring progress against benefit targets, continuously.

Count the proportion. One step in five is the modelling. Four in five are about whether the number means anything and whether anyone acts on it.

The chart that gives the game away

The best page in the whole method is the worked example of the validation step, and it is a bar chart. Conversion ratio plotted across a set of local branches, with an average line drawn at twenty-two percent. The average is a demonstration figure, no branch on it is a real branch, and the document says so on its face.

What matters is the annotation beside it. The observation reads: large variation in conversion rate between branches. The prescribed next move is to invite the best and worst performers into the same room and make them explain each other. Then two candidate causes are printed underneath.

The first is that a branch does not register its activities properly in the customer system. The second is that a branch has a different customer base.

Only the second of those is a performance story. The first is a measurement artefact, and it is listed first.

That ordering is the entire discipline compressed into two bullet points. Before you tell a branch manager their conversion is poor, you are required to entertain the possibility that their conversion is fine and their logging is different. Nothing in the data can tell you which, because both produce the same bar. Only a definition, written down in advance and applied the same way in both branches, can separate them.

Getting the same answer twice

You find it everywhere in the document.

The data preparation page puts its own question to the room: what is the most critical factor in getting the right data to prove a hypothesis. Four options. Extract, transform, load. Validation sessions. Database keys. Structured data. Three are plumbing and the fourth is a meeting. Not one is a model, or a technique, or anything from the catalogue two pages later. The question asks what stops a hypothesis from being provable, and every answer on the table is about whether the data can be assembled the same way twice.

The validation step is the same idea wearing a different hat. An insight is treated as a claim about the business until somebody who runs that business has contested it. The page even makes the choice of chart part of the method, on the grounds that the visual decides whether the expert engages at all. A finding nobody argues with has not been validated. It has been ignored politely.

The last step is not analysis either. It is cadence: a recurring conversation, a named sponsor, a live tracker, a targeting motion. Four organisational mechanisms, none of them analytical, because the analysis finished two steps earlier and every unit of value after that comes from operating rhythm.

Our vocabulary has moved on. We say lineage now, and feature store, and data contract, and we build machinery to enforce what this document expected a person to write down. The thing being described has not moved at all. It is a number that can be reproduced by somebody who was not in the room when it was first produced.

16
Distinct techniques in the catalogue, across five headings
1
Lines given to neural networks
1 of 5
Steps in the method that are the analysis
5
Verbs in the definition of data preparation

Why it reads differently now than it did then

Two things have changed since that document was written, and they pull in opposite directions.

The expensive half of the list got cheap. Fitting a model is close to routine in most stacks. Pipelines are assembly work. Deployment is a config file and a review, and the machine learning operations tooling that used to be a project is now a dependency. The first language model pilots crossing my desk this winter stood up in days rather than quarters. The half of that catalogue that was hard to staff a decade ago is the half you can now rent by the hour.

The other half never got cheap, because it was never technical. Nobody has automated the argument about what a resolved case is, or the decision about which of two systems is the record for a customer, or the moment a branch manager says the number is wrong. Those cost what they always cost, and they cost it in senior attention, which has not become more abundant.

So the ratio inverted. One page of techniques once sat inside four pages of definitional work because the techniques were scarce and the discipline around them was the price of admission. Today the techniques are abundant and close to free, and the definitional work costs what it always did. Those four pages are now almost the whole cost of an analytics programme, and they are the four every roadmap I see quietly skips, on the reasonable-sounding grounds that the modelling is the hard part. It is why a RealAI Platform engagement opens on definitions and joins rather than on model selection.

The page that shows a miss

The document closes on a benefit tracker. It runs a sales funnel and a conversion ratio into a target, then a stated potential, then what has been realised to date, then the benefit booked. Four stages on one artefact, and every number printed on it is a demonstration rather than a result.

The shape is what I would keep. The potential in the example is eleven points of conversion. The realisation shown beneath it is seven. Whoever drew that sample chose to draw a shortfall, and to keep it on the same page as the promise, where the sponsor has to look at both at once.

That is rarer than anything in the technique catalogue. Most benefit dashboards are built to show a target being approached. This one is built to show the distance between what was modelled and what arrived. It is also the only part of the method that cannot be satisfied by declaring a pilot a success.

Re-read as a whole, the document is not really a list of what an analytics function could do. It is a list of the conditions under which anything it does can be believed. The techniques dated fast. The conditions did not date at all.

Drawn from the delivery method set out in an analytics proposal put to an Austrian retail and corporate bank: its technique catalogue, its five-step sequence and its worked examples. That document is a proposal, and its charts are marked illustrative or sample on their own faces, so every figure quoted here is a demonstration rather than a client outcome. Reading the shape of the method as a statement about where analytics work costs money is ours.

Sixteen techniques fitted on one page. The four steps around that page were about whether the number could be reproduced, and whether anyone acted on it.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?