In July 2014 a global management consultancy running an internal knowledge and information management programme circulated a draft change plan for its own firm. Version 0.1, marked for discussion within the team. Three slides in, before any change theory, sits a status table. Seven rows, five columns, one date stamp. It covers engagements closed between 1 January and 25 June that year, snapshot as at 2 July, and reports what share of proposals, shared client documents and track records reached a place where a colleague could find them.
Read down the rows it looks like a league table, and league tables invite the wrong response, which is to congratulate the top and lean on the bottom. Read across two axes instead and the same seven rows become a rough decomposition of where behavioural variance actually lives in a firm that has already bought and deployed the tool.
Seven rows, two axes
The table carries a header row naming its columns, so the numbers read cell by cell without guessing at layout. Call the two practices Practice One and Practice Two. Practice One appears four times. Denmark at 68% of proposals captured, 17% of shared client documents, 8% of track records, over 59 closed engagements. Sweden at 39%, 26% and 17% over 23. Norway at 24%, 0% and 35% over 17. The Netherlands at 52%, 16% and 16% over 31.
Practice Two appears three times. Denmark at 77%, 14% and 11% over 73 engagements. Sweden at 67%, 17% and 28% over 18. Norway at 36%, 0% and 14% over 14.
Three countries carry a row in both practices. That is the two-by-two: hold the practice fixed and move the country, then hold the country fixed and move the practice, and see which move changes the number more. The eighth cell, the second practice in the Netherlands, has no row, and the source does not say why. The firm target for the year was 70% of track records, proposals and client documents uploaded by the end of 2014.
Exhibit 1 puts the table back into the shape it was always in. The translucent plane is the 70% target, and one of the twenty-one measured values stands above it. The two Norwegian zeros sit on the floor as flat marks rather than as gaps, and the cell with no row at all is drawn hollow, because those are different facts about a unit and a summary line renders them identically.
The country axis repeats
On proposals the three shared countries fall into the same order inside both practices. Denmark first, Sweden second, Norway last. Inside Practice One that runs 68%, 39% and 24%. Inside Practice Two it runs 77%, 67% and 36%. Different absolute levels, identical ranking. Those are the two ridges drawn over the proposal bars in Exhibit 1, one per practice row: the same fall in the same order, set at two different heights, and the vertical distance between the ridges is shorter than the drop along either one of them.
That repetition is the finding. An ordering that survives a change of practice is not an accident of one team's management. Something located in the country is producing it, and the table cannot say what, because it measures outcomes rather than causes. The plan itself notes that take-up and feedback varied by location and that several markets were on hold or struggling to start, so the firm knew the local factor was real without having isolated it.
The client document column makes the point harder. Both Norwegian units recorded 0%. Not a small number, none at all, across 17 closed engagements in one practice and 14 in the other. Two practices, two reporting lines, one country, the same zero.
The country axis also covers the greater distance. Inside a single practice, the gap between its strongest and weakest country on proposals is wider than the gap between the two practices in any country where both appear. Move a unit from one practice to the other and the number shifts. Move it from one country to another and it shifts more.
The practice axis exists, and then reverses
Practice Two led Practice One on proposals in every country where both operated, at 77% against 68%, 67% against 39%, and 36% against 24%. Consistent, and easy to read as a verdict on the two practices.
Then look at shared client documents. In Denmark, Practice One recorded 17% against Practice Two's 14%. In Sweden, 26% against 17%. The ordering has flipped, and whatever advantage Practice Two held on proposals did not carry to delivered client work.
Track records break the ranking altogether. The highest track record capture in the table, 35%, belongs to Practice One in Norway, the same unit with the lowest proposal capture at 24%. The unit with the highest proposal capture anywhere, Practice Two in Denmark at 77%, files track records at 11%. The strongest and weakest units swap places depending on which column you read.
This is why measuring three document classes separately rather than publishing one blended adoption figure was the most valuable decision in the whole table. A single number would have averaged the 24% and the 35% into an unremarkable middle and hidden the fact that one unit was both best and worst in the same period.
The platform explains none of it
Every unit in the table was running the same thing. From November 2013 the programme's first phase had been under way in every location the programme covered but one, a market held back for a combined launch later in 2014, and that phase covered per-engagement sites, search and a document upload process, deliberately scoped as tools only. A second phase, aimed entirely at behaviour, was planned separately for each geography, and the document containing this table is the first draft of that design.
So the technology is a constant across all seven rows, and the outcome ranges from 24% to 77% on the best-covered document class and from 0% to 26% on the worst. A constant cannot explain a variance. The platform bought the possibility of capture and nothing beyond it.
The change plan's own answer names three axes to segment by: practice, location and rank. The table shows the first two. The third is invisible in the data and specified in the design, which writes target behaviours as separate first-person commitments for partners and directors, for managers and bid leads, for consultants and analysts, and for corporate roles.
The plan also leaves a measurement question open. It names an audit tool, counting documents saved per engagement site, as the instrument that measures how the process is performing, and gives mid-August as the earliest availability of the reporting breakdown that tool is meant to produce. The status table is dated 2 July, six weeks earlier, and the source does not say where its numbers came from. Whatever produced them, the plan does not connect them to the measure it names.
The grid is now the agent's map of your firm
Take that table forward twelve years and it stops being a management report and becomes an inventory of what an AI system can know about the firm.
A retrieval agent, or an agent drafting a proposal, answers from what was filed. Nothing else is available to it. Ask such an agent for the firm's delivery precedent in Norway and it has no client deliverables to draw on in either practice, because both recorded zero. It will not fail. It will answer from proposals, captured at 24% and 36%, and give a confident account of what the firm sold in that country rather than what it did. The blind spot has a shape, and the shape is geographic.
That changes what an adoption dashboard has to report. Most AI rollout reporting in 2026 is a blended percentage: seats issued, weekly active users, queries served. It is the exact metric the 2014 table refused to publish. A firm-wide assistant adoption figure cannot tell you whether one country sits high and another low, and the low one is where your agent goes quiet. Report it as a grid, by unit and by artefact class, or you are managing an average no agent ever experiences.
It also changes what improvement means. Push the blended average up through the units that were already good and you have moved the number without touching the hole. The country that filed nothing still files nothing, and every future answer about that market still comes from the sales material. Agents care about variance, not averages.
The interesting part is how much faster this whole problem moves now, and it is not because the software got better at storing files.
Start with the measurement. The 2 July table was a single snapshot of one fixed six-month window, unattributed to any named instrument, and the reporting breakdown the named audit tool was meant to produce was not available until mid-August at the earliest. The same cut now runs off engagement events, split by practice, country, artefact class and rank, recomputed as engagements close, with the residual attributed rather than eyeballed. An autonomous agent owns that cut, watches for a cell going to zero and raises it the week it happens, not six weeks later when the breakdown lands. The reason a snapshot ages is that somebody has to commission it, and nobody has to commission this one.
Then capture itself. All three columns in that table are document classes, and every one arrives as a file rather than a form: a proposal deck, a signed client deliverable, a track record written after the fact. Computer vision reads those pages as pages, so the engagement, the client, the dates, the document class and anything that has to be masked before a colleague can see it come off the artefact itself rather than off a metadata form a consultant closing an engagement was never going to fill in. Scanned signature pages, exhibits and slide graphics are legible to the same model that reads the body text, and the shared client deliverable, the class where both Norwegian units recorded 0%, is the class that arrives most often in exactly that shape. An agent drafts the track record from the engagement's own artefacts and a partner corrects or approves it, which deletes the filing task no incentive was ever going to pay for.
Individually those are features. What makes them quick is wiring them into one pipeline: engagement close triggers extraction, extraction feeds classification and a check against what is already filed, filing updates the grid, and the grid triggers the rank-level prompt that the plan could only write as a first-person commitment for partners, managers and analysts. Every handoff is a machine handoff. The firm-wide campaign becomes a running system, and agents carry a real share of how the firm operates rather than sitting to one side as an assistant somebody remembers to open.
That is loop and harness engineering, and it is the part that sets the timeline. The model is the least of it. The work is the scaffolding around it: the tools each agent may call, the permission on each tool, the escalation path when a document will not parse, and an evaluation set built from documents a human already classified, so the loop scores itself against a known answer and gets better between runs rather than between programmes. Built that way, most of what the 2014 plan asked people to remember to do stops depending on them remembering, and the rank axis it could only address through written behaviour statements and an accountability chain becomes enforceable in the workflow rather than requestable.
One part is not easier. The local factor that produced the same country ordering inside two independent practices is still local, still unexplained by any system, and still the largest term in the equation. Software did not create it in 2014 and will not remove it in 2026. Someone has to go to the country that returned a zero and find out what is true there.
- 24-77%
- Proposal capture across seven practice-and-country units
- 0%
- Shared client documents captured in both Norwegian units
- 8-35%
- Track record capture range, same seven units
- 70%
- Firm capture target set for end of 2014
Move a unit from one practice to another and the number shifts. Move it from one country to another and it shifts more. The platform was identical in every row and explained none of it.
Where to start
Rebuild your adoption reporting as a grid before funding another rollout. One axis for the organisational unit, one for the artefact class the AI will actually read, and a row for every unit including those that returned nothing, because a missing row and a zero row look the same in a summary and mean opposite things.
Then attribute the variance instead of ranking the units. If a pattern repeats across an axis, as the country ordering did in both practices here, you have a cause you can go and investigate. If it reverses, as the practice advantage did between proposals and client documents, the problem belongs to one artefact class and should be fixed there rather than through a firm-wide campaign.
And judge the programme on the worst cell, not the average. An agent built on that corpus inherits the empty cells, and it will speak just as fluently from the places where you have nothing.
“The same platform, the same window, the same upload process, and capture ran from 24% to 77%. Whatever moved those numbers was not in the software, and a single adoption percentage would have hidden all of it.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
