In July 2014 a global management consultancy running an internal knowledge and information management programme circulated a change plan for its own firm. The cover marked it Version 0.1, for discussion within the team. It is not a results document and never claimed to be. It is something rarer: an honest attempt to write down every part of the operating model that had to move before a piece of software could produce the value someone had already promised.
The software was the part that had already shipped. From November 2013 the firm had been rolling out per-engagement job sites, search, and a document upload process to every location except its home market, with take-up and feedback described as varied and several markets on hold. That first phase was deliberately scoped as tools only. A second phase, aimed entirely at behaviour, was to be planned separately for each geography, and this document is the first draft of it. Reading that sequencing in 2026 is uncomfortable, because it is exactly the sequencing most AI programmes are running right now.
What the tools bought on their own
A status table dated 2 July 2014 covered engagements closed between 1 January and 25 June that year, across seven practice-and-country units. It measured three document classes separately rather than reporting one blended adoption figure, and that choice is what makes it worth reading.
Proposal capture ran from 24% to 77%. Shared client documents ran from 0% to 26%. Track records ran from 8% to 35%. Unit workload over the same window varied from 14 closed engagements to 73. Only one unit passed the firm's 70% target, and only on proposals.
Two units recorded 0% on shared client documents. Not a low number. None. Across 17 closed engagements in one and 14 in the other, no delivered client work reached a place where a colleague could find it. They were two different practices sitting in the same country, on the same platform, in the same window, with the same search and upload process.
One of those two is the most instructive row in the table. The same unit that filed nothing at all on delivered client work posted the lowest proposal capture of the seven, 24%, and the highest track-record capture of the seven, 35%. Rank it on one document class and it is the worst unit in the firm. Rank it on another and it is the best. A single blended adoption score would have averaged those two extremes into one unremarkable figure and described neither.
That is the finding worth carrying forward. The technology was constant, and the outcome ranged from 24% to 77% on the best-performing document class and from nothing to 26% on the worst. Whatever was moving those numbers was not in the software, and it was not a general willingness to file either.
There is a second, quieter problem in the same table. Elsewhere the plan names the instrument for this: a job audit tool that counts documents saved per job site, with earliest availability given as mid-August 2014. The status table is dated 2 July. Capture percentages were being circulated roughly six weeks before the measure the plan had specified for them was due to exist.
The table is also a partial view. Ten locations are listed as part of the first phase; the table carries seven units drawn from four countries. Three of the missing markets were on hold and one was described as struggling to kick off, and not one of them has a row. The locations the plan itself flags as stalled are precisely the locations the status report does not cover.
The seven things
Against that backdrop the plan did something most technology programmes skip. It mapped knowledge sharing onto a cultural web and wrote a destination state for each of seven elements, under a vision the firm expressed as treating its knowledge the way it treats its cash.
How people are managed. Induction, development, performance management, reward, recognition and career progression, all restated in terms of knowledge sharing. The design specified that every appraisal and every mid-year and year-end submission review knowledge-sharing performance explicitly, and that promotion criteria carry a minimum standard.
Structures and capabilities. Roles defined and described rather than assumed. Practice heads accountable for capture across their practice, supported by knowledge officers inside the practice for compliance and quality. Engagement leads accountable per engagement, delivery managers responsible for the working method, consultants responsible for contributing and for flagging where the behaviour was weak.
Control processes and governance. An executive owner of the process accountable to the chief executive, a business sponsor drawn from the partnership to represent the operating side, and a prioritisation forum accountable for evolving the process against business need and gathering user feedback. Lines of authority, decision rights, and the degree of freedom people had were all treated as design variables rather than as background.
Signs and symbols. Saying "we" rather than "I" about the firm's work. Sending links instead of attachments. Every engagement team using its site as the place the current version lives. Small, visible, cheap to observe, and each one a proxy for whether the change was real.
Routines and behaviours. Written out longhand and then per role. The one worth stealing: on any new lead or proposal, the first activity is to search the repository. Not a suggested step, the first one. Senior staff were additionally expected to review the knowledge-sharing performance of their teams at least monthly.
Communication and information. Reporting on site usage delivered directly to the people accountable, success stories carried on the internal channel, and a tone the plan fixed deliberately: positive about the commercial opportunity, and unambiguous that not sharing knowledge was unacceptable. Good and bad behaviours were both to be stated, so that expectations were legible.
Leadership. Eleven leadership populations named separately, from the management committee down to informal influencers inside practices, each carrying its own line of expectations. All leaders were to store their own documents on the sites, encourage their teams to use the job search, and be conscious that what they say to a new joiner is part of the operating model. Account leaders got the sharpest instruction in the set: build searches on the account site rather than creating folders. That is a filing philosophy in one clause, and it is the same argument anyone standing up a retrieval index has to win today.
The one they could not close
Read the seven together and one is visibly harder than the rest. Under reward, the plan states plainly that cash is king, because utilisation is the predominantly weighted element of the scorecard, and then asks what minimum percentage knowledge sharing should carry. The number is not filled in. It also floats a link between partner bonus retention and completion rates on their engagements, again without a figure.
This is the most honest slide in the document. Six of the seven elements can be designed by the programme team. The seventh requires someone senior to move a percentage from one line of the scorecard to another, and that had not happened. Everything upstream, the champions network, the training updates, the appraisal wording, sits on an incentive that still paid people for billable hours and not for what they left behind.
The plan also separated two kinds of measurement on purpose: an audit count of documents captured per site, and a staff survey scheduled for November 2014 to read perceived improvement. Counting artefacts and asking people whether anything got better are different questions, and running both is the right instinct.
The same seven, in 2026
Change the nouns and this is a current document.
The tool phase now is an assistant licence, a retrieval index over the document estate, and now an agent with permission to act. It ships in weeks. It is easy to fund, easy to demonstrate, and it generates exactly the same class of status report: seats issued, queries served, an adoption percentage that describes logins rather than value. The behaviour phase, as in 2014, is planned for later and tends to stay there.
What is different in 2026 is that the corpus is no longer just a place people look. It is the ground truth an agent reasons from, and its shape propagates. A retrieval system built over the estate in that 2014 table would have been fluent about proposals and close to silent about delivered work, because that is what people filed. It would not have failed loudly. It would have answered every question confidently from the sales material and almost never from what the firm actually did, and no one would have seen the gap unless they were already measuring capture by document class.
So the first move is the diagnostic that table represents. Before switching on retrieval, measure completeness by document type, not by volume, and give the units that returned no data an empty row rather than no row at all. Volume hides the bias. A corpus that is well covered on one class and close to empty on another is not partly complete, it is skewed, and skew is what an agent inherits.
The second move is that the seven elements now have a mechanical edge they lacked in 2014. Search-first was written as a routine: on any new lead or proposal, search the repository before anything else. A routine has to be remembered by the person doing it. The same rule built as a harness sits inside the proposal workflow, where an agent runs the search the moment a lead is created and hands back what the firm already has, and nobody is relying on discipline. Capture moves the same way. Rather than a consultant uploading finished material to a job site after an engagement closes, an agent drafts the track record from the engagement's own working material as the work ends, and a human corrects and approves it. That takes out the filing task the incentive was never going to pay for.
The class that ran from 0% to 26% is also the least uniform, and that is the likeliest reason it ran last. A proposal and a track record are firm-shaped artefacts with a template and an owner behind them. Delivered client work is whatever a given engagement happened to produce, in whatever form the client accepted it. The 2014 process handled that variety by asking a person to open each document, judge what it was, and decide where it belonged. Document AI takes that judgement off the person now, reading page structure and labels well enough to classify and route a deliverable without someone deciding first. Which matters here because that judgement step is precisely the work the scorecard priced at zero.
Chain those steps together and the seven elements stop being a change programme and start being wiring. The job audit tool that was not due until mid-August becomes a query that runs at every engagement close, by document class, by unit, with the units that returned nothing appearing as an empty row rather than as no row. Reporting to the people accountable runs continuously instead of arriving as a table someone assembles by hand before the measure behind it exists. The monthly review of team knowledge-sharing performance that senior staff were asked to perform arrives already written, with the gaps named. Each of those is a loop: an agent acts, the result is measured, the measurement changes the next action, and the loop runs whether or not anyone remembers it. That is what autonomous means in practice here, and it is where agents start carrying a real share of how the firm runs rather than sitting beside it as an assistant someone chooses to open.
Which makes the honest reading of this plan in 2026 not that it was wrong but that it was slow. It set a capture target for the end of the year and put the behaviour phase into a second document, to be planned separately for each geography. Built as loops with a harness around each one, instead of as instructions people are asked to follow and managers are asked to audit, six of those seven changes land in weeks rather than in the next planning cycle. The seventh does not compress at all.
The third move is the one automation does not touch. Agents raise the cost of the unresolved scorecard question rather than lowering it, because an agent acting on a thin corpus does damage faster than a colleague reading a thin intranet. If contribution to the shared record is still unrewarded, the record stays thin, and the assistant built on it will be confidently unrepresentative of your own work. That is not a model problem and no procurement decision fixes it.
The 2014 plan deserves credit for knowing this. It shipped the tool, went back and counted what the tool alone had bought, and then wrote down the seven other things that had to move. More than a decade on, most AI programmes manage the first step, mistake a seat count for the second, and never write the list at all.
- 70%
- Capture target set for end of 2014
- 0-26%
- Range on shared client documents, seven units
- 24-77%
- Range on proposals, same units, same period
- 7
- Cultural elements the plan required to change
Six of the seven elements a programme team can design. The seventh needs someone senior to move a percentage on the scorecard, and that is the one that decides whether the other six hold.
“The tool had been live in almost every office since November 2013. By the following summer, capture of delivered client work still ranged from zero to twenty-six percent against a seventy percent target. Nothing was broken. Nothing was being filed either.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
