Every retrieval programme I have been asked to look at this year arrives with a dashboard already built. Accuracy against an evaluation set. Response latency. Cost per thousand queries, watched closely, because that one lands on a budget somebody defends.
All three numbers are about the model. Not one of them is about the corpus.
I have started asking the same question in those rooms, and I have not yet had an answer. Of the documents you loaded into the vector store, which ones have ever been retrieved? Of the ones retrieved, which were quoted back to a user? Of the ones quoted, which changed what that user then did? The knowledge was the expensive half of the build. Somebody wrote it, somebody reviewed it, somebody approved it. It is the half with no instrument on it.
The plainest statement of that missing instrument I know of was not written for machine learning at all. It sits in an internal digital strategy proposal from a global pharmaceutical company, put to a governance forum, on the page where the group defines its own role in the operating model.
The half everybody builds
The performance half of that charter needs no defending. Dashboards, measurements, indicator tracking, targets, at group level and regional level and market level, pointed at whether the campaigns worked. Every organisation I work with has some version of it, usually several versions, usually arguing with each other.
It is also the half that gets rebuilt first whenever a new capability arrives. Put a retrieval-augmented assistant in front of a body of internal guidance and within a quarter you will have a reporting pack for it. Queries per week. Deflection rate. Tokens consumed. Score on a held-out evaluation set, tracked release over release. That pack answers one question well, and it is whether the machine is behaving. Nothing in it answers the question the second charter asks.
The half that is written down and then not built
Strategic conformance and course correction is an odd phrase to meet in a marketing governance model, and it is doing precise work. Conformance asks whether the strategy that was published is the strategy the organisation is actually running. Course correction asks what you change when the answer is no. Neither is a performance question. Both are questions about whether an instruction, once issued, made it into practice.
The things named underneath it tell you how the group intended to answer. Content utilisation first. Then market case studies, service enhancement pointers, and increased use of published assets. Read together, those are not measures of output. They are measures of uptake, and they are pointed at the organisation's own published material rather than at its customers.
That distinction is the whole finding, so it is worth stating flatly. A programme that produces guidance and measures how much guidance it produced is measuring the wrong side of the transaction. Production volume rises fastest exactly when the material is not being used, because unused guidance generates requests for more guidance.
Honesty about what this document is: the charter states an intent, and the deck records no measurement against it. No utilisation baseline, no count of markets that had picked up a given asset, no figure for how often a published guide was opened. It is a proposal to a governance forum, and like most governance pages it specifies the mandate and leaves the instrument to be built later. What makes it worth carrying is that it named the mandate at all. Most of the equivalent pages I read do not.
Why this matters more now than it did then
When published guidance was read by people, non-use was at least visible in the ordinary way. Somebody would notice that a market was doing its own thing. A regional lead would mention that nobody had opened the deck. The signal was weak and late and social, but it existed.
Retrieval systems remove even that. The corpus is consumed by a machine on behalf of a user who never sees the shelf it came from. If a document is never selected, nobody experiences its absence, because the assistant answers anyway. It answers from whatever else it found, or from what the base model already carried, and that answer looks exactly like the one the approved material would have produced. Non-use becomes invisible by construction.
That is a governance problem before it is a technical one. The organisation believes it has published a policy into the system. What it has actually done is add a file to a store that may or may not surface it, and the gap between those two statements is where copilot programmes quietly stop being governed. No number in the standard reporting pack detects it.
- Two mandates
- One analytics capability, chartered for performance and for conformance
- Content utilisation
- The first thing named under the conformance half
- Five verbs
- Source, reuse, manage, monitor and maintain, the content charter
- No figure
- Utilisation measurements recorded anywhere in the document
What the second instrument actually measures
The measures follow directly once you accept that a corpus is a product with users rather than a pile of files.
Retrieval coverage. What share of the corpus has been selected at least once in the last period. The shape is always the same: a small head of documents doing nearly all the work, a long body never surfacing at all. That body is not neutral. It is material somebody was paid to write, review and approve, sitting in the store misleading everyone about what the system knows.
Citation, separately from retrieval. A chunk can be pulled into context and contribute nothing to the answer. Retrieval counts are the easier number and the less interesting one.
Consequence. Whether the answer changed a decision, a draft or a next step. In most organisations this can only be sampled rather than logged, which is fine. A sample is a measurement. An assumption is not.
Contradiction. Which documents get retrieved together and disagree. Two approved guides that conflict were always a problem; a retrieval system makes it an operational one, because whichever chunk ranks higher today becomes the policy today.
Conformance, at the level the source document meant it. Take the guidance the organisation has actually published and ask whether the assistant's answers are consistent with it. That is an evaluation set built from your own published position rather than from generic benchmarks, and it is the only version of accuracy that means much to a regulated business. With political agreement on the EU's AI Act now reached, the question of what an answer was based on is heading toward being asked formally rather than only internally, and lineage from answer back to source document is what will be asked for.
A corpus nobody retrieves is not a knowledge base. It is a filing cabinet with an API in front of it.
The verb that carries it
The most transferable line in the source is not on the analytics row at all. It is in the ownership statement for content, which charters that function to source, reuse, manage, monitor and maintain. Five verbs, and only one of them is about making anything new.
Compare that with how a knowledge base for an assistant typically gets funded. A project to write it, a project to clean it, a project to load it, and then nothing. No owner, no review cadence, no monitor verb anywhere in the plan. The corpus becomes a one-time capital event, which is the surest way to have it decay into a source of confidently delivered stale answers.
The same document shows what governed content looks like on the publishing side. A backlog of a dozen or more guidance documents, sorted into status bands so that half-finished material is visibly half-finished. Dates of the kind technology delivery gets: a usable draft estimated within thirty days, final publication committed within sixty. And open questions put to the partner functions about what other guidance they would need and what format would work best, which is the only demand-side question on the page.
Every one of those moves transfers to a retrieval corpus without modification. Status on each document. A named owner. A publication date and a review date. A standing question to the people whose work it is supposed to change.
What I would do on Monday
Instrument retrieval before you tune anything. Log the document identifier for every chunk that enters context and every chunk that reaches the answer, keep it joined to the query, and hold it long enough to see a pattern. This is MLOps plumbing of the least glamorous kind, it costs a fraction of one accuracy sprint, and it is the first thing a RealAI Platform engagement puts in, ahead of any work on the model itself.
Then publish two numbers next to the accuracy number, monthly, to the same audience: share of corpus retrieved at least once, and share of answers that cited approved material. Both will be uncomfortable at first, and both tell you whether the knowledge investment was real.
Then give the corpus the five verbs. One owner, a monitor cadence, a review date on every document, and a rule that anything unretrieved for two review cycles gets rewritten, retired or explained.
None of that requires a better model, which is the awkward part, because all of it has to be true before a better model is worth buying.
Details are as set out in an internal digital strategy and roadmap proposal at a global pharmaceutical company, presented to a governance forum: its governance operating model page, its guidance publication backlog and its publication commitments. That deck is a plan put forward for approval, not a record of results, and it records the analytics mandate without recording any measurement against it. Reading that mandate as the missing instrument in retrieval programmes is ours.
“A corpus nobody retrieves is not a knowledge base. It is a filing cabinet with an API in front of it.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
