Most organisations I work with now buy assurance before they commit. An evaluation run before a copilot reaches the frontline. A review of an agent harness before it is granted write access to anything that matters. A code review of the platform a programme is about to build itself on. This is progress, and I would rather argue with a team that commissions reviews than one that does not.
The failure I keep meeting is not that these reviews are done badly. It is that they are done well, against a question that was settled before the reviewer arrived. The scope sits in a short paragraph near the front of the report, or in a footnote nobody reads aloud. At the moment of commissioning it sounds like precision. Read six months later, it is the whole story.
The cleanest example I know is not from an AI programme at all. It sits in an independent readiness check on a European energy retailer that had decided to rebuild the sales, contracting and billing platform behind its business-customer operation. Rather than buy a package, it chose to extend a bespoke system already running inside an affiliated business. The decision document came first. Some months later an outside software firm delivered a review of the code of the system being borrowed. The readiness check cites that review by name and date, and then attaches two notes to it that carry more weight than the citation.
Two notes in the margin
Strip the names out and the notes read like this:
The outside firm has only focussed on the scalability of the existing code, not taking the new programme's scope into account.
The outside firm has only focussed on the existing code, not on all the prerequisites of the extended system.
Nothing in either note is a criticism of the reviewer. A code review that examines a codebase as it stands and reports on how it scales and how it can be maintained is a legitimate piece of work. The readiness check does not question the quality of this one. It questions its reach. The review answered the question it was given, and the question it was given was about the code that already existed.
The question the organisation actually needed answering was different in kind. Will this codebase carry a system roughly a third larger, serving customer types it has never served, wired to systems it has never touched, at data volumes nobody has modelled, built by teams whose delivery rate is not yet measurable? That question was never put to the reviewer. It could not have been answered by reviewing the existing code, because the thing at risk did not exist yet.
So the file closed with a delivered review in it. An assurance exercise produces confidence as a by-product of having happened, whether or not it covered the risk, and the scope line at the front of the report does nothing to slow that down. That is the part that should worry people.
The decision the reviewer never gets to make
The reviewer answers the question in the terms of reference, thoroughly, on time, within budget. By the time the work starts, the useful range of the answer has already been fixed by whoever wrote those terms, usually in an afternoon, usually under time pressure, usually by the party with the most to lose from a wide scope.
The readiness check is explicit that the platform choice was made under severe time pressure, with a limited view of the difference between what the borrowed system did and what the new one would have to do, and that the selection was not the product of a full and objective process. It records that challenging questions were asked before the decision and were not fully answered when the decision was taken. That is the environment in which review scopes get written. The scope of the code review was not an act of evasion. It was the natural output of a hurried decision that needed reassurance quickly, and quick reassurance is easiest to obtain about the thing that already exists.
This is why I treat scoping as the review rather than as preparation for it. Everything after the scope is execution, and execution is the part that is easy to buy.
The same shape, in every AI programme I see
Change the nouns and this is the most common assurance failure in agentic work.
An evaluation set is built from traffic the assistant has already handled. It is a real evaluation set, carefully labelled, and it certifies performance against the distribution of yesterday. Then the retrieval corpus gains three new document sources, the assistant is pointed at a second business line, and the certified number travels along with it as if it still meant something. Nobody lied. The set answered its question, and its question was about the old corpus.
Or the harness is reviewed while the agent is read-only, and reviewed well: prompt injection paths, tool permissions, logging, the lot. Then autonomy is graded upward, write access is granted, and a loop that was safe to observe becomes a loop that can act. The review said the system was sound. It said so about a system with different powers.
Or model risk management documentation is prepared and signed off describing the system as at sign-off, while the interesting risk arrives through the interfaces added afterwards. This is the same failure as the thirty percent of new code and the unexamined connections to sourcing and the general ledger. The reviewed component stayed sound. The composite that got built was never reviewed at all, because at review time it was not there.
The EU AI Act, now in force, makes this sharper rather than softer. Obligations attach to a system's intended purpose, and in an agentic programme the intended purpose is the thing most likely to move between the assessment and the deployment. Conformity work scoped to the system as documented, in a programme where the documentation trails the build, produces a folder that is accurate and beside the point.
- ~30%
- Estimated share of new code in the extension, flagged as an estimate wanting analysis
- 50+
- Issues a month on the existing platform before extension
- None
- Detailed fit-gap analysis on record
- Not determined
- Performance impact of the new data volumes
Four questions I now put in every terms of reference
None of these requires a specialist to answer. All of them have to be settled before a specialist is worth hiring.
What will this thing be carrying at the moment it matters, expressed in load, data populations and connections rather than in adjectives. Review that, not the demonstrator.
What is the delta, named in its own units. A share of new code. A count of new interfaces. A customer population the system has never seen. The delta is where the risk lives, and it is precisely what a review of the current artefact cannot reach.
What would count as a failure, written down before the work starts. A review with no defined failure condition returns observations, and observations are absorbed rather than acted on.
What is out of scope, and who is picking it up. This is the discipline the readiness check modelled well. It did not pretend the code review had covered the programme scope. It said plainly what had been covered and what had not, and left the gap visible on the page.
A reviewer answers the question in the terms of reference, thoroughly, on time. Whoever wrote those terms had already decided how much the answer would be worth. Scoping the review is the review.
What happened to the gap
The readiness check recommended an independent assessment of whether the chosen platform was a suitable foundation for what was going to be built on it, and whether it would still be one later. That is the question the code review had not been asked. When the report closed, the assessment was a recommendation with a date pencilled next to it rather than a piece of work anyone had done.
I am careful about what that does and does not tell us. This was a review, so what it produced was findings and recommendations ahead of a delivery that had not happened. It is not a record of an outcome, and I am not claiming one. What it is, is an unusually clear photograph of a moment every programme passes through: the moment when a competent piece of assurance is sitting in the file, the room feels covered, and the question that would have changed the decision has not been written down anywhere.
The practical work is unglamorous and it happens before anyone is engaged. Say what you are building toward. Say how far it is from what exists. Then scope the review to the distance between them, and treat any review of the current artefact as evidence about the starting point rather than about the journey. Our Consult team spends more time on that paragraph than on the assessment that follows it, and the same habit is wired into how Agentic OS deployments are gated: each increase in autonomy re-opens the evaluation rather than inheriting the last one.
Detail as recorded in an independent readiness check on a European energy retailer's platform rebuild: its findings on the technical basis of the programme, and its own notes on what an earlier third-party code review did and did not cover. That review produced findings and recommendations ahead of delivery, not results. The reading of it as a scoping failure, and the application to agentic assurance, are ours.
“A reviewer answers the question in the terms of reference, thoroughly, on time. Whoever wrote those terms had already decided how much the answer would be worth. Scoping the review is the review.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
