Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

What a Readiness Check Costs

RealAIApr 9, 20248 min read
Delivery AssuranceEnergyProgramme AssuranceRetrieval-Augmented GenerationData Readiness

Assurance reviews get bought the way insurance gets bought. Somebody on a steering committee is uneasy, a number gets agreed, four weeks later a deck lands with findings in it, and nobody asks what the four weeks were spent on. The price is discussed. The instrument is not.

So it is worth taking one apart. I have in front of me a pre-delivery readiness review of a core platform programme at a European energy retailer, which was rebuilding the customer and billing core behind its business supply, building rather than buying, on a platform that already ran elsewhere in the group. The review was run jointly by an outside firm and the group's internal audit function, and it produced findings and recommendations rather than results. What makes it useful now is not the verdict. It is that the review documents its own method well enough to cost it, and once you cost it you can see which lines a machine has already taken.

The register is the honest part of the method

Most reviews describe their document work in one clause. This one printed the list, which is a gift to anyone studying the method.

What is on it is exactly what a programme sheds as it runs. Steering committee packs numbered in sequence with their action and decision lists. Risk logs. Business cases, including the calculation workbook behind one. Sprint plans, governance structures, organisation charts, works council filings, investment approvals, and working documents on billing, metering, credit checking and product range. A commercial fact pack set that predates the programme. Prior review artefacts, one of them a proposal to do the very work this review was doing.

Read it as a sourcing problem and it is full of the same document more than once. A risk log appears at three successive versions. A governance structure appears twice at one version and again at the next. One steering pack appears at two versions dated a day apart, the later one carrying the earlier date. The register is not a library. It is a snapshot of a shared drive, carrying the shared drive's contradictions inside it. That matters more now than it did then.

Two legs, two yields

The detail appendix is where the instrument shows its hand: it lays out the questions asked and the answer each one returned.

Take the panel on development, test and release. Nine questions, walking the build chain from how software gets written to what has to be true before operations will accept it. Six of the nine answers begin with the phrase no actual findings, on my count, off a panel the reviewers themselves stamped work in progress. Four of those six say nothing else at all. One adds that nobody raised the topic in interview. One adds that the plan carried no slack. The three answers with content describe things a reader could have inferred from the sprint plans already in the register.

Now take the panel on the organisation's readiness. Seven questions, all of them about the business on the receiving end: whether it wants the thing, is shaped to take delivery of it, and runs a portfolio the thing can be slotted into. Five carry an explicit risk flag and one of those carries two. Among them: that the programme sat deliberately apart from the business it was rebuilding, with few business people inside it. That the organisation was not on board. That sourcing was not being kept simple, a concern the panel attributes to a steering committee member. That scope was undefined to the point where the question could not honestly be answered. That the same benefits might be counted twice across programmes, which the review flags and then declines to have examined.

Same four weeks, same reviewers, same client. One leg returns a clean bill and the other returns the reasons the programme was in trouble. That asymmetry is what the instrument is for, and it has been hiding inside a cost line nobody itemised.

What the document leg looks like once a machine does it

Sweeping ninety five artefacts, at three versions each in places, is precisely the shape of work retrieval-augmented generation now handles. You index the register into a vector store, retrieve against the reviewer's own question list, and get back per-question summaries with the source document carried alongside every claim. Weeks become hours, and the reviewer's day starts from a drafted evidence pack rather than a folder.

Two conditions decide whether that is worth having. The first is lineage, and this register is a warning about it. If three versions of the risk log are indexed flat, retrieval will blend a risk closed in one version with a risk reopened in the next, and the summary will read fluently and be wrong. Version and date have to be first-class metadata, superseded documents marked as superseded, and every returned claim has to name the version it came from. That is ordinary data readiness work, applied to a document set instead of a warehouse.

The second is an evaluation set. Every question on the panel is a query, and you already have the ground truth: the answers a human reviewer wrote for a comparable engagement. Run the retrieval, score it against those, and you learn where the drafting is trustworthy and where it hallucinates comfort. A copilot that returns a confident no actual findings against an unexamined area is worse than an empty page, because an empty page stops a reviewer and a false clean bill travels.

4 weeks
Stated elapsed time of the full review
19
People interviewed, programme staff and steering committee
95
Document register rows, my count, and a floor not a total
6 of 9
Build-and-test answers opening with no actual findings, my count

The leg that does not compress

Three things in this review could not have come from any document, and all three are why the verdict reads as it does.

The first is that a key person was leaving a stream that had barely started. The review says so in one clause and moves on, because everyone in the room already knew. It is not the kind of thing a risk log carries, because a risk log is a public artefact inside a programme and that is not a thing anyone writes down.

The second is attrition already suffered. Specialists on the chosen platform, a group the review describes as very limited to begin with, had already left, and it names pressure and cultural strain as the cause. A document sweep would have found the staffing plan and the ramp-up chart. It would not have found the leavers, or the pressure and the culture the review names as the reason they went.

The third is a pair of numbers a document sweep would swallow whole. The plan was said to be reliable to about fifty percent, development velocity still unknown under a newly introduced way of working, and a footnote attributes that figure to programme management's own estimate. Beside it, the estimate that the platform would need roughly thirty percent new code is qualified in the review's own words as just an estimate needing supporting analysis. Retrieval reads both as findings. A reviewer in a room reads them as claims and asks who produced them, which is how they came to be footnoted at all.

Every finding that changed the verdict came out of a room. Every finding that came out of the files was already known to the people who filed them.

The dividend is more interviews, not fewer reviewers

The obvious conclusion is that assurance gets cheaper. I think that is the wrong read, and taking it will produce worse reviews.

Nineteen interviews in four weeks, alongside a ninety five item sweep, is a reviewer at the limit of the calendar. It shows: one detail panel in the pack I have is still marked outstanding, two more are marked work in progress, and the assurance plan proposed for the year ahead carries six checkpoints listed as still to be determined. That is not sloppiness. It is what four weeks buys when a large share goes on reading.

Give that reading back and the same four weeks buys more people in the room, including the ones a nineteen person list leaves out: developers rather than their leads, the business people kept at a distance, the specialists thinking about leaving. It buys second conversations, where the first careful answer gets revised. And it buys a reviewer who arrives already knowing what the documents claim, so the hour goes on the gap between the claim and the room.

That is the reallocation we build for on the RealAI Platform, pairing a retrieval layer over the client's own programme record with an interview plan sized against the hours the layer frees. The European rules on AI have reached political agreement and have not yet arrived as obligations, so the assurance question in front of most boards is still an internal one: what evidence does your own review rest on, and which half of it did a person have to be present to collect. Early autonomous experiments belong on the reading half, carefully scoped, with a reviewer signing every claim. Not near the interview.

The instrument does not get replaced. One of its two legs gets cheap, and the other one, which was always the leg that found things, finally gets the time it needed.

Method, counts and quoted phrasing are as recorded in a pre-delivery readiness review of a core platform programme at a European energy retailer, run jointly by an external firm and the group's internal audit function. That review produced findings and recommendations, not delivered outcomes; the register count is mine, and reading the instrument as a two-leg cost structure is ours.

Every finding that changed the verdict came out of a room. Every finding that came out of the files was already known to the people who filed them.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?