Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

Report Early, Then Review Yourself Against It

RealAINov 28, 20238 min read
Delivery AssuranceProgram GovernanceEnergyMLOpsData Readiness

The default shape of an assurance review is one arc. Somebody commissions it, a team goes in for a few weeks, and a report lands at the end. Between the last interview and the steering meeting where the findings are read aloud sits an interval in which the reviewers know things the organisation does not. Nothing moves in that interval. It is dead time by design, and everybody has agreed to it so completely that no one prices it.

Then the report arrives complete, and the meeting that receives it accepts the findings, assigns owners and schedules a follow-up. The response begins after the reviewers have gone, so the one question worth asking about a review, whether anything moved because of it, is answered by nobody who was there.

I keep returning to one document that refused that shape. It is an independent readiness check run on a European energy retailer's business and SME division, which was part way through rebuilding the sales, contracting and billing platform behind its business-customer operation on a bespoke system borrowed from an affiliated business. Four weeks of fieldwork, jointly staffed by the parent group's internal audit function and an external assurance firm. Nineteen interviews spanning programme staff and steering committee members, plus a pass over the programme's own documentation. A deliberately partial scope: several assessment areas were taken now and the rest deferred to a later review. It produced findings and recommendations about a programme still in front of everyone, not an account of how it went.

What makes it worth reading is structural rather than substantive. It reported twice.

The two passes

The mechanism is plain enough to walk past. Roughly two weeks in, with the interviews mostly done and the documentation reviewed, the team wrote down what it had and handed it to the people running the programme. Not a corridor warning, not a heads-up to the sponsor. A document, in draft, with findings in it.

The programme did what programmes do when handed a credible list early enough to act on: it acted on some of it. The report records that management acknowledged the findings and responded, and that the response produced measures for the programme overall. Then the second half of the review ran with that response as part of its subject matter.

The closing report then states its own purpose in one line, and the line is unusual. It is the result of validating the first findings against the progress the programme had made since. Read that as a definition of the deliverable. The report is not a description of a programme. It is a description of a programme plus its reaction to being described.

The reviewers also proposed that the same team monitor and report on how the recommendations were followed up, which is the same idea extended past the end of the engagement. Grading the response was not an afterthought. It was the design.

What the second pass could see

The detailed finding tables carry three columns: the question asked, the finding, and the movement between the two passes. That third column is the review looking at itself.

In the organisation table, one row records that the programme's deliverable could not be judged justified because the scope had not been defined. Beside it, the movement column records a session with the steering group to define that scope, attributed in the table to the outcome of the first review. The finding was not resolved. It was answered with a meeting, and the meeting is on the record with its cause attached.

In the delivery table, a row on standard development approach records that agile working was being introduced, coaches had recently started and teams had recently scaled. The movement column beside it says the coaches were up and running. That is a small thing to write down and a hard thing to reconstruct later. Anyone reading the closing report alone would see a team with coaching in place and no way to know whether that was true when the reviewers arrived or became true because they did.

Be honest about what that column is: it records movement, not outcome. A scope session held is not a scope agreed, and coaches working is not delivery velocity. Its value is that it separates two things a single report fuses, the state of the programme and the responsiveness of the people running it.

The split that early reporting produces

Twelve recommendations went out against four headline findings, distributed three, two, five and two. Three carried the mark for already addressed. Nine carried the mark for not addressed yet.

It would be comfortable to read three out of twelve as a weak result. I read the composition instead. The three that moved were the ones the programme could move on its own: approving scope, goals, timelines, stages and budgets in its final plan of approach; developing an architecture and business design with a roadmap attached; building a staged plan on top of a roadmap picture that had recently been delivered. All three are documents, all three were already in motion, and all three sit inside the programme's own authority.

The nine that did not move mostly needed something the programme could not supply by itself. An independent assessment of whether the borrowed platform could carry the load, which required commissioning someone. A fallback route if it could not, which required money and a decision nobody wanted to take. A stakeholder plan reaching into two other organisations that had deliberately been kept at arm's length. A lagging workstream brought up to speed, which required a team and a plan of approach it did not yet have. Levers by which the steering committee could actually monitor and control the benefits case.

That sort is the real product of reporting early. A single report gives you twelve items ranked by the reviewer's sense of importance. Two passes give you twelve items ranked by observed tractability, which is a ranking no one can produce from an interview. The programme graded its own recommendations by acting on some of them, and the reviewers only had to write down which.

4 weeks
Total review, reporting twice rather than once
19
Interviews across programme staff and steering committee
3 of 12
Recommendations already addressed when the report closed
9 of 12
Not addressed yet when the closing report went in

Why this belongs in machine learning delivery

The single-arc review is the standard shape in our work too, and it fails the same way.

A data readiness assessment runs for six weeks and reports on an estate that has since had two pipelines rewritten. A model evaluation lands after the deployment window it was meant to inform. Lineage work is scoped against a feature store that changes underneath it. Process mining pictures a process the owner has already begun rearranging, partly because they knew someone was looking. Each report describes a photograph, and the organisation is asked to act on it weeks later, when the thing photographed has moved.

The two-pass structure fits this work better than it fits programme assurance, because machine learning delivery moves faster than platform delivery and the interval costs more. Write the first findings at the halfway point, hand them over in draft, and let the team respond while you are still measuring. Then spend the second half asking a different question. Not what is wrong with this estate, but what this organisation did in two weeks when it was told what was wrong with it.

The answer to the second question is what I would want before signing off any model into production. A team that closed three definitional gaps in a fortnight will close the rest. A team that closed none, having been told plainly and early, is telling you something about its capacity that no maturity discussion will surface. The early language model pilots crossing my desk this year are almost all being run by teams whose real constraint is response speed rather than technical skill, and response speed is invisible to a review that reports once.

A single report tells you what was true when the reviewers stopped looking. A second pass tells you what the organisation did when it found out, and only the second of those predicts anything.

What the reviewers proposed next

The closing recommendation set includes an operating model for assurance rather than another review. A short quality report into every steering committee meeting, in the same rhythm as the programme's own reporting, carrying urgent findings, the status of follow-ups, and new risks. Alongside it, a small number of deeper assessments at defined moments in the delivery calendar, each producing its own findings for the committee to decide on.

That is a proposal, not a delivered outcome, and the report marks it as such. Whether the division adopted it is not in the document and I will not invent it. What is in the document is the argument for it, which the review had already run on itself over four weeks: assurance that reports continuously turns findings into a clock, and a clock is what makes people move before the deadline.

Our Consult team now writes the interim report into the scope of every assessment we run, with a date on it, before the first interview is booked. It is unpopular for about a week, because handing over an incomplete finding feels like handing over a weakness and the draft is always argued with. Then the arguing turns into work, and by the closing report there is something to measure that was not there at the start.

Findings, counts and the two-pass structure are as recorded in an independent programme readiness check on a European energy retailer's business and SME division, jointly staffed by the parent group's internal audit function and an external assurance firm. That review produced findings and recommendations on a programme still in delivery, not results. The reading of its two-pass design as an operating model for machine learning assurance is ours.

A single report tells you what was true when the reviewers stopped looking. A second pass tells you what the organisation did when it found out, and only the second of those predicts anything.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?