Every readiness review I have run or been handed ends with a list. Six items, nine, fourteen. Each one true, each one uncomfortable, and the list arrives carrying an implicit instruction to go and fix them. Boards read the list, programme directors count it. Almost nobody interrogates the thing that decides whether the review was worth commissioning at all, which is how the list was sorted.
I keep returning to one review because of what it did on a single page, which is still the most useful assurance writing I know for what is now asked of AI programmes.
The engagement was an independent readiness check on a European energy retailer that had chosen to rebuild the sales and operations platform behind its business-customer operation in-house, on a bespoke system borrowed from a sister business rather than a commercially available one. Four weeks of work. Nineteen interviews across programme staff and steering committee members, plus a documentation review. It was commissioned as the programme was meant to be entering delivery. It produced findings and recommendations before that delivery happened, not an account of how it went.
What the three groups actually were
Read as filing, three headings over six items is administration. Read properly it is a routing decision, because each group closes only through a different set of people. Take each group by what is in it.
The platform pair was evidential. The reliability of the borrowed base system as the foundation for the new one had not been assessed sufficiently or proven, which the review called a major liability. That system had been selected under time pressure, with a limited view of the delta between what existed and what was needed, and the review says plainly that the choice was not the product of an extensive and fully objective selection process. Neither of those closes in a meeting. Somebody has to go and produce evidence: an independent assessment of the base system, and a designed alternative route to take if the answer comes back badly.
The governance trio was decisional. Scope was still being argued between the business owner, programme management and the steering committee, with the outcome bearing directly on the plan and the requirements. The steering committee had no levers to monitor and control the benefits case, and no clear link between the phases of the programme and the benefits it was supposed to produce. And the two organisations whose people would have to live with the result were being kept at a deliberate distance from the content and the purpose, by choice, with no plan yet for how the gap would be bridged at delivery. None of that is evidence work. All of it is a sponsor and a steering committee choosing.
The remaining item, plan and design, was a design authority problem. The end state and the intermediate states were not clearly scoped or defined, and the development velocity was unknown, which the review says makes the planning indicative at best. Later detail sharpens it: the extension was estimated at around thirty percent new code, and the review records that figure as an estimate needing supporting analysis rather than a finding. The existing base system was running more than fifty issues a month, with the effect of extending it unclear. Programme management's own estimate put the plan's reliability at about half.
Three groups, three rooms, three cadences: an engineering assessment, a steering committee decision, a design authority. Nothing about the severity of the six items tells you that. Only the category does.
Why a category beats a severity score
Most reviews rank. Red, amber, green, and the ranking gets treated as the intelligence in the document. Ranking is useful for one thing, deciding what to worry about first. It is silent on the other question that matters, which is who can act.
Put a scope decision and an unproven platform in the same tracker at the same amber and they get the same attention, which is the wrong attention for at least one of them. The scope decision can be closed by the three parties already arguing it, once someone convenes them. The platform question cannot be closed by anybody inside the programme at all, because the programme is not a disinterested party to it. Both facts are invisible in a severity column and obvious in a category.
That is also why the verbatim carry-through matters more than it looks. When the same words that named a gap on the summary page open the detailed finding pages later, the reader can see the sort is load-bearing rather than decorative. Nothing has been quietly reassigned between the executive page and the analysis. The one-page version and the long version are the same argument at two resolutions, and a steering committee reading only the first is not reading a different document.
The finding that would not sort
Four headline findings, six gaps, three of the findings fed by the page. The fourth stands apart and is the one I think about most.
That finding was about the shape of the programme itself. The set-up was entrepreneurial by design, built for speed around a small group of experts with a high do-it-yourself profile. Teams were assembled largely from individual external contractors working together for the first time. The pool of people who knew the borrowed system was very limited, and some of them had already left. The choice to build in-house rather than bring in a system integrator was, in the review's words, not documented, with no in-depth risk analysis available.
None of that is a gap. A gap implies a hole with a defined shape that closing work can fill. This was the operating model, and the review is careful to say so: it required a high risk appetite and the measures to run and control it, not a fix. Sorting a list is useful partly because of what refuses to sort. The item that will not go into any of your three columns is usually the premise, and the premise is not on the improvement plan.
- Six
- Readiness gaps named on a single page
- Three
- Categories they were sorted into
- Four
- Headline findings in the report
- Twelve
- Recommendations, split into in progress and not yet addressed
The same page, rewritten for agentic programmes
I am seeing this exact list shape again, with the nouns changed.
An AI readiness assessment today comes back with something close to: no evaluation sets worth the name, a retrieval corpus nobody has curated or given an owner, no agreed limits on what an agent loop may do unsupervised, no named owner for model risk, monitoring that stops at uptime, and a business function that has heard about the copilot but has not been asked anything. Six items, every one true, ranked by severity and handed to a programme board.
Sort them instead and the same three groups fall out.
The evidential ones, whether the retrieval layer can actually answer the questions the work requires and whether the model performs the task at the quality the process needs, close only when somebody builds an evaluation set against real cases and runs it. That is engineering work with a deliverable, and like the platform assessment in the energy programme it cannot honestly be marked done by the team whose plan depends on the answer.
The decisional ones, graded autonomy in particular, belong to a sponsor and a risk committee and nowhere else. Which actions an agent may take unsupervised, which need a person in the loop, and who is accountable when a loop acts, are not engineering questions with technically correct answers. They are the same category as scope: a choice, owned by named people, with a date. With the EU AI Act now in force, that conversation also has a counterparty outside the building, which is new, and it means deferring the decision to a later model risk review is no longer a neutral act.
The design ones, what the end state looks like, what the harness around the agent is, what done means for a workflow that used to be a queue, need a design authority holding one picture. In our own work this is the layer we group under Agentic OS, because a programme with no declared intermediate states will measure velocity it cannot interpret, exactly as the energy programme was about to.
A severity score tells you what to worry about. A category tells you who can act. Programmes drown in the first and starve for the second.
The reason to do the sort before costing anything is that the three groups run on different clocks. Evidence takes weeks and money. Decisions take one meeting and a sponsor willing to hold it. Design takes an authority that in most organisations does not currently exist and has to be stood up. Cost a list that has not been sorted and you produce a single number that hides three schedules, and the one that slips is always the one nobody inside the programme could close anyway.
Detail is as recorded in an independent programme readiness check on a European energy retailer, ahead of the delivery phase of an in-house rebuild of the platform behind its business-customer operation: its summary page, its four detailed findings, its recommendation set and the interviews run alongside. It produced findings and recommendations, not delivered results. Every figure quoted is the review's own or the programme's; reading its sort as a routing decision, and mapping it onto agentic programmes, is ours.
“A severity score tells you what to worry about. A category tells you who can act. Programmes drown in the first and starve for the second.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
