Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

The Plan Is the Artefact Under Pressure

RealAIAug 1, 20248 min read
Delivery AssuranceOperating ModelAI GovernanceData ReadinessMLOps

Every delivery review I have worked on has found problems. That is not a compliment to the reviews, it is a description of deliveries. Point a careful reviewer at any programme of size and they come back with a list, and the list is never empty, and the length of it tells you more about how long they were given than about how the work is going.

Which is worth holding on to when you design anything that watches a delivery, because a watcher built to answer whether something is wrong is answering a question whose answer is already known.

I found the better version of that question in an old document, written for people rather than software.

A compound condition, read slowly

Some years back I led the technical workstream of an independent post-mortem review of a failed offshore application build, commissioned by a European banking group's IT services subsidiary. The build had lost its quality and then its schedule, there had been a replan partway through, and the work had been taken back in house before the review opened.

Among the recommendations was a template of status definitions, proposed so that both parties on a future assignment could agree in advance what the words meant. A sibling piece takes up a different part of that template. This one is narrower still, and sits inside a single line of the entry conditions for its most serious state.

That state can be entered by more than one route, and the first of them is the line worth reading slowly, because it has two halves joined by an and. One half is serious concern about whether the work can deliver in its current shape. The other is that no remedial plan has been agreed. Not one or the other. Both.

Read it as an operations document and it is a rule about when to sound an alarm. Read it as an instrument, and it is doing something more interesting: it is discriminating between two projects with the same problems. One of them has an agreed answer and one does not, and only the second satisfies the condition.

What that condition is actually measuring

Take that seriously and a single line of a status definition stops being a severity threshold.

Read against it, a delivery can be late, over its estimate and carrying quality it should not be carrying without meeting that condition at all, provided somebody has produced a route back and somebody with standing has accepted it. What the line turns on is not the difficulty. It is whether an agreed answer to the difficulty exists.

That is a more useful thing for a line of a definition to do than grade severity, because severity is the part everyone is already arguing about. In a troubled delivery the argument is about defect rates, inflated estimates, slipped milestones. Meanwhile the object that decides how this line reads is not the problem at all. It is the response to it, which is usually the artefact getting the least attention in the room, written last, by the person with the least time.

Problems are the constant in any delivery of size. The variable is whether anyone has agreed what to do about them, which means the object under pressure is never the problem. It is the plan.

Why this is the harder thing to detect

Now put a piece of software next to that delivery record and ask it to help.

Finding problems in a project archive is close to free. The archive is made of them. Risk logs exist to hold them, minutes record the ones raised out loud, test reports count them, escalation mail is nothing else. Point a retrieval-augmented assistant at that corpus with a question about what is going wrong and it will return a great deal of material, all of it genuinely present in the documents, none of it telling you anything you did not know on the way in.

Finding an agreed remedial plan is a different task, and harder in three specific ways.

It is sparse. A delivery record holds far more description of difficulty than it holds of considered response to it.

It is relational. A plan only counts as an answer to a particular concern, so what has to be recovered is a link between two documents that were probably written weeks apart by different people, in different systems, using different words for the same thing. A retrieval layer that ranks passages by similarity to a query will happily return a good plan and a serious concern that have nothing to do with each other.

And it is conditional on acceptance. The agreement is the load-bearing word in that definition, and agreement almost never lives in the plan document. It lives in a minute, a mail, an approval in a workflow tool, or a name in a slide footer. A plan nobody accepted is a draft, and by the plain reading of the criterion it does not count.

So the useful output of a watcher over a delivery is not a list of problems. It is a much shorter list: concerns raised, with no accepted response attached to them, and a person's name beside each so somebody can say whether the gap is real.

Existence is machine-checkable. Credibility is not.

Here I have to be careful about what the source will support, because the same review that produced the template also shows the limit of the idea.

The commissioning side's project documentation was, by the review's own reading, hard to fault. Every expected document was produced. Meetings were minuted, mail was logged, the statement of work was detailed and the non-functional requirements were clear. Concerns were escalated through the agreed route when they arose. And there was a replan. Documents existed, responses existed, and the delivery still ended with the work taken back in house.

A record can be immaculate and still fail the test, because the criterion asks for a plan that has been agreed, and behind that word sits a judgement nobody has automated: whether the plan would have worked. Credibility is not a property of text. Two documents can be equally well formed, equally signed off, and one of them changes what happens next while the other restates the schedule with more confident adjectives.

That is the boundary I would draw for anyone building this. Software can establish, reliably and at a scale no reviewer manages, whether a response exists, whether it is linked to a specific concern, whether it has an owner and a date, and whether anyone with standing accepted it. Those are structural facts and they are checkable. Whether the response is any good is a judgement, and it belongs to the people who will have to live with it.

Which is a comfortable division of labour, and an unusually clean one. The tedious half of delivery assurance is the half that fails most often: nobody has the hours to walk the whole archive and check that every raised concern has an accepted answer tied to it, so it does not get walked, and the gaps that matter sit in the record undisturbed. Hand that walk to software and you have not automated judgement. You have delivered the judgement calls to the people paid to make them, sorted, with the evidence attached.

What to build first

Three things, in this order, none needing a model to start.

One, decide what an accepted response looks like in your own estate, written down as fields rather than a description: what it must name, who must have accepted it, where that acceptance is recorded. Until that exists there is nothing to retrieve and nothing to evaluate against.

Two, build the evaluation set out of the right pairs. Not passages labelled as concerning or reassuring, which is the easy dataset and the useless one, but concerns paired with the responses that were genuinely accepted for them, and concerns marked as having none. That set is expensive to assemble, and it is the whole asset. Anything that reads a delivery record can be judged against it.

Three, keep the lineage. Every link the system proposes between a concern and a response should carry where both came from, so a person can open the two documents and disagree. This is where a RealAI Platform engagement starts on work like this, and it is not a compliance decoration. A link you cannot open is an assertion, and assertions are what the whole exercise was meant to replace.

The template I keep coming back to did none of this. It was a template in a recommendations pack, proposed for assignments that had not begun, and I have no results from it. What it had was a correctly framed question, which is rarer than a working system and harder to arrive at.

Drawn from an independent post-mortem review of a failed offshore application build, commissioned by a European banking group's IT services subsidiary, on which I led the technical workstream. The status definitions described here were proposed for future engagements: the review produced findings and recommendations, not delivered results. Reading one criterion's compound structure as a design brief for automated delivery assurance is ours.

Problems are the constant in any delivery of size. The variable is whether anyone has agreed what to do about them, which means the object under pressure is never the problem. It is the plan.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?