Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

Was It Really Red

RealAIApr 9, 20248 min read
Delivery AssuranceVendor ManagementFinancial ServicesData ReadinessMLOps

A delivery review is normally commissioned to answer a question whose shape somebody has already decided. Why did this cost more than the plan said. Why did the date move.

The most interesting page in one review I worked on carried a question nobody had commissioned. It sat in an appendix, in the reviewer's own shorthand, addressed to the reviewer rather than the client. Was the project really red, or did it just get rated red?

The engagement was an independent post-mortem of an outsourced application build at a European banking group's IT services subsidiary, already taken back in house by the time we looked at it. We examined project management, the delivery approach including its technical side, and the supplier's performance. What we produced was findings and recommendations, not a verdict on anybody.

That appendix page was laid out to hold a comparison the review intended to build: the monthly status colour each party had reported, the client organisation on one row, the supplier on the other, across the engagement. Underneath it were three questions. Whether the two organisations held different definitions of the colours and different understandings of severity. Whether the team could have been more collaborative and flexible to get the thing finished. And whether the project was genuinely red, blocked and not moving forward, or whether the supplier's team had in fact been making progress.

A colour is testimony, not a measurement

Everything else in a governance pack reconciles to something. The cost line is checked against invoices, the date against a plan baseline, headcount against a payroll or a purchase order. The status colour is checked against nothing. It is entered by a project manager, reviewed by people who report to the same outcome, then treated downstream as though it were an instrument reading.

It is not a reading. It is a claim about work, filed by the people the claim is about.

That is not an accusation of dishonesty, and I want to be careful here, because the review was careful: it put the point as a question and asserted nothing about whose rating was wrong. But a colour compresses a messy month into one of three values, and every compression needs a rule. Where no shared rule exists, the compression takes its shape from whatever else is going on. A supplier that has just asked for more days does not want a red on its own report, and a client that has lost confidence does not hand a green back easily.

Notice which way the question cuts. Had the audit found the supplier shipping steadily behind a rating that said otherwise, the finding would have been unwelcome to the organisation that commissioned the review. A reviewer who writes that question down has already accepted that the answer might not be the one anybody ordered, which is the whole value of asking it.

The evidence was already in the room

What makes this page more than a rhetorical flourish is that the same document, a few appendix pages away, lists everything you would need to settle it.

The review kept an evidence register: fifteen classes of artefact, each recorded as obtained and separately recorded as read. The contract and statement of work, the requirements, the design documents, the risk log, the issue tracker export, the test plan, the change, defect and code fix logs, the monthly headcount bills, the communications plan, and both parties' weekly reports, obtained from both sides rather than only from the supplier. The plan itself came in two variants, an optimistic one and a contingency one, with a note that the team never moved off the optimistic one.

Every one of those is residue. Nobody wrote a defect record to influence a governance forum. Nobody logged a code fix to make a case. The bills exist because someone had to be paid, and a second organisation checked them because it was paying, which is a stronger guarantee than any status report carries. That is why the register's own comment on the billing records notes an absence of architect billing, a fact about how the project was staffed that appears in no status colour anywhere.

Put the colour beside the residue and the question becomes answerable.

Were defects closing during the months rated red. Was the code fix log active or quiet. Did change records cluster where the rating moved, or somewhere else. Did the weekly reports from the two sides describe the same week. Where colour and residue agree, the rating was doing its job. Where they diverge you have found either a reporting problem or a management problem, and either is worth more than the argument the two organisations were actually having.

Why nobody ran the audit

The honest reason this was not done properly is cost, and the review's own draft state proves it. The grid was designed, laid out and left to be confirmed. The question was cheap. The answer meant reading both organisations' weekly reports across the whole engagement, plus the defect, change and code fix logs and a stack of monthly bills, all of it needing a person to read it and hold the timeline in their head.

That is weeks of skilled time, spent after the project is over, to establish something nobody has an incentive to want established. So it does not get done, the colour stands unexamined, and the record of what actually happened is used for one purpose only, which is billing.

Meanwhile the same colour becomes an input. It feeds portfolio dashboards, supplier scorecards, escalation thresholds and, increasingly, the summarisation layer on top of all three. A field that reconciles to nothing is a poor foundation for a measurement practice and a worse one to point a model at.

What is different now

In that review the three letters meant red, amber and green. In most of my current conversations they mean retrieval-augmented generation, a coincidence I have started to enjoy, because the second sort is what makes the first sort checkable.

The reconciliation I described is a retrieval problem with a small reasoning step on top. Put the artefact corpus into a vector store, keep the lineage from every extracted fact back to the document and page it came from, and the reading job that used to take weeks becomes a scheduled run. Ask it narrow questions with checkable answers. Which artefacts moved this month, which of them are consistent with the reported colour, which are not, and quote the line. A copilot for a delivery assurance lead, not a replacement for one.

Two things have to be right for that to be worth having. The first is an evaluation set. Take closed programmes where the outcome is now known, run the reconciliation over them, and check whether it flags the months a human reviewer would have flagged. A retrieval assistant will reconcile whatever you hand it and sound confident either way, so the only way to know whether yours is any good is to test it where you already know the answer. The second is lineage per claim. A divergence note that cannot point at the document line that produced it is one more opinion, and the point of the exercise was to stop adding opinions.

Done that way the audit stops being an exercise and becomes a standing job. Early and carefully scoped autonomous experiments are reasonable here because the task is narrow, the sources are fixed, and the output is a draft for a named human to accept or reject. Nothing is being decided. The machine is asking the colour to show its work.

It lands at a useful moment for a regulated group. The European AI Act has reached political agreement, and while none of it binds a delivery assurance function today, the direction is clear enough: systems that inform decisions will be expected to show where their inputs came from.

A status colour is a claim about work, filed by the people the claim is about. Every other claim in a bank gets checked against a record. That one gets typed into a slide and believed.

What to fix before you automate any of it

None of this needs a data scientist, and all of it has to land before one is any use.

Write down what each colour means, agreed by both parties, before the work starts. That was the review's own recommendation, offered for the next engagement rather than as a criticism of the last.

Name, for each colour, the artefact that would corroborate it independently. If none can, that colour is a mood rather than a status.

Keep the artefacts joinable. One work item identifier surviving from the plan into the defect log, the change log and the invoice line is worth more than any dashboard built on them.

Build the evaluation set from your own closed programmes and keep it, because it is the only thing that will tell you whether the reconciliation is reading your estate correctly. Then keep the colour with a human and put the machine on the evidence.

That sequence is how a RealAI Platform engagement opens on delivery assurance work, and it is why we ask what your governance data reconciles to before anyone talks about models. Answer those and you have automated nothing yet. You have only made the automation possible.

Which leaves the interesting part. When the reconciliation cost weeks, not running it was a resourcing decision. It does not cost weeks any more.

Drawn from an independent post-mortem review of an outsourced application build at a European banking group's IT services subsidiary: its evidence register and the working questions the reviewers put to themselves. That review produced findings and recommendations, and the status comparison it designed was still unfilled in the draft I worked from. No status ratings appear here, because that grid carried none. Reading the open question as a measurement problem is ours.

A status colour is a claim about work, filed by the people the claim is about. Every other claim in a bank gets checked against a record. That one gets typed into a slide and believed.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?