Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking technology

Case study
Banking technologyA European banking group's IT services subsidiary

For four consecutive months the client reported red and the supplier reported green or amber, and neither party was lying

An independent post-mortem review of a failed offshore application build laid both parties' monthly status colours side by side across about ten months. The client ran green for four months, amber for one, then red for the remaining five. The supplier ran green for six months, amber for three, and reached red only in the final month, about two months before the statement of work was terminated. The divergence was not concealment. It was a definition: for the client, red meant the delivery date would be missed; for the supplier, red meant there were blockers to moving forward. Those measure different variables, so one project could sit in two colours on the same day with both reports accurate. The review rated the entire documentary estate, sixteen artefact classes including the plan, the risk log and both parties' weekly reports, as good quality, and still found no evidence of any method behind how the supplier set its flag. The recommendation was to clarify definitions anywhere the project documentation risked different readings. What was produced was a diagnosis and a phased set of recommendations. The build had already been taken back in house.

4 monthsLag between the client's first red and the supplier's
Client
A European banking group's IT services subsidiary
Duration
Independent post-mortem review, findings and recommendations
AI · RIDGE E86.4 N12.5ρmax 1.00
~10 monthsStatus reporting compared side by side, both parties
2Incompatible definitions of red inside one shared vocabulary
16 of 16Documentary artefacts the review rated good quality

Two organisations ran one project. They reported its health every month, in the same three colours, on the same cadence, into the same governance forum. Neither ever filed a report it believed to be false. For four consecutive months one of them was describing a project heading for a missed date while the other was describing work with little or nothing standing in its way, and both descriptions were accurate.

A European banking group's IT services subsidiary had contracted the build of a customer-facing application to an offshore development supplier. The build lost its quality, then its schedule, then the willingness of either side to believe the other's numbers. The statement of work was terminated and the work was taken back in house. I led the technical workstream of the independent post-mortem review that followed, and one appendix in that report is the cheapest lesson in the whole document, because fixing it would have cost an hour.

The challenge

The review put both parties' reported status on a single row each and lined them up month by month, across about ten months of the build. Read as one picture, the two series behave like a slow-motion collision.

The client ran green for the first four months, went amber in the fifth, and stayed red for the remaining five. The supplier ran green for the first six months, moved to amber for three, and reached red only in the final month of the comparison, roughly two months before the statement of work ended. The client's first red and the supplier's first red are four months apart. In the month where the gap is widest, the client was reporting red and the supplier was reporting green.

The obvious reading is that somebody was managing the news. It is the wrong reading. The two organisations were working to different definitions of the colour. For the client, red meant the project would miss its delivery date. For the supplier, red meant there were blockers to moving forward.

Those are not two strengths of the same signal. They are two different variables. A build can be entirely free of blockers and still miss its date, because the remaining work is simply larger than the remaining calendar. A build can be full of blockers and still hit its date, if it started with float or if the blockers sit off the critical path. The client was reporting a forecast about an outcome. The supplier was reporting a condition of the work in front of it. Both were answering their own question correctly, every month, and the governance forum was reading the two answers as though they were points on one scale.

What makes this worth writing down is how well everything around it worked. The review read sixteen classes of project artefact: the statement of work and its addendums, the master agreement underneath it, the requirements, the high and low level designs, the architecture diagram, the plan and its milestones along with the alternative schedules drawn up in case delivery slipped, the weekly reports from both sides, the risk log, the issue log, the test plan and acceptance criteria, the change, defect and code fix logs, the headcount bills, the communications plan and two code drops. Every one of the sixteen was rated good quality. The risk log had been maintained throughout. The paperwork was not the problem. The paperwork was in order and it still transmitted a false picture, because the schema underneath it was ambiguous and nothing in the estate defined the terms.

The review also recorded that there was no evidence of a method behind how the supplier arrived at its flag at all. Not a bad method. No visible method. The colour was a judgement, rendered monthly, in a vocabulary that had never been agreed.

The approach

The comparison itself is the technique worth stealing. Rather than audit either party's reports for accuracy, the review laid the two organisations' practices against each other on one timeline and treated the divergence as the finding. The same device was applied to how each side tested: the supplier tested each sprint in isolation, the client expected each sprint to re-exercise everything before it. Neither approach is wrong on its own terms. Run side by side without anyone naming the difference, they guarantee that dependency defects arrive late and arrive all at once.

The evidence base was deliberately balanced: fifteen interviews across both organisations, roughly half on each side, every one closed with notes reviewed and follow-up completed, plus direct inspection of the code.

The recommendation that came out of the status matrix is two lines long. Clarify definitions in the parts of project documentation where there is a risk of different interpretations, including interpretations that diverge for cultural or language reasons. In the phased plan attached to the report, that recommendation was not scheduled first. The plan has three bands. The earliest, measured in weeks, holds exactly one item: a joint workshop to walk both organisations through the lessons learnt. Aligning the reporting mechanism was filed in the last band, beyond three months, next to staff selection and estimation methodology, while governance and meeting structure sat in the middle band. The cheapest item on the list, and the one that would have surfaced most of the others earlier, was sequenced after the expensive ones. That is what a definition defect looks like from the inside. Nothing is broken, every report is filed on time, the paperwork is good, so it never presents as urgent.

The part that transfers to how these systems get built today is narrower than it looks, and sharper. If you fed both parties' entire reporting history into a language model and asked it to find the disagreement, it would find nothing, because there is no disagreement in the text. Each pack is internally consistent, correctly worded, honestly filed. The defect lives in the schema, in the meaning attached to a field, and the meaning was never written anywhere the model or the reader could see it. Summarisation over documents that were individually true cannot recover a definition that no document contains.

That is the same discipline a feature store enforces and for the same reason. A field is not a name and a value. It is a name, a value, a definition, an owner and a lineage back to the event that produced it. Status colour is a feature. It was being computed by two organisations from two different inputs and stored in one column.

The outcome

Nothing in this review repaired the engagement. The statement of work had already ended and the build had already come back in house. What the work produced was a set of findings, a recommendation against each one and a phased sequence for the next engagement. The status matrix is the piece of it I have used most often since, because the failure it describes is not rare and is not detected by any of the controls organisations usually buy.

Two moves close it. The first is a page: name every term the reporting vocabulary uses, define it in a sentence, have both parties sign it before the first report, and re-read it at the first divergence rather than the fourth month of one. Ambiguity across a supplier boundary is not a soft problem to be handled with better relationships; it is a defect in the control system, and it responds to specification.

The second is to stop declaring the colour. Most of what these two organisations were arguing about was already sitting in their systems: issue ages in the tracker, defect arrival and closure rates, how much of the plan's remaining work had a completed test behind it. Process mining over that event trail computes the schedule forecast and the blocker count as two separate, named, sourced numbers, refreshed as often as anyone wants, with a lineage trail back to the events that produced them. Two numbers, two definitions, both visible. The argument then happens over evidence rather than over a colour. Our Platform work builds that pair before anyone is asked to pick a colour, because a status that is declared rather than computed leaves nothing to argue with except the person who declared it.

Our Consult engagements now open by asking what each number on the client's reporting pack is computed from, and how often the room finds a field that means two things. It is a short conversation with an uncomfortable habit of not being short.

The status pack said green. The status pack said red. The reports agreed on the facts and disagreed on the language, which is the one failure mode a well-run governance process will carry, undetected, all the way to the end.

NEXT STEP

Ready to make AI real?