Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking IT services

Case study
Banking IT servicesA European banking group IT services subsidiary

Sixteen of sixteen artefacts rated green and fifteen of fifteen interviews closed complete, and then eight of the ten questions the review was hired to answer came back No

A European banking group IT services subsidiary commissioned an independent post-mortem review after an offshored build failed to reach user acceptance testing. The review published its evidence base: sixteen artefact types, each marked evidenced and reviewed, each rated green for document quality; fifteen interviews, each tracked through six stages from request to follow-up, each closed at the register's top state on a three-state legend whose two trouble states went unused. Then it answered the ten questions it had been commissioned to answer. Eight came back No, covering contractual delivery, solution design, software quality, requirement conformance, process compliance and maintainability. One was an explicit split. The single unqualified Yes was that the client's own requirement specifications met common international practice. The output was a findings pack and a set of recommendations, not a repaired application.

16 of 16Artefacts rated green for document quality
Client
A European banking group IT services subsidiary
Duration
Independent post-mortem review of a terminated offshore build, findings and recommendations
AI · RIDGE E63.6 N12.5ρmax 1.00
15 of 15Interviews closed complete, including every participant from the supplier
8 of 10Commissioned questions answered No
4 monthsLag between the client's first red status and the supplier's

A review can rate every document in an engagement as good and still conclude that the thing those documents describe does not work. That is not a contradiction, and it is not a sign the rating was wrong. It is what happens when an organisation checks whether a process ran and reads the answer as evidence the process worked. Those two questions look alike on a register and measure different things.

A European banking group IT services subsidiary commissioned an independent post-mortem review after an offshored build failed to reach user acceptance testing and the statement of work was ended. The review built its evidence base in public, which is the right instinct. It listed sixteen artefact types, marked each evidenced and reviewed, and rated each for document quality. All sixteen came back green. It listed fifteen interviews, tracked each through six stages from request to completed follow-up, and rated each on a three-state legend. The two states that would have signalled trouble went unused. All fifteen closed at the top state, including every participant from the supplier whose work the review was about to condemn.

Then it answered the ten questions it had been hired to answer. Eight came back No.

The challenge

The sixteen artefacts were not a token list assembled to look thorough. They covered the statement of work and its addendums, the master agreement, the business, technical and non-functional requirements, the high and low level design documents, the architecture diagram, the project plan, the weekly reports, the risk log, the issue tracker export, the test plan, the change, defect and code fix logs, the monthly headcount invoices, the communications plan, and two builds of the code itself.

The green rating was defensible. The documents existed, were current, and were coherent enough for a third party who had been in none of the rooms to follow what had been agreed. Several yielded findings on their own reading. The monthly headcount invoices carried no billing line for an architect, which is an invoice fact rather than an opinion about staffing. The non-functional requirements carried a line saying mobile capability need not be a concern, flagged as something that could have been misread in isolation. The project plan carried a contingency variant alongside the base plan, and the project stayed on the base plan throughout.

So the register was doing real work. It was also, quietly, answering a different question from the one its colour implied. Green meant the artefact was present, legible and good enough to work with. It did not mean it had governed anything. A contingency plan that exists and is never invoked is a good document and a failed control. So is a weekly report filed on time every week describing a project the other party does not recognise.

The verdict sheet was answering the other question. The chosen approach did not fit the expected outcome. The development process was not compliant with the usual procedures. The supplier did not deliver the output specified in the contract. The solution design and its implementation did not meet contractual obligations. Software quality did not meet contractual obligations. The delivered software did not meet the requirement specification, and was not maintainable. The procedures did not meet common international practice. Eight questions, eight No. One, whether project management had been optimal at all levels, was answered with an explicit split rather than a hedge in prose. And one came back Yes: the client's own requirement specifications met common international practice.

That single Yes is the load-bearing one, because it removes the standard defence. When a build fails, the first thing said in the room is that the brief was bad. The review checked that first and cleared it, which is what makes every other No stick. Elsewhere the application is described, in a figure the report quotes rather than measures, as seventy percent complete, error-prone and unfit for acceptance testing.

The approach

What let the review reach the second set of answers was that it declined to treat any artefact as its own evidence, and it published its limits on its method page: only a subset of the code was assessed, and the expert review was high level, touching key aspects rather than the whole. A review that states what it did not read survives being argued with.

Three moves did the work. The first was reading each artefact for content rather than for existence, which is where the missing architect line and the ambiguous requirement came from. The second was reading the status history and asking what its words meant. Ten months of parallel monthly reporting had been filed by both sides, in the same three colours on the same cadence, on incompatible definitions of red. For the client, red meant the delivery date would be missed. For the supplier, red meant blockers to moving forward. A project can be free of blockers and certain to miss its date, so both sides could file honestly and describe different worlds. The colours bear it out. The client went green for four months, amber for one, then red and stayed red. The supplier stayed green for six months, amber for three, and turned red once, four months after the client's first red and roughly two months before the statement of work ended.

The third move was reading the software: two builds, ten weeks apart, the second produced after a grace period granted specifically to fix what the first review had found. No improvement sufficient to allow acceptance testing was observed.

The outcome

What the engagement produced was a findings pack, a set of recommendations and a proposed lessons-learned workshop. No application was repaired inside it, and the numbers above are ratings and verdicts, not results. The gap at the heart of it was never written down as a finding in the report's prose. It appears only when you read the colour of one column against the answers on another page, which is how this defect survives inside a live programme.

The prescription is not more governance. It is a second instrument beside the first. Let the register keep counting artefacts, and put beside it a measurement drawn from the work rather than from the filing. Most of that already exists as exhaust. The issue tracker, the defect log, the code fix log, the build history and the monthly invoices are event streams with timestamps. Process mining over them answers what a register cannot: how long a defect sat before anyone touched it, whether the fix rate kept pace with discovery, whether the roles named in the statement of work ever appeared on a monthly bill. That last check reconciles two documents the review read separately, and it would have run every month for the price of writing it once. The Platform work we do wires those streams into standing checks, so the second instrument runs beside the register rather than after the termination.

The same applies to the definitions under the colours. Before a dashboard is shared across two organisations, the words on it need a dictionary both sides sign: what counts as done, what counts as blocked, what red means. Without one, structured reporting produces agreement rather than information, and the agreement holds right up to the failure. Our Consult work starts there, because it is the cheapest item on the list and almost never in place.

Two closing notes. A control that fires when its condition occurs beats one reviewed at a monthly meeting, and the contingency plan nobody invoked here failed because nobody had defined the trigger that would switch to it. And the current wave of cautious language-model pilots in document handling changes none of the arithmetic. A model that reads a statement of work well and drafts a status report well makes the artefact estate faster to produce and no more predictive of the outcome. It raises the count of documents that would be rated green.

Sixteen documents were good. Fifteen interviews were complete. The software was not compliant, not conformant, not maintainable. An organisation that grades itself on artefact completeness will pass its own audit the month before it terminates the contract.

NEXT STEP

Ready to make AI real?