Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking IT services

Case study
Banking IT servicesA European banking group IT services subsidiary

The review named four root causes behind the defect profile, and the fix on the table answered one of them

A European banking group IT services subsidiary commissioned an independent review of an application it had outsourced. The code review pinned its findings to twelve artefact IDs in the shared issue tracker across seven design sub-aspects, with duplication drawing three corroborating tickets on its own and caching escalating to the executive summary because what was cached was customer data. Behind that profile the review named four root causes: the delivery team was new to the framework, nobody wrote the requirements, the team was junior, and there was no technical architect on the project. The most concrete remedy in the report, four weeks of upfront framework training, was written by the criticised vendor and adopted word for word. It addresses one of the four. The other three are contracting, staffing and role decisions that no training course reaches. What was delivered here is a review with findings and recommendations, not a remediation.

1 of 4Root causes addressed by the proposed fix
Client
A European banking group IT services subsidiary
Duration
Independent delivery review, findings and recommendations
AI · RIDGE E81.8 N12.5ρmax 1.00
4 weeksUpfront framework training proposed as the remedy
12Tracker artefact IDs cited as evidence across seven design sub-aspects
0Architect lines found in the vendor's monthly staffing invoices

A review can be right about every one of its findings and still hand back a remedy that reaches a quarter of the problem. That is what happened here, and the distance between the diagnosis and the cure sits in the same report, a few pages apart, without anyone in the room appearing to notice.

A European banking group IT services subsidiary commissioned an independent review of an application it had outsourced. The review worked through the code, the plans, the contract and the people, and named four root causes: the delivery team was new to the front-end framework the application was written in, nobody authored the requirements, the team was young and inexperienced, and there was no technical architect on the project at any point. Separately, in its own words, the vendor proposed the fix. Four weeks of upfront training on any new software before starting a future project. That is a genuine answer to the first cause and no answer at all to the other three.

The challenge

The technical section did not assert that quality was poor and leave it there. It named a design property, then pointed at the ticket in the shared issue tracker that evidenced the breach. Seven design areas were assessed: responsibility boundaries, layering, event flow, coupling, duplication, caching, and how the development environment was set up. Twelve distinct artefact IDs were cited as evidence. Duplicated code drew three corroborating tickets on its own, more than any other area examined. Caching drew two, and caching is the finding that travelled upward, because what was cached was customer data. The executive summary named it as one of two defects that should not have existed at handover. The other was the absence of error handling.

Read that profile back against the four causes and it stops looking like carelessness. Duplication is what a team writes before it knows the abstraction the framework wants; you copy the block that worked because you cannot yet see the shape it belongs to. Coupling and layering breaches accumulate when nobody holds the boundaries between parts of a system. Caching a customer record is a local decision taken quickly to get a screen to respond, with nobody in the room whose job is to ask where that data is now allowed to live. The code was the symptom and the staffing was the disease.

The cheapest evidence for that came from an unglamorous place. The review's evidence register covered fifteen artefact classes. One row was the monthly staffing invoices, and the comment against it records that the absence of billing for an architect was noted. Not thin architect coverage. No architect line at all. The client separately reported that the architect had been available for transition and not for the project, and recorded against itself that it should have sent its own architect in once it doubted the vendor could deliver.

That is a control anyone can run monthly without a forensic exercise: reconcile the roles you contracted against the roles you are billed for, and treat a missing role as a defect forecast rather than a saving. The Platform work we do wires that in as a standing reconciliation across the staffing record, the issue tracker and the commit history. Assembled afterwards by interview it costs a five-week engagement. Assembled continuously it costs a query.

The approach

The review's instrument on the project management and delivery findings pages was a three-column grid: the client's account of a question, the vendor's account of the same question, then the independent observation. The criticised vendor had a column of its own on each of them. Six contractual test questions defined the delivery assessment, ending with the one no acceptance test measures and every buyer pays for over the following decade: is the delivered software maintainable.

The symmetry held on the client side too, which is what makes the document credible rather than merely adversarial. Five adverse observations were recorded against the vendor: insufficiently qualified staff, contested code quality, no requirements authorship, issues not taken seriously enough, and over-confidence in the client relationship. Four were recorded against the organisation that commissioned the review: schedule pressure that predated the engagement, a contract it may have rushed into, a testing approach it could have run differently, and its own architect kept out of the room. A review that only finds fault outside the building has not been run properly.

That format produced the report's most useful admission and its most misleading one in the same column. The vendor's own feedback said the collaboration could have been much better, that the testing approach should have been agreed at the start, and that it should have done four weeks of framework training upfront. The reviewer took the training line and carried it forward; it appears three times across the document. It is the only quantified remedy anywhere in the review, and it was authored by the party under criticism.

Set it against the four causes one at a time. Framework inexperience: answered, and answered well. Requirements authorship: not touched, and the source is specific about where the failure was not, because the client's own scope and non-functional documents were assessed as detailed and clear. The specification existed. What was missing was anybody converting it into buildable requirements on the delivery side, a contracting decision about who owns that step rather than a knowledge gap a course closes. Team seniority: not touched, because seniority is decided when the team is assigned. The architect: not touchable at all, because you cannot train an absent role into existence. Someone has to be named, staffed and paid.

Of the four causes, exactly one is technical, and it is the only one a training budget can buy its way out of. The other three are decisions about contract scope, staffing mix and role ownership, taken by people who will never be in the training room.

The outcome

What this engagement delivered was a review: an evidence register, findings pinned to tickets, four root causes, and eight recommendations, of which five required the commissioning organisation to change its own behaviour, two required the vendor, and one was joint. No remediation programme is reported, and the technical section states its own limit plainly, that it was a high-level review rather than a deep-dive code audit. It is a diagnosis, and should be read as one.

Three things would be done differently now. The first is that duplication, coupling and layering are continuously measurable from the repository, so the staffing signal that took a review to surface can be read weekly. Clone density is a read on the team, the tooling and the time pressure, not on individual discipline, and it moves before the defects do.

The second is that the evidence discipline the review set for itself was partly unexecutable. Every finding was supposed to carry an artefact reference, and the identifier column in the register was never populated, so the chain the reviewer was told to produce had nothing to point at. Process mining over the delivery workflow, with the tracker, the build record and the staffing record joined into one lineage trail, produces that chain as a by-product rather than as homework.

The third is that the current wave of cautious language-model pilots in code review changes less of this than the demonstrations suggest. A model reads a diff well and flags duplication cheaply. It does not decide who owns the abstraction that should have existed, and where nobody owns it the flag lands in a queue and the duplication stays. Our Consult engagements start by asking who holds the boundaries and whether that person appears on the invoice, because a finding with no owner will be found again next quarter.

The review was right about all four causes. The organisation was offered a fix for one of them, priced in weeks, written by the party being reviewed, and it was the most concrete thing in the document. That is usually how the wrong remedy wins.

NEXT STEP

Ready to make AI real?