Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking IT services

Case study
Banking IT servicesA European banking group's IT services subsidiary

The review judged the escalation correct on timing and named it a possible reason the supplier grew hesitant about admitting timeline problems

A European banking group's IT services subsidiary outsourced the first phase of a customer-facing application to an offshore development partner, ran into quality and schedule trouble, escalated early through the governance route both sides had agreed, and eventually took the work back in house. The independent post-project review that followed found the client's process close to faultless: all documentation and expected governance present and evidenced, collaboration good in the majority of areas, and correct use of the hierarchy to raise the issue. In the same summary it recorded that escalation was called early and rightly so, but that this may have discouraged the supplier and made it hesitant about admitting timeline issues. A separate working page found that reaching for the contract in the working relationship was not helpful in building it. The review produced seven findings and eight recommendations, five of which asked the client rather than the supplier to change, including an alternative approach to deadline management that drops penalty-clause language on the reasoning that a supplier not being measured for punishment gives more honest estimates. What was delivered was a diagnosis and a set of recommendations for the next engagement, not a recovered project.

5 of 8Recommendations from a supplier review that asked the client to change
Client
A European banking group's IT services subsidiary
Duration
Independent post-project review, findings and recommendations
AI · RIDGE E81.8 N81.3ρmax 1.00
7Findings across project management, technology and delivery
4 weeksUpfront framework training the supplier itself said it should have had
1 of 8Recommendations addressed to both parties together

There is a version of this project in which the client does nothing wrong and the outcome is bad anyway. It is not a comforting story, which is roughly why it so rarely gets written down.

A European banking group's IT services subsidiary outsourced the first phase of a customer-facing application to an offshore development partner. The build hit quality and schedule trouble. The client escalated, early, through the governance route both parties had agreed in advance. The work was eventually brought back in house, and the subsidiary then commissioned an independent post-project review to establish what had happened.

That review had to decide what to say about the escalation, and it said two things in a single sentence. Escalation was called early, and rightly so. And it may have discouraged the supplier and made it hesitant about admitting timeline issues.

Both halves are in the source. Keep only the first and you have a client vindication. Keep only the second and you have an argument against escalating at all, which is worse advice than the practice it criticises. The finding means something only held whole: the governance behaviour was correct, and doing it correctly is a plausible reason the bad news became harder to get.

The challenge

Start with how well the client came out of its own post-mortem, because the finding depends on it. All documentation and expected project governance was present and evidenced. Collaboration in the majority of areas was very good, and the subsidiary tried extremely hard to make the project work. A later page of the same report credited the client, without qualification, for using the management hierarchy properly to raise the issue. On process, this was a client doing the textbook thing.

Against that sat the delivery picture. The supplier was stretched by a new framework and a newly hired team at once, which set it back against the agreed timeline. Its own retrospective view was that it should have set aside four weeks of upfront framework training before writing production code. Testing plans existed and had apparently been agreed by both parties, but were not followed to the standard expected.

So the client had a real problem and raised it the moment the agreed thresholds said to. An escalation route is a mechanism for moving information upward: a problem that passes an agreed threshold stops being handled where it sits and goes to someone senior. It is a good mechanism for that. It says nothing at all about what it does to the incentives of the person the information has to come from.

What it can do, and what this review reached for as an explanation, is convert a shared problem into an attributed one. The moment a problem becomes a formal escalation, somebody owns it. A supplier that expects its next disclosure to be logged as further evidence against it will disclose later, in softer language, with more of the uncertainty resolved in its own favour. That requires no bad faith. It is the ordinary behaviour of a party that has worked out what reporting costs it.

The review found the same mechanism from a second direction. In its working notes on managing the offshore relationship it recorded, bluntly, that reaching for the contract was not helpful in building the relationship. And in its recommendations it proposed an alternative approach to deadline management that drops penalty-clause language entirely, on the stated reasoning that a supplier not being measured for punishment is more likely to be realistic and honest about what it can actually deliver inside the timeline. A finding, a working note and a recommendation, on three different pages, all point the same way: the review kept associating formal pressure with worse information coming back. Not one of the three is stated as a measurement. All three are stated as a risk.

The approach

The reason this review could hold both halves of the finding is a choice about evidence, made before any of it was gathered.

The report format put three columns on the same question: the client's account, the supplier's account, and the reviewer's own observation. The supplier was given a voice inside the client's post-mortem of the supplier. That is an uncomfortable design for the party paying for the review, and the only design under which you learn what the escalation looked like from the other end. A single-roster review would have recorded strong governance and early escalation, marked both as strengths, and stopped there.

The second choice was evidential discipline. Working observations were sent back with an instruction to show the artefact each claim rested on. A statement about how a supplier behaved after an escalation decays into anecdote unless it is pinned to something in the record.

The outcome

What the engagement produced was a draft report: seven findings across project management, technology and delivery, and eight recommendations. The distribution of those recommendations is the most honest thing in the document. Five ask the client to change: build more experience of the sourcing geography, cultural awareness training and visits to the supplier included; collaborate more on site, up to placing client engineers alongside the delivery team; manage deadlines without penalty-clause language; gauge supplier capability by practical test rather than by interview; and make sure the supplier understands the end-to-end deliverable rather than its own slice. They were written for that engagement and its particular supplier, not as a checklist to lift. Two ask the supplier to change. One is addressed to both parties together, and it is the one that answers the escalation problem: agree joint pass and fail definitions for testing, and joint status definitions, up front, so that both sides are working to the same standard of what good and bad look like.

That is the actual fix, and it is quieter than it sounds. A status downgrade agreed against a definition both parties signed is a reading. A status downgrade issued against no shared definition is an accusation, and it gets answered like one. The review also drafted a status-definition template for future engagements, and offered it with an explicit disclaimer that it did not imply the client’s past ratings had been wrong. The point of the disclaimer is the point of the template: agree what a status means before anyone needs it to settle an argument. Worth being precise about tense here: that template was proposed for adoption, not evidence of a practice that had run. This review produced findings and recommendations for the next engagement. It did not recover this one.

The pattern generalises past outsourcing, and it is already visible in AI delivery. A bank running copilots against its own document estate, retrieval-augmented generation over a knowledge base whose lineage nobody has finished mapping, or an early and carefully scoped autonomous experiment in a back-office process, has the same information problem in a shorter loop. Someone has to report that the evaluation set is not moving, that retrieval is returning the wrong document class often enough to matter, that the data readiness work is further behind than the plan says. That person is usually inside the programme whose funding depends on the answer being encouraging. Governance that treats every such disclosure as an exception to be owned will get fewer of them, later, and will mistake the silence for progress. The direction of European AI regulation sharpens this rather than solving it: the documentation and record-keeping obligations under discussion reward organisations that can show what a system actually did, not organisations that reported confidently.

The way out is the same as it was on this build, and it is not softer governance. It is reducing how much of your view depends on a self-report. Evaluation results, retrieval quality, defect ageing, build health and lineage are all recordable as they are produced, written to a trail rather than summarised into a colour by the party they reflect on. That is standard MLOps discipline pointed at a governance problem instead of a deployment one, and it is what the Platform work we do is for. Our Consult engagements now open by asking which delivery signals a client can see without asking anyone, because that number, not the escalation policy, sets the ceiling on how early bad news can possibly arrive.

The client in this review escalated correctly and paid for it in candour. Nobody designed that trade. It falls out of any oversight system where the only sensor is the party being measured.

NEXT STEP

Ready to make AI real?