Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking operations

Case study
Banking operationsA European banking group IT services subsidiary

A replan partway through the build put the total past 3,047 person-days against 2,576 allocated, and the project carried on the plan it started with

A European banking group IT services subsidiary outsourced the first phase of a customer-facing securities advisory application. The work order allocated 2,576 person-days in total, of which 996 sat in the development phase. A revised plan produced partway through put the total above 3,047 person-days, at least 471 above the allocation, while the build was still running. The project continued on the original plan and was eventually taken back in house. The independent review that followed found every governance artefact present and every escalation correctly raised. Fifteen classes of project document were requested and all fifteen were produced. The failure was not detection. It was that a forecast breach triggered nothing, because the two sides had agreed no status definition a forecast could fail, and because nothing in the record attached a condition to the contingency plan that would have switched the project onto it. What this engagement produced was a findings and recommendations report, not a recovered project.

3,047 vs 2,576Forecast person-days against person-days allocated
Client
A European banking group IT services subsidiary
Duration
Five-week independent review, findings and recommendations
AI · RIDGE E59.1 N18.8ρmax 1.00
471 PDGap between the revised plan total and the contracted allocation
1 of 2Plan variants the project ever ran on
15 of 15Artefact classes requested, produced and reviewed

Most organisations that have run a large build have a story about a budget that went. Fewer have the document saying it was going to go, written while there was still time to act on it.

A European banking group IT services subsidiary outsourced the first phase of a customer-facing securities advisory application to an offshore development partner. The work order allocated 2,576 person-days in total, of which 996 sat in the development phase. Partway through, a revised plan put the total above 3,047 person-days, at least 471 above the allocation and close to a fifth more than the contracted figure. It was written down while the build was still running.

The project went ahead on the original plan. It was eventually insourced in full, after quality and delivery deteriorated far enough that the subsidiary took the work back. I led the technical workstream of the independent review that followed.

The overrun is not the finding. Builds overrun. The finding is that this one was predicted in writing, by the people running it, and the prediction changed nothing.

The challenge

The review looked for the failure in the usual places and did not find it there. Fifteen classes of project artefact were requested and every one was produced: the statement of work and its addendums, the master contract, the requirements, the design documents and the architecture diagram inside them, the project plan, weekly status reports from both parties, the risk log, the issue tracker export, the test plan, the change and defect and code fix logs, the monthly headcount bills, and the communications plan. All fifteen were evidenced and reviewed. Nothing was missing.

The governance was not missing either. Meetings were minuted, correspondence was logged, and the risk log was updated throughout rather than abandoned after the first month. When the client project manager developed concerns, they were escalated early through the agreed route while the project kept running. The review found it hard to fault the client-side project management on documentation or process. It also noted the second-order effect of that early escalation: raising the alarm promptly may have made the supplier more guarded about admitting timeline trouble.

So this was not an organisation that could not see. It saw, it wrote it down, and it filed the writing. What it lacked was the mechanism that converts a forward-looking number into an obligation.

Two details make that concrete.

The first is that the project plan existed in two variants, an optimistic one and a contingency one. Someone had done the work of imagining the branch before the project started, and the review found that the team stayed on the optimistic variant for the life of the project. The contingency plan was not absent. It was unreachable: nothing in the review's evidence base named the reading that would have moved the project onto it. A plan with no trigger attached is an essay.

The second is that nothing the two sides had agreed carried a threshold a forecast could cross. The review's remedy for the next engagement was to settle shared status definitions with the supplier before work starts, which is a remedy only because no shared definition existed on this one. A revised plan heading past 3,047 person-days therefore failed no test, because there was no agreed test for it to fail, and a number that fails no test belongs to nobody.

The approach

The review method was ordinary, and the ordinariness is the point. Five weeks of it. Interviews on both sides. A full read of the fifteen artefact classes. A sample technical review of the code against a checklist covering design, architecture, testing, performance, error handling, security and deployment.

The most useful finding came out of the least interesting document. The monthly headcount bills were read line by line rather than totalled, and what they turned on was a role that never appeared. No technical architect was billed across the engagement, and the absence was noted against the bills themselves. The client had said in interviews that the architect needed to be available throughout and not only during transition, and the review's own technical findings named the missing architect as a cause of what went wrong in the code. Those were judgements. The bills were a record.

That is the shape of nearly every finding here. The evidence was already inside the project's own paperwork, in a form nobody had been asked to read against anything else. The revised plan and the original work order sat in the same repository, as did the billing records and the architecture complaints. Both parties filed a weekly status report for the same weeks, using the same colour scale with different thresholds for what counted as serious, and nobody during the project put the two side by side.

The outcome

What this engagement produced was a findings and recommendations report, not a recovered project. The build had already been taken in house before the review began, so nothing here is a delivery result. The recommendations were aimed at the next engagement: agree joint pass and fail testing definitions and shared status definitions with the supplier before work starts, make continuous architectural ownership an obligation the supplier carries rather than an intention both sides share, and test supplier capability with practical exercises rather than interviews. The supplier itself said four weeks of upfront training on the front-end framework would have helped, an admission that is only useful before a contract is signed.

Reading it back now, with retrieval-augmented systems sitting in front of document collections, the part that has genuinely changed is narrow. Nobody needed a model to predict this overrun. The prediction was already in the file, in plain prose, produced by competent people doing their jobs. What nobody had was a reader.

Point retrieval at the programme's own record: the contract, the work order, both sets of weekly reports, the risk log, the issue tracker, the change log, the bills. Ask one question ahead of every governance meeting, which is whether anything filed since the last one contradicts the plan of record. A revised plan totalling more than 3,047 person-days against a work order for 2,576 is not a subtle inference. It is a comparison of two numbers that no human had been made responsible for holding together.

Three conditions decide whether that works, and none of them is the model. Data readiness comes first: this collection happened to be complete, all fifteen artefact classes present and legible, which is rarer than it sounds. Lineage comes second, because a plan revision has to be a versioned event with a named predecessor rather than a fresh file with a fresh name. An evaluation set comes third, built from past programmes where both the forecast and the outcome are known, so you can measure whether the reader surfaces the contradictions that mattered rather than flooding a steering committee with the ones that did not.

None of this is autonomy, and it should not be sold as any. It is a copilot for the programme office, scoped to reading and flagging, with a person still deciding. The autonomous experiments worth running are narrower again, and should be scoped that way deliberately: watch for the trigger condition on a contingency plan and raise it when it fires, a job with one defined input and one output. Anywhere a system starts making the consequential call instead of surfacing it, documented human oversight stops being good manners and becomes the thing a regulator asks to see, and a reader that only flags contradictions stays well clear of that line.

The Platform work we do starts where this review started, with an inventory of what a programme already writes down and where the writing lands. Consult engagements that open by asking what an organisation already knows and cannot see tend to be shorter and more uncomfortable than those that open with a model.

The lesson is not that these people were careless. They documented more thoroughly than most, escalated earlier than most, and produced the forecast that would have saved them. They had a machine for generating warnings and none for receiving one.

NEXT STEP

Ready to make AI real?