Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking technology

Case study
Banking technologyA European banking group's IT services subsidiary

The review told the client to stop taking the supplier's word on who could code and interview the pool itself

A European banking group's IT services subsidiary commissioned a post-mortem review of an offshore software build that ended before user acceptance testing. Among the possible reasons the review gave for the state of the code was that the supplier had staffed the project with inexperienced developers and had not worked to the coding standards set out in the architecture document. The client's own record shows the concern was live during delivery rather than discovered afterwards, and it was still recorded as a concern, not a fact, because nobody on the client side had any way of checking. The review's recommendation on that point was addressed to the client. Revise the intake: a fifteen to twenty minute yes-or-no interview across the pool of developers offered, then deeper competency testing differentiated by role for those who pass, then continuous monitoring through onboarding, execution and churn. All three are recommendations. The contract had already ended, and no such screen was operated on this build.

15 to 20 minThe first-stage screen the review told the client to run on the supplier's developer pool
Client
A European banking group's IT services subsidiary
Duration
Post-mortem project review, findings and recommendations
AI · RIDGE E54.5 N12.5ρmax 1.00
Yes or noWhat that first stage is asked to produce, before anyone tests for depth
20 to 30Developers on a single ramp-up, which the review tied to teams that stopped seeing the whole
RecommendationStatus of the three-stage intake; the contract had ended before the review reported

An extended workbench arrangement is a way of buying hours. It is not a way of buying capability, and the two get confused because the invoice looks the same either way.

A European banking group's IT services subsidiary commissioned a post-mortem review of an offshore software build that ended before it reached user acceptance testing. The review read the contract, the project record, an interview programme and the code. Its account of why the code came out the way it did includes a clause most readers skim: among the possible reasons for the quality problems was that the supplier had put inexperienced developers onto the project, and had not followed the coding standards set out in the architecture document.

The interesting part is where the review sent the remedy. Most of the technical recommendations went to the supplier, as you would expect. This one went to the client. If who writes your code determines whether the build works, the review said in substance, then checking who writes your code is your job, and you should build an intake to do it.

The challenge

The relationship was described in the review as an extended workbench: the supplier bid for individual projects inside the subsidiary rather than holding a strategic partnership with it. In that arrangement, staffing is the supplier's business by design. The client buys a team, not named people, and the assurance that the team can do the work is a contractual promise rather than an observation.

That promise held right up until it did not, and the record shows the client watching it fail in real time without an instrument to prove it. A note beside the programming fundamentals findings says that there was a problem in the initial sprints because the team was new, that this was raised and resolved, and that fundamental errors persisted anyway. Read that sequence carefully. Individual defects were logged, tracked and closed by agreement. The thing producing them was not a defect and no ticket could close it.

The second pressure was churn. One ramp-up took the development team from twenty to thirty people, and the review connected that scaling directly to teams that turned inward, each developing its own modular piece and none holding the shape of the whole. Adding people to an offshore team is the easiest lever a supplier has and the one with the least visible cost to the buyer, because the extra names arrive as a line on a staffing report rather than as a fall in coherence.

The third is the one that should worry any buyer of delivery capacity. Late in the project, with the backlog piling up sprint after sprint, the client still had concerns about whether inexperienced developers were on the project. Concerns. Months into a build it was paying for, the client could not answer from evidence a question as basic as who was writing its software and what those people could actually do. The review recorded it as a transparency and trust problem, which it was, but a transparency problem is only a problem if you have no independent way to look.

The approach

The recommendation the review put to the client has three stages, and the sequencing carries most of the argument.

The first stage is a fifteen to twenty minute interview conducted across the pool of developers offered, producing a yes or a no and nothing finer. Two properties do the work here. Twenty minutes is cheap enough to run across an entire pool rather than a shortlist, and a screen you can only afford for four people is not a screen, it is a formality performed on whoever the supplier put at the front. And a binary output resists the pressure that kills every scored intake: a number can be argued down when the start date is close, whereas a no has to be overturned by someone willing to sign their name to it.

The second stage tests depth, and it tests differently by role. Project managers, architects and developers were each to be examined on their individual competencies, in writing or orally, across basic, moderate and advanced skills. Role differentiation is not administrative tidiness. On this build the architect was the role the client kept asking for and waiting on, and an architect who scores well on developer questions is a good developer with a title. Testing everyone against one paper hides exactly the shortfall that hurt most.

The third stage is the one organisations always agree to and never operate: continuous monitoring through onboarding, execution and churn, so that knowledge is created, retained and handed on without a break in service. It exists because the first two stages screen a moment. Nothing in a good interview survives the replacement of half the team six months later, and on this project the team did change size and shape while the work was in flight.

One detail keeps the whole thing from being unilateral gatekeeping. The review tied the intake to success factors defined upfront with the supplier, so that both sides know in advance what the bar is and what evidence clears it. An intake the supplier has never seen is a trap. An intake it helped calibrate is a standard, and the supplier can staff against it rather than around it.

The outcome

Nothing here was operated. This engagement produced a post-mortem, and the intake is a recommendation inside it, offered for future projects to a client whose contract on this one had already ended. No pool was screened, no competency test was written, no monitoring ran. The build had failed before anyone drew the remedy, which is the honest and slightly bleak shape of most assurance work: the instrument that would have caught the problem gets specified after the problem has finished being expensive.

What has changed since is that the same failure is now easier to have and harder to see. When a supplier's team builds a screen or a report, weak capability shows up as a crash, a slow page, a defect somebody can log. When the same team builds machine learning pipelines, it shows up as a model that trains, deploys and returns plausible numbers that are quietly wrong: a feature computed differently in serving than in training, a leak from the label into the inputs, an evaluation split that shares customers with the training set. None of those raise an exception. They pass user acceptance testing. The failure surface has moved from something a tester can see to something only a competent reviewer can find, which raises the value of screening the reviewer.

That is also why data readiness and lineage belong in the intake conversation rather than beside it. If the pipelines your supplier builds record what each model saw and where every feature came from, a capability problem becomes traceable after the fact. If they do not, you are back to trusting a staffing report. Our Consult engagements open on that pairing now, and the lineage half of it is what our Platform work is built to leave behind, because the two questions, who is building this and what will their work leave behind as evidence, are one question that most organisations ask of two different departments.

The cautious language-model pilots currently running in document handling and code assistance change the pressure but not the principle. A weak developer with a fluent assistant produces more code that looks right, faster, which makes the twenty-minute screen more useful, not less.

The lesson survives the technology. Capability assurance is the one part of a sourcing arrangement that cannot be delegated to the party being assured.

NEXT STEP

Ready to make AI real?