Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking risk and compliance

Case study
Banking risk and complianceA European banking group IT services subsidiary

Twelve elements on the drawn system view, and only two of them are the advice

An independent post-delivery review of an outsourced securities advisory platform at a European banking group IT services subsidiary began with a problem that had nothing to do with code: no single document showed what the system contained. The reviewer drew the system himself and asked for the high and low level design documents to confirm the drawing. It came out at twelve labelled elements, of which the business rules and the calculation engine are the advice and the rest are data supply, client holdings, calls to other systems, two databases, branding assets, marketing documents, a configuration tool, a workflow tool and a signed-on web front end. The total allocation was 2,576 person-days, of which 996 sat in the development phase, and a revised plan took the total to 3,047. The two defects the review's summary named at handover as things that should not have existed were both outside the advice logic: caching of customer data, and missing error handling. This was a review. It produced findings and recommendations, not a fixed platform.

2 of 12Drawn system elements that produce advice rather than evidence
Client
A European banking group IT services subsidiary
Duration
Independent post-delivery review, findings and recommendations
AI · RIDGE E50 N56.3ρmax 1.00
2,576Person-days allocated, of which 996 sat in the development phase
3,047Person-days in the total after the plan was revised
15Classes of document in the review evidence base

The reviewer could not find the system. The code existed and the system had been built. What was missing was any single document that showed what the thing was made of. So he drew it himself, and then asked for the high level and low level design documents in order to check the drawing against them, which is the reverse of how that is meant to work.

The drawing came out at twelve labelled elements. Securities data, meaning prices and risk levels. A mainframe holding client data and holdings. Calls out to other systems for data. Two databases. Images and branding. Marketing documents in printable form. Business rules. A calculation engine. A configuration tool. A workflow tool. A web user interface with secure sign-on.

Read that list once and the shape of a regulated advisory platform is already visible. Two of the twelve are the advice: the business rules and the calculation engine. That classification is mine rather than the report's, and it is the whole argument, so it is worth being explicit about what the other ten are doing. They are there because a bank that tells a customer to hold one instrument rather than another has to be able to reconstruct, afterwards, what that customer was shown, what data it rested on, who was allowed to change it, and what it looked like on the screen at the time. Not one of those ten is a feature anybody would demo.

The challenge

A build plan gets written against the two. The estimate, the sprint board, the demo script and the conversation with the sponsor all attach themselves to the part that produces the recommendation, because that is the part that has a business owner who can describe it. The other ten arrive as assumptions, and assumptions do not get sized.

The numbers in this engagement carry the same lopsidedness. The total allocation for the first phase was 2,576 person-days, of which 996 sat in the development phase. Most of the commitment was therefore somewhere other than writing the software, and a revised plan later pushed the total to 3,047. Nobody in the review treated that split as the anomaly, and it is not one. It is what a platform of this shape costs. The anomaly is that a plan built around the advice logic tends to discover that split rather than start from it.

The written requirements did not close the gap either. The review read the business, technical and non-functional requirements together and noted that the non-functional set said mobile capability did not need to be worried about, a line the reviewer flagged as something a reader could misunderstand if it were read on its own. That is a small observation with a large tail. When the inventory of the system lives in nobody's head as a whole, individual requirement lines stop being interpretable, because there is no picture for them to sit inside.

The evidence base for the review says the same thing from a different direction. Fifteen classes of document were pulled and read: the statement of work with its addendums, the master services agreement, the requirements in all three flavours, the high and low level design documents, the architecture diagram, the project plan with its optimistic and pessimistic variants, weekly reports from both sides, the risk log, an export from the issue tracker, the test plan and acceptance criteria, the change log, the defect log, the code fix log, the monthly staffing bills, and the communications plan setting out committees, governance and approved decision makers.

That is a properly assembled record of how a project was run. Thirteen of the fifteen describe process. Two describe the artefact, the design documents and the architecture diagram they contain, and those were used to confirm a drawing rather than to supply one.

The approach

The technical half of the review worked the codebase against a checklist, and what that checklist reached into is itself instructive. Whether information could leak out of the application or be forced out of it. Whether the system stayed correct when several things happened at once, and whether it let go of what it had taken. Whether it noticed its own failures and wrote them down. Whether a release could be put out and taken back again. Whether anyone could support it after handover. Whether it held up in more than one language and on more than one size of screen.

Not one of those questions asks whether the advice was right. They are all about the ten, and a reviewer with a fixed budget spent it there because that is where a regulated platform fails in ways that matter to a regulator rather than to a product manager.

The confirmation came at handover. Two defects were named in the summary as things that should not have existed at the point the software changed hands: caching of customer data, and an absence of error handling. Neither is a fault in the recommendation logic. One is a confidentiality question about where client information was allowed to rest, and the other is the difference between a system that can tell you it failed and a system that cannot. Both sit squarely in the ten. The product was described in the same summary as impressive and about to go live, which is the point: a platform can be impressive on the two and unacceptable on the ten at the same time.

The outcome

What this engagement produced was a review: findings, observations and recommendations across project management, delivery approach and software quality. The build had already been brought back in house before the review began, and the review fixed no component of the platform. The person-day figures are allocations and plan revisions, not measured effort. The defect observations are the report's, drawn from the tracker and the logs listed above.

The honest question is what would be done differently on the same platform today.

The drawing would not be drawn. A component inventory is derivable from the build: deployment manifests, dependency graphs, the set of external endpoints the application actually calls, the tables it actually reads. Producing that by hand at the end of an engagement is expensive and produces a picture that is stale on the day it is finished. Producing it from the pipeline gives the same picture, refreshed, with a source attached. Half the value of the reviewer's drawing was simply that it existed, and existing is the cheap part to automate.

The second change is that the ten are now the part that grows. If a calculation engine gains a model rather than a rule table, the advice logic barely changes shape while the evidence half doubles. Something has to record which version of the model ran, which features it saw, which data vintage those features came from, and what the customer was shown as a result. That is lineage, and it is the same discipline that model deployment work already demands, arriving in an advisory platform under a different name. The Platform work we do starts at that inventory rather than at the model, because a model with nowhere to write its reasoning is a demonstration.

The third is that the early and cautious language-model pilots now running in client communication land in exactly this part of the diagram. A model that drafts the rationale a customer reads is writing into the branding and marketing boxes, and those boxes are already governed. Our Consult engagements begin by asking where the output lands and who signs it, because that question separates a pilot that can be approved from one that cannot.

The inventory is the argument. Most of a regulated advisory platform is not advice, and a plan that does not say so on its first page is a plan that will discover it later, with the schedule already spent.

NEXT STEP

Ready to make AI real?