Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking risk & compliance

Case study
Banking risk & complianceA consumer and SME banking group operating across several European markets

The platform is designed to go live at week twelve and the automated controls at week twenty-four, and the plan names the twelve weeks in between as manual work rather than leaving them blank

A consumer and SME banking group commissioned a board-level design covering governance, organisation and skills. One page of it is a twenty-four week implementation playbook in three phases: mobilise in weeks zero to four, stand up a platform and run first pilots in weeks five to twelve, put the regulatory checks into the delivery pipeline as policy as code in weeks thirteen to twenty-four. The platform milestone is designed for week twelve and the first release of control automation for week twenty-four, which leaves a designed twelve-week window in which governance is performed by people. The playbook states that window as an activity rather than leaving it as a gap. Every figure here is a target the design commits to. Nothing had been run when the plan was written.

24 weeksPlanned horizon from mobilisation to the first release of automated controls (design target)
Client
A consumer and SME banking group operating across several European markets
Duration
Design deliverable, twenty-four week implementation plan
AI · RIDGE E0 N0ρmax 1.00
Weeks 5-12Window the plan designs to run on manual control gates
Week 12Planned platform milestone: registry, feature store, delivery pipeline
2-3Pilots the plan schedules for the manual-gate phase (planned, not delivered)

Most AI programme plans carry a compliance bar that runs the full width of the page: controls, unbroken, from the first week to the last. The bar gets drawn that way because nobody wants to be the person who writes down the weeks in which the controls are people rather than code.

A consumer and SME banking group operating across several European markets commissioned a board-level design covering governance, organisation and skills. One page of that deliverable is an implementation playbook for the target operating model: twenty-four weeks, three phases, each with named activities, named outputs and, in one case, a named dependency. Nothing on the page had been run when it was written. It is a plan, and the numbers below are what the plan commits to, not what it produced.

The shape is small enough to hold in your head. Weeks zero to four, mobilise. Weeks five to twelve, baseline: stand up a platform and run the first pilots. Weeks thirteen to twenty-four, embed: put the regulatory checks into the delivery pipeline as policy as code, open the communities of practice, move pilots into production. The platform milestone sits at the halfway point, and the first release of automated controls sits twelve weeks behind it.

That distance is the design decision, and the playbook does not hide it. The baseline phase carries an activity written in plain words: implement manual control gates for pilot models. Twelve weeks of governance performed by humans, deliberately, printed on the page a board is being asked to approve.

The challenge

The EU AI Act is in force, and the assessment behind this plan splits the group's AI estate roughly ninety to ten by role. About 90% of it arrives as a deployer, through vendor tools and assistants bought and switched on inside business processes. About 10% arrives as a provider of models the bank builds itself, such as SME credit scoring. That split drives the sequencing, because deployer duties land the moment a licence is activated while provider duties land when a model is built. A bank can be in scope before it has written a line of its own code.

The design then sets four numeric compliance targets: full coverage of the high-risk inventory, zero prohibited practices in operation, a drift tolerance under 5%, and incident reporting inside 24 hours. All four are targets chosen on paper, not readings taken off a running estate. Each is also a quiet claim about a mechanism. A drift tolerance under 5% means something measures drift. A 24-hour reporting clock means something detects the event that starts it. Full inventory coverage means a registry someone keeps current.

None of those mechanisms exists in week one, and there are only two honest answers to that. Delay the platform until the controls that govern it are automated, which pushes the first validated model past week twenty-four and leaves the deployer nine tenths of the estate ungoverned while everyone waits. Or stand the platform up earlier and put people in the gate until the pipeline can carry the load.

The playbook takes the second route and dates it. Writing the manual months into the plan is what turns the schedule into a commitment rather than an ambition.

The approach

Phase one, weeks zero to four, contains no models at all. Appoint a head of AI for the centre and business AI leads for the units. Charter the oversight committee and finalise the risk-tiering policy. Select the top three use cases. The stated outputs are a signed RACI and an approved funding model. Four weeks, two artefacts, nothing shipped. Choosing the use cases is a phase-one decision here rather than a phase-two discovery, which is the difference between a programme that starts and one that keeps meeting.

Phase two, weeks five to twelve, is where the platform arrives: a model registry, a feature store and a basic delivery pipeline, described as a first workable version rather than a finished estate. Two to three pilots run on it. Governance is at version one, meaning manual control gates on pilot models. One dependency is flagged on the whole page, and it is cloud environment access. A single named dependency is itself a statement of belief: the plan expects the binding constraint to be infrastructure access, not talent or model quality.

The manual gate is where a plan like this either earns its keep or quietly fails. A gate is not a placeholder for automation that has not arrived. It is work, and it has to be specified as work: who signs, against what evidence, filed where, at what frequency. Elsewhere in the same deliverable the control library gives every control an accountable owner role, a named evidence artefact and a stated frequency, with the evidence landing in one place rather than scattered across inboxes. A manual gate built on that produces a file an auditor can read. A manual gate without it produces a meeting.

Phase three, weeks thirteen to twenty-four, is the conversion. Integrate the AI Act checks into the delivery pipeline as policy as code. Launch communities of practice for data science and for prompt engineering. Move the pilots into production and open the second wave. The stated outputs are the communities and control automation, version one.

Note that version number, because it is doing real work. The design does not promise finished controls at week twenty-four. It promises a first release of them, which means the manual gate does not vanish at week twenty-four either. It narrows to the cases the pipeline cannot yet decide, and the interesting metric from that week onward is how fast that residue shrinks.

The outcome

Nothing in this piece was run. What exists is a dated plan with named outputs, a named dependency and a board decision about whether to fund it. The week twelve platform milestone is a target, the week twenty-four control automation is a target, and the pilot counts are allocations. Read any of them as an achievement and you have invented a result.

What does generalise is the sequencing logic, and it carries a price the plan is right to make visible. A human gate has a throughput ceiling. Two to three pilots in the baseline phase reads less like modesty than like a capacity number, since a manual review load has to be carried alongside the reviewers' day jobs. Wave two waits for week thirteen, which is also the week the gate is designed to stop being the narrowest point in the system. Read that way, the pilot count is a capacity calculation wearing the clothes of an ambition statement.

The second price is variance. The same gate applied twice by two reviewers produces two answers, and the difference between them is invisible unless someone records the reasoning. That is precisely what an evidence file has to survive later, when a supervisor asks why one model passed and a similar one did not.

Which suggests the one thing worth doing differently. The manual phase should be instrumented from its first day even though the judgement inside it is human: log the submission, the evidence attached, the reviewer, the decision and the stated reason. Do that and the twelve manual weeks stop being dead time before the automation and become the specification for it, with a labelled record of what a pass actually looked like. That is the design our Agentic OS work starts from, because graded autonomy is only meaningful when someone has written down what the ungraded version decided and why. An evaluation harness needs a corpus of judgements, and the manual phase is the only place that corpus gets made.

The honest part of this plan is not the endpoint. Plenty of decks promise controls in a pipeline. The honest part is the twelve weeks in the middle, named as an activity, owned by people, with a date on which they are expected to stop being the only thing standing between a model and a customer.

NEXT STEP

Ready to make AI real?