Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesRetail banking

Case study
Retail bankingA consumer and SME banking group operating across several European markets

Twenty-five cells, four of them strong, and a regulatory ceiling already scheduled

A review graded a consumer and SME banking group across five capability domains and five delivery units. Four of the twenty-five cells sit in the strong band, nine in the critical band, and three are marked not applicable. Prompt and agent engineering is empty in every unit. The same deliverable carries the deadline that makes the grid urgent: the already scheduled date when full EU AI Act obligations reach credit scoring, risk assessment and biometric identification. What was delivered is a design, not a result.

4 of 25Capability cells in the strong band
Client
A consumer and SME banking group operating across several European markets
Duration
Assessment and design, pre-build
AI · RIDGE E54.5 N62.5ρmax 1.00
0 of 5Units with prompt and agent engineering
20 + 5Analysts and risk officers to upskill
ScheduledDate full high-risk obligations bite

Capability behaves like a branch network. Capacity gets sited where the demand was when somebody drew the map, and it stays there after the traffic moves. A review of a consumer and SME banking group operating across several European markets drew exactly that map: five capability domains against five delivery units, twenty-five cells, each painted in one of three bands. Four cells came out strong. Nine came out developing. Nine came out critical. Three were marked not applicable.

Read on its own, that grid is a scorecard. Read against the deadline set out elsewhere in the same deliverable, it is a map of where the work now arrives and where nobody is standing.

The challenge

The grid crosses five domains, data engineering, machine learning, prompt and agent engineering, MLOps and platform, model risk and governance, with five units: the central centre of excellence acting as hub, retail banking, SME banking, risk and compliance, and IT and operations.

Data engineering is the group's oldest strength. Mature at the hub, strong in risk and compliance, mature in IT and operations, and only partial in the two customer-facing lines. Machine learning is thinner: hiring at the hub, pilots in SME banking, developing in risk and compliance, and an outright gap in both retail and IT and operations. Prompt and agent engineering is critical at the hub and none in all four other units, the only row where every cell sits in the critical band. MLOps and platform is building at the hub, planned in IT and operations, and marked not applicable in retail, in SME and in risk and compliance, none of which own a platform. Model risk and governance is strong in exactly one place, risk and compliance, while the hub rates itself only defining, retail and SME sit at awareness, and IT and operations has a process.

Two things fall out of that shape. The centre grades itself less mature on model risk than the second line it intends to govern with. And capability sits where a function already existed. Data engineering, machine learning, model risk: each of those had a department to inherit. The one domain with no predecessor department, prompt and agent engineering, is the one with nothing under it anywhere.

The deadline is on the regulatory track in the same pack. The Act entered into force first, six months later the prohibitions took effect and AI literacy for staff became a legal obligation, and six months after that the general-purpose model rules arrived. A further twelve months on, full obligations reach credit scoring, risk assessment and biometric identification, tagged in the pack as banking core impact, and that step is the one still to come. The readiness track runs ahead of it: a control framework built across three quarters, then a one-quarter pre-compliance window for a dry-run conformity assessment, a technical file audit and registration in the EU database.

The group's own posture is stated plainly: a deployer of high-risk systems in creditworthiness and biometrics. Creditworthiness evaluation for retail and SME sits on the high-risk list. So does remote biometric identification at digital onboarding. The two units that own those regulated decisions are the two graded partial on data engineering, gap and pilots on machine learning, not applicable on platform, and awareness on model risk.

That is the reason to stop reading the grid as colours and start reading it as distance.

Loading exhibit
Exhibit 1Readiness debt is a distance, not a colour.Five capability domains across five delivery units as a terrain surface: ground height is the grade recorded in the review, read onto a nought-to-four scale, so the strong corners rise and the empty ones crater. Each column measures the climb from that ground to the flat production plane at level 4, the level the deadline demands, coloured in the review's own three bands and pulsing where the grade is zero. Hollow ghost columns are the three cells the review marked not applicable, which is a different fact from a gap, so the ground under them is held level instead of cratered. Hover a column for one cell, or let the ceiling band step through a domain at a time.

The columns are not evenly spread. They are tallest under a single row, and that row is the one nobody in the group has built.

The approach

The talent answer splits three ways. Two strategic hires carry the core: a head of AI in Q1 and a lead MLOps engineer in Q2. Upskilling carries delivery: 20 data analysts moving into AI practitioner roles and 5 risk officers moving into model risk management. Partnering covers the rest, a platform provider for commodity infrastructure and external counsel for generative AI legal work. Twenty-five people re-pointed internally, two hired.

The money is split 40 percent to platform and governance, 40 percent to use case delivery, and 20 percent to upskilling and change. How that money reaches a team is left as three options rather than one decision, each tagged with the phase it suits: a pooled innovation fund for the first six months, value-based co-funding from month six, where the centre pays for governance and platform while the business pays for delivery resources, and usage-based chargeback held back for mature platforms in year two. Co-funding is the one recommended. Capital releases against three gates: risk tiering confirmed, value case signed by the business, data privacy cleared.

The operating machinery is where the design gets precise. The lifecycle runs seven stages, plan, build, validate, deploy, monitor, improve, retire, and every stage carries exactly one control and one metric. Business case approval against a projected return. Canary release and security scan against P99 latency. Drift alerts wired to automatic ticket creation, measured against accuracy versus baseline. Data purge and a final audit log against systems decommissioned. Only one stage is a hard gate: validation, where model risk sign-off and EU AI Act compliance meet.

Classification works the same way. A business owner submits a use case, logic gates propose a tier, legal and compliance verify only the high-risk ones, the system is logged in a central inventory with an identifier, and controls are assigned from the tier. Adoption is targeted numerically across three waves, 20 percent of staff, then 50, then 85.

The outcome

What was handed over is a design. A graded baseline, a talent plan with quarters and headcounts, a funding split with release gates, a lifecycle with named controls and metrics, a classification workflow, and an adoption curve with targets. The grades describe the week they were collected. The percentages are goals, not measurements. Nothing on those pages has yet been run.

The honest question now is not whether the design is right. It is how much of it still has to be done by people.

Start with the grid itself. Twenty-five cells graded through interviews and judgement, fixed to one date, stale within a month of delivery. A harness computes the same grid instead of surveying it. The artefacts already named in the lifecycle are the evidence: registry entries, eval runs, gate outcomes, drift tickets and how fast they close, purge logs at retirement. Point a scoring loop at those and each cell carries a grade with a source, refreshed weekly, on the same five-by-five frame. The capability review becomes a running instrument rather than an engagement.

Then look at the row with nothing in it. Prompt and agent engineering scored critical at the centre and none everywhere else, and that capability today is not writing prompts. It is building loops, an agent that plans, acts against real tools, and checks its own output against tests, and building the harness around the loop: fixtures, evals, a permission boundary and an audit trail that make it safe to run unattended. Staff that row and the not-applicable cells stop mattering, because the business lines never needed platforms of their own. They needed one harness with their data and their controls inside it.

The lifecycle is closer to a harness specification than to a diagram. One control and one metric per stage, one human gate. Write those controls as executable checks and the loop runs itself: plan, build and validate execute as a pipeline, the run stops at model risk sign-off, and the reviewer opens an assembled case with the evidence attached instead of a queue of tickets. The design already wires drift alerts to automatic ticket creation. The same wire, extended, has the alert trigger a re-evaluation, draft the retraining change, and post the result back to the gate for a human decision.

Compliance work is text and image work, which is why it moves fastest. The classification process already assumes partial automation: logic gates propose a tier and people verify only the high-risk cases. An agent reads the use case description, proposes the tier with its reasoning, drafts the record-keeping and transparency documentation that high-risk obligations demand, files the registry entry, and keeps the technical file current every time the model changes. The one computer-vision system on this group's high-risk list, remote biometric identification at digital onboarding, is also the one where post-market monitoring cannot be an annual exercise. Vision systems drift against cameras, lighting and document stock, so the evidence a dry-run conformity assessment needs is either emitted continuously by the pipeline that serves the model or assembled by hand in the one pre-compliance quarter left before the deadline.

The three pieces connect into one line, which is the part a branch network never had. An intake proposes a tier, the tier assigns controls, the controls become the checks in the lifecycle, the checks produce the evidence, the evidence grades the capability cell and prices the decision. Cost to serve stops being something finance surveys before it can invoice a business line and becomes a number the pipeline emits per decision, which is what the chargeback option in this same deck quietly depends on.

The capability map in this review is accurate and it is already out of date, because it was drawn by hand at a point in time. The build that follows it should be designed so the next version draws itself.

NEXT STEP

Ready to make AI real?