Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking operations

Case study
Banking operationsA consumer and SME banking group operating across several European markets

The agent governance session was booked before the bank had an agent worth governing, which is the only time that session is cheap

A consumer and SME banking group operating across several European markets designed its executive AI governance programme as eleven decision sessions, roughly fourteen hours of room time, each one specified as a duration, three named activities, a required input and a single output artefact. Agentic systems were given their own ninety-minute block, separate from the ninety minutes spent classifying systems against EU AI Act categories. That block maps autonomy levels, designs kill-switch mechanisms and settles whether oversight sits in the loop or on the loop, and it is sequenced after the session that charters the governance committee, so a kill switch has a named holder and a working escalation path before anyone specifies the button. Executives arrive having read a primer, because ten to twelve hours of mandatory pre-work displaces the levelling-up out of the room. What this snapshot records is a designed and booked programme, not a delivered control: pre-work stood at sixty per cent complete and the execution meter at zero.

90 minutesExecutive time booked to map autonomy levels and design agent kill switches
Client
A consumer and SME banking group operating across several European markets
Duration
Executive governance programme design, pre-workshop snapshot
AI · RIDGE E77.3 N81.3ρmax 1.00
14 hoursTotal room time across eleven decision sessions
10-12 hoursMandatory pre-work per executive, so room time is decision time
0%Execution status at the snapshot: designed and booked, not yet run

An executive committee can approve an agentic pilot in twenty minutes. What it cannot do in twenty minutes, or at any point after the thing is running, is agree on who is allowed to stop it. The approval is easy because the system is small and the demonstration is charming. The disagreement is expensive because by the time it surfaces, the agent sits inside a process, the process sits under a customer promise, and switching it off costs something that someone has to own on the spot.

A consumer and SME banking group operating across several European markets built its executive AI governance programme as a series of decision sessions rather than a briefing. Eleven of them, roughly fourteen hours of room time. One block, ninety minutes, is held for a single question: how do you govern a system that acts without a person approving each step. It was written into the agenda by a programme that had not yet picked its first three pilots.

The challenge

Most AI governance programmes have one risk session. Systems get sorted into categories, controls get attached, a heatmap comes out, and the committee moves on feeling the subject has been covered. That is a reasonable shape for a model that scores something and hands the score to a person, because the person is the control and everyone in the room understands what a person is.

This programme does carry that session. Ninety minutes, mapping the bank's AI systems against the EU AI Act's own categories, separating the higher-risk uses from the more limited ones, scoring impact and deciding which controls are mandatory. The Act is law and that classification is not optional work.

The design decision worth copying is that a second ninety-minute block was written for agentic systems, and kept separate. A classification exercise asks what could go wrong and which control reduces it. An agentic session asks something the classification grid has no column for: at what point does this system stop asking permission, what happens when someone takes that permission away, and whose hand is on the mechanism at three in the morning on a Sunday.

Fold that into the general risk hour and it dissolves into a bullet under mitigation controls, phrased as human oversight, agreed by everyone because nobody disagrees with human oversight in the abstract. The disagreement appears only when you ask which human, over which decisions, with what latency, and whether the system keeps running while the escalation is in flight.

The approach

Every session is specified the same way: a fixed duration, exactly three named activities, one required input the participants must bring, and one named artefact the session must produce. Nothing is on the agenda that does not end in a document, and each of the eleven blocks carries its input and its output beside it in the plan.

The agentic block follows that pattern. Its required input is a primer on agents, circulated in advance. Its three activities are autonomy level mapping, kill-switch mechanism design, and an explicit oversight decision between keeping a person in the loop and keeping a person on the loop. Its output is a control plan.

Take those in order, because the order is the argument.

Autonomy level mapping exists to kill the binary. Executives arrive with two categories, the tool that suggests and the system that acts, and almost every real deployment sits between them. Mapping the levels forces the committee to say out loud where the line falls for each candidate use: which steps an agent may take unreviewed, which it may take and then report, which it may propose and not execute. That map is what makes the next two activities answerable at all.

Kill-switch mechanism design is three questions wearing one name. What does the switch stop, the single run or the whole capability. At what granularity, one customer's case or every case in flight. And what happens to work already half done when it fires, because an agent halted mid-sequence leaves a state someone has to reconcile. A committee that has not answered the third question has designed a button, not a control.

The oversight choice decides staffing and cost. A person in the loop approves each consequential step, and the throughput of the system becomes the throughput of that person. A person on the loop watches a stream and intervenes, which scales, but only if the intervention is fast enough to matter and the watcher can see what the agent is actually doing. That is an instrumentation requirement disguised as a governance choice, which is why the session output is a control plan rather than a policy.

Now the sequencing, which is the part a reader can copy without any of the rest. The agentic block does not come first. It sits after the session that charters the governance committee, and that session settles the committee's mandate, the difference between voting and advisory membership, the escalation paths, the meeting cadence and the quorum. All of it inside its own ninety minutes.

Put those two blocks the other way round and the kill-switch conversation has nowhere to land. You specify a mechanism, then discover there is no agreed holder, no escalation path for when the holder is unreachable, and no quorum rule for switching something back on. In this order, the mechanism is designed into an authority structure that already exists on paper, and who may pull it stops being a philosophical question and becomes a line in a charter the same room has already drafted.

The primer matters for the same practical reason. This committee is not asked to learn what an agent loop is during the ninety minutes it is meant to be deciding something. Pre-work across the programme runs to ten or twelve hours per participant, capped item by item: a reading pack of three documents, a ten-question readiness self-assessment built with the bank's own team so it interrogates the projects that actually exist, and a one-page written position. Materials go out a week ahead, with a minimum lead of forty-eight hours and a deadline twenty-four hours before the session. Completion is tracked daily, and at this snapshot it read sixty per cent.

The outcome

Be precise about what this is. The snapshot records a designed and booked programme, not a control in production. The execution meter on the follow-through plan reads zero per cent, labelled as awaiting completion of the workshop series. No agent is under governance yet, no incident stopped, nothing measured. What exists is an agenda with an artefact chain, a specification for a control plan, and ninety minutes of scarce executive attention held open for it.

The follow-through carries the same discipline as the pre-work. Decisions, action items and a sponsor debrief inside forty-eight hours of the room closing. The final report and a detailed action plan in the first two weeks, alongside signed personal commitments from each executive. And by the end of the second week, the thirty, sixty and ninety day reviews in calendars with their agendas already written, because a review booked later is a review that gets moved.

The general point is about price, not about banking. A kill-switch conversation held before the capability lands costs ninety minutes and produces a document. Held after an agent is embedded in a live process, it costs an incident, a regulator's question, or a standoff between two functions who each believe the other has the authority. The Platform work we do runs into this constantly: the instrumentation that makes on-the-loop oversight real has to be designed into the agent harness rather than bolted to it, long before a governance committee is asked about it. Our Consult engagements now open by asking who holds the stop for each automated process already running. The answer is usually a name nobody has told.

An organisation that books this session while the topic is still theoretical gets to be reasonable about it. An organisation that books it after gets to be fast, in a room where somebody is already angry.

NEXT STEP

Ready to make AI real?