Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking operations

Case study
Banking operationsA consumer and SME banking group operating across several European markets

Mandatory AI training capped at two hours, and mostly spent on refusals

An enablement design for a consumer and SME banking group operating across several European markets. Against the instinct to build a long academy for everyone, the mandatory frontline curriculum was capped at two hours across five modules, weighted towards refusals and detection: what not to paste into a prompt, how to spot a hallucination or an AI-written phishing attempt, and when to stop and call a manager. Capability came fifth. The stated targets, due to be measured two quarters after rollout, are 85% weekly adoption of approved tools and a 15% reduction in errors and compliance breaches.

2 hoursTotal mandatory frontline curriculum
Client
A consumer and SME banking group operating across several European markets
Duration
2 hours, five modules
AI · RIDGE E59.1 N56.3ρmax 1.00
5Modules, one sitting
85%Weekly adoption target, two quarters on
-15%Error and breach reduction target

A consumer and SME banking group operating across several European markets had to decide how much AI training to make compulsory for the people who sit in front of customers. The answer, set out in the enablement design that followed, is two hours. Five modules, one sitting, mandatory. Four of the five are spent on things not to do.

The challenge

The group was standing up a full AI academy with several audiences at once: a board and executive track, technical tracks for data and machine-learning practitioners, a platform and MLOps track, an enablement track for product owners and business analysts, and a literacy layer for everyone else. Each of the specialist tracks could justify weeks. The literacy layer could not.

Frontline staff are the most exposed population in a retail and SME bank. They handle customer personal data, they field the disputes that follow an automated decision, and they are the ones a well-written phishing message reaches. They also have the least discretionary time. A compulsory curriculum long enough to teach real capability would not be completed; a curriculum short enough to be completed would not, on the usual assumptions, change anything.

The design had a second constraint. The same programme carries the group's regulatory work, with EU AI Act compliance sitting inside the responsible-AI track and control-evidence collection written into the product owner and business analyst track as regulatory mapping and audit trail. Literacy for the wider workforce had to produce something a regulator would accept as evidence, not a completion certificate.

The approach

The curriculum was capped at two hours and the five module slots were spent in a deliberate order.

Module 1, safe copilot use, covers interacting with AI assistants without over-reliance, understanding their limitations, and human-in-the-loop basics. Module 2, privacy and data handling, is explicit about what not to paste into prompts, along with handling personal data, redacting customer information, GDPR compliance and a zero-retention policy. Module 3, spotting red flags, teaches identification of hallucinations, bias in outputs and AI-enabled phishing attempts, with verification technique as the through-line. Module 4, escalation protocols, sets out when to stop and call a human manager and how to handle customer disputes involving an AI decision, and hands the staff member a decision tree.

Only module 5 teaches capability. It is prompt libraries for common tasks such as summarising meeting notes and drafting emails, and it carries the design's only productivity number: a target of four hours a week saved.

The ordering is the argument. Modules 1 through 4 are restraint, detection and escalation, and only the fifth adds capability. Module 3 puts hallucination spotting and AI-written phishing in the same lesson, which treats both as one skill: verifying output that sounds plausible. That turns the literacy layer into a security control with a number attached rather than a soft-skills exercise.

The goal isn't to turn everyone into data scientists, but to make everyone AI-safe and AI-efficient.

Internal AI centre of excellence · Consumer and SME banking group

The two hours sit inside a wider delivery architecture built on the 70-20-10 model with five channels matched to audience. Executives and leaders get one to two day in-person workshops for strategic alignment and governance. Data scientists and engineers get cloud-based sandboxes for practising MLOps safely. The frontline and operations population gets self-paced micro-learning, on-demand video and quizzes, delivered through the learning system and on mobile. Senior engineers mentor juniors through paired programming on agentic patterns. Cross-functional case days give teams 24 to 48 hours on a real business problem.

Certification is tiered across three levels and mapped to job families. The foundational tier covers all employees with an internal AI-awareness certificate on safe usage, data privacy and phishing awareness, plus an ethics basics module. The applied tier covers analysts, product owners and managers, pairing an external cloud-provider fundamentals credential with internal prompt engineering for banking. The expert tier covers data scientists and MLOps staff and cannot be passed with coursework: alongside an external AI engineering credential and an internal responsible-AI qualification, it requires a capstone delivering one production-grade AI agent. Every badge expires. Renewal is mandatory every 18 to 24 months, capped at 24 months maximum validity, and it is tied to keeping access rights.

The outcome

What exists is a specified curriculum and a target set, not a measured result. The design states two behavioural targets to be read two quarters after rollout: an adoption score of 85%, defined as the percentage of staff actively using approved AI tools weekly, and a 15% reduction in manual data entry errors and compliance breaches. At academy level the measurement plan has three layers on a single dashboard: knowledge lift of plus 25% against a pre-training baseline, a training experience net promoter score of 55 or better, and a behaviour test at 90 days where the target is 80% of participants still using the new tools weekly.

85%
Weekly adoption target, two quarters on
-15%
Errors and breaches target
80%
90-day behaviour target

The 90-day layer is the one that decides whether the two hours worked, because it measures retention rather than attendance.

The curriculum has a visible seam, and seeing it needs no hindsight, because both sides of it are in the same programme. All five modules assume a human is typing the prompt and reading the answer. That is what a copilot is. The group's own deeper tracks assume something else: agent orchestration, agentic workflows running on cluster infrastructure with drift detection, paired programming as the transfer mechanism for agentic patterns, and a top-tier capstone whose qualifying evidence is one agent in production. The specialist half of the academy is building systems that act without a person in the loop for every step, while the mandatory half trains people to supervise a system that waits for them.

Each of the four non-capability modules changes shape under that shift. Human-in-the-loop basics stop meaning read the answer before you send it and start meaning approve a plan before it runs, then set the bounds it runs inside. What not to paste into a prompt becomes what an agent is permitted to read, which is a tool-permission and data-scope question settled at configuration time rather than by a person's judgement at the keyboard. Verification stops being a check on one visible answer and becomes a check on a trace: which tools ran, in what order, on whose data. Phishing detection moves in the same direction module 3 already named, because the message that has to be distrusted is no longer only the one a person opens; it is the one an agent reads and acts on before anybody sees it. The escalation decision tree stops being advice a staff member follows and becomes an interrupt path the agent itself has to offer, with a defined hand-back point and a response time.

The adoption target is the metric most exposed. Weekly active use of an approved tool is measurable while people open the tool. When agents run in the background of a servicing queue, usage becomes ambient and 85% stops meaning much. The two numbers that survive are the ones about consequences: the 15% reduction in errors and compliance breaches, and the 90-day behaviour check. The product owner and BA track already points at this by defining AI KPIs as time-to-value, cost-per-query and user trust scores instead of accuracy, and by requiring three artefacts per AI product: a requirements document written around data sources and model behaviour, an evaluation test plan with golden datasets and adversarial cases, and an operational playbook holding the human-in-the-loop procedures and escalation paths.

There is a second seam, and it is wider. All five modules are written for work that begins at a keyboard, and a consumer and SME frontline runs a large queue that is never typed: account opening packs, identity papers, signed mandates, proof-of-address scans, the dispute evidence a customer photographs on a phone. Wherever that queue is machine-read, it reaches the staff member already converted into fields, and its failure mode is not the one module 3 drills. A hallucination arrives as a sentence that reads well and says something untrue, and a person can be trained to distrust it. A misread field on a scanned mandate arrives already looking like data, sitting in a form beside fields that are correct. The through-line module 3 already names, verify the plausible, carries over intact. What has to change is the evidence the staff member is given: the confidence on the extracted field, the crop it came from, and a way to send it back.

The direction the specialist tracks point at is not a better assistant but a chain of them: intake reading the document, an agent pulling case history out of the servicing systems, another drafting the customer reply, a checking agent holding that draft against policy, and a person reached only where the chain stops. Once servicing runs on interconnected pipelines like that, agents carry a working share of how the bank operates, and the frontline job stops being use of a tool. It becomes supervision of something autonomous between the points where it asks. Reading a trace, judging where the chain halted and knowing what each step was permitted to touch is a different skill from writing a good prompt, and it is the skill the mandatory two hours will eventually have to teach.

The badge expiry is the design's most durable decision. Tying certification validity to access rights, and capping it at 24 months, concedes that a permanent AI credential certifies a version of the tooling that will not exist by the time it is claimed. The two-hour ceiling holds for the same reason. A short curriculum built on refusals, detection and escalation is cheap enough to rewrite when the systems being supervised change from ones that answer to ones that act.

The rewrite is where this gets faster than the design assumes, because it does not have to be an instructional-design project at all. The loop and harness engineering that puts an agent into production already produces the teaching material as a by-product. The loop leaves traces of what the pipelines actually did, which refusals they actually raised, which cases actually reached a human and how long the hand-back took. The harness is the thing that scores those runs against the policy they were meant to hold, and it accepts adversarial cases just as readily as the evaluation test plan the product owner track already requires. Point that harness at the frontline instead of at the model and the modules write themselves from real refusals, real escalations and real misread fields, current by construction rather than current as of the last instructional-design cycle. The group has budgeted two hours of a frontline worker's time. What decides whether those hours are worth anything is how quickly the content behind them can be made to match the systems that exist by the morning the worker sits down.

NEXT STEP

Ready to make AI real?