Almost every AI governance deck in banking contains the phrase "human in the loop," and almost none of them say what would have to be true for that human to count as a control. A person is named. A step is drawn. A box appears in the workflow with an approval stamp in it, and the box is treated as the answer to the question a supervisor is going to ask. It is not an answer. It is a promise that somebody will look, with nothing behind it that would tell you whether they did.
A consumer and SME banking group operating across several European markets designed its target-state credit workflow around one line: "AI proposes, Human disposes." The agent executes the sub-tasks and drafts the decision. The human expert keeps full authority to approve, reject or modify before anything becomes final. That much is conventional. What is not conventional is that the design attaches two numbers to the approval step, and both of them are about the human rather than the model.
The challenge
The starting point was honest about how little was automated. Roughly 65 percent of back-office processes and roughly 80 percent of risk reporting were still fully manual, with technology producing dashboards and no recommendations. Underwriters were reading PDF statements and retyping the figures into risk tools. In financial crime, more than 95 percent of alerts were non-fraudulent, and working a single one meant an analyst opening four or more systems by hand.
That baseline creates an obvious pull toward automation and an equally obvious governance answer: keep a person in front of every output. The rung below the target state in this design does exactly that. Its guardrail is 100 percent human review, the AI has read-only access to core systems, it cannot execute a transaction or modify customer data, and every prompt and response is logged for audit. As a way to cap blast radius, that is sound. As a control on the human, it has a defect that only shows up later. The 100 percent is a coverage requirement rather than a measurement, and the logs capture what the model said, not whether anybody read it. "The user must verify the output" is a policy about intent. It emits no figure of its own, so it cannot fail an audit and cannot pass one either.
The EU AI Act, in force over this design, is specific that high-risk systems carry human oversight obligations, and the governance model designed alongside this workflow puts the technical file, the data quality and the execution of human oversight on the business owner in the first line of defence, with conformity assessment and bias testing sitting in the second. Read those two together and the gap is plain. The second line is being asked to challenge the effectiveness of a control that the first line has not instrumented. Oversight that leaves no trace is not something a challenge function can test. It is something they have to take on trust, which is the arrangement internal audit exists to prevent.
The approach
The target-state design routes work through four stages: data input, agent analysis producing a score and its reasoning, a human gate that verifies and approves, and a final decision. Two workflows carry it. In SME risk scoring, the agent aggregates bank statements and bureau data, calculates a probability of default and drafts the credit memo, leaving the underwriter to weigh the qualitative material, the strength of a business plan among it, and to validate the score. In KYC file preparation, the agent retrieves beneficial ownership data, screens against sanctions lists and packages the file, leaving the compliance officer to adjudicate the high-probability matches and clear the false positives.
Three prerequisites are named ahead of any of it: structured data clean enough to score on, explainability sufficient to say why a recommendation was made, and a feedback loop so the system learns from overrides. That third one is the hinge. An override is the most informative event the workflow produces, because it is the only moment where a human expert has looked at a specific model output and said no with a reason. Most deployments throw that away. The gate records a decision and the disagreement evaporates into a case file.
Instrumenting it changes what the gate is. The design names the override rate as the metric for the approval gate, with a target of staying below 20 percent, and reads a rate inside that band as evidence that the people using the system trust it. Beside that sits a fairness measure, a disparate impact ratio computed on the decisions flowing through, printed with a value and marked compliant. Two numbers, pointing in different directions on purpose. One bounds how often the human disagrees with the machine. The other bounds who the agreement is falling on.
They have to be read together, and this is the part worth carrying out of banking into any regulated workflow. An override ceiling on its own can be satisfied the cheap way. A rate that drops toward zero is exactly what trust looks like and exactly what rubber-stamping looks like, and the two are indistinguishable from the counter alone. Volume per reviewer, time spent per case and the fairness ratio on the passed flow are what separate them. The design gets the pairing right. What it does not do is state the floor the fairness ratio has to clear, which leaves a value on a slide with no test behind it, and a threshold nobody has written down is a threshold nobody can breach.
The outcome
What this engagement produced is a design, and every figure in it should be read that way. Nothing had been run. The 80 percent reduction in time-to-yes is a projection against the manual baseline, not a measurement of a deployed workflow. The productivity figure attached to the underwriting case, an expected 25 percent lift in analyst capacity, is the same kind of number. The fairness value is a pass mark the design expects to hold rather than an audited result. And the override figure printed beside the target is ambiguous on the face of the slide between an expected operating point and an illustration of what a compliant rate looks like. Since no cases had passed the gate, it cannot be a measurement, and calling it one would be inventing a client outcome.
The unambiguous commitment is the shape, not the values: this gate has a definition of working, written down before anyone built it. That is rarer than it sounds and it is the whole reason this deliverable is worth publishing.
The design also makes the next rung coherent, which unmeasured oversight never does. Once the workflow moves from approving every case to monitoring exceptions, per-case review is replaced by a random sample of 10 percent of autonomous decisions with a kill switch behind it, and eligibility is bounded by hard exposure conditions rather than by confidence alone. That handover is only defensible if the gate produced a record. An override history is the evidence that says which slices of the flow the model has earned autonomy on and which it has not. Without it, the move to lighter oversight is a decision taken on optimism.
This is where the work stops being a slide and becomes engineering, and it is the work our Agentic OS is built to carry: the override taxonomy that distinguishes a correction from a policy disagreement, the fairness computation running on the same cadence as the queue rather than quarterly, and the feedback path that turns disagreement into a training signal instead of an audit note. None of that is a model problem. All of it is loop and harness design, and it is what decides whether the human in your workflow is a control or a courtesy.
