A credit portfolio has a shape nobody sits down and designs. The largest exposures get a named officer, a file review on a fixed cycle and a place on the watchlist. Everything underneath gets an annual renewal and a payment history. The line between the two is not drawn by risk. It is drawn by how many hours a credit team has, and it moves whenever somebody resigns.
A consumer and SME banking group operating across several European markets wrote that line down. Portfolio reviews ran quarterly. Deep monitoring reached the top twenty percent of exposures. The reason given on the page for the twenty percent is analyst capacity, stated plainly and without embarrassment. The remaining four fifths of the book were monitored in the weak sense that a missed payment eventually appears somewhere.
The challenge
Three problems are named, and they are usually treated as separate complaints from separate people.
The first is cadence. A quarterly review is a lagging indicator, and the deliverable says so in those words: often too late to prevent default. A borrower's receipts thin out over weeks. A review arriving once a quarter reads the deterioration after it has finished happening, at which point the available responses have narrowed from restructuring to recovery.
The second is coverage. Twenty percent is not a risk judgement, it is a capacity budget wearing a risk judgement's clothes. Concentration is a real argument for watching large exposures closely, but it is not an argument for watching the rest never.
The third is what the scarce capacity is actually spent on. Analysts spread financials by hand out of covenant documents. The hours that decide how much of the book can be watched are consumed by transcription.
Put in that order the three stop being separate. Transcription consumes capacity, capacity sets coverage, and the quarterly cadence is simply what the remaining capacity can afford. The binding constraint on early warning at this bank was clerical, which is an uncomfortable thing to find under a risk function and a very cheap thing to attack.
The approach
The workflow has four steps and the first one is the whole argument.
Continuous ingestion runs across transaction flows, covenant documents and macro and sector news. Anomaly detection sits on top of it, checking covenant breach, cashflow deterioration and negative sentiment. Where something trips, the agent drafts a risk memo that cites the specific transaction drop or covenant miss that triggered it. A risk officer then decides between three options: watchlist, restructure, or dismiss.
Notice what the design does not ask for. The agent does not restructure anything, does not reprice, and does not move a name onto the watchlist by itself. It sits on the augmented rung of the deliverable's own graded autonomy ladder, where the machine assembles evidence and a human decides. The deliverable classifies the system as high risk under the EU AI Act on the grounds that it informs credit decisions, and attaches three controls to it: a defined human review threshold, full audit logging, and bias testing.
That combination is the reason coverage is the right first move rather than an ambitious one. Extending monitoring to the whole book is a read-only act. The agent's output is a memo with citations, and a memo can be wrong in front of a person whose job is to disagree with it. Compare that with the autonomy cases elsewhere in the same deliverable, where acting without a human requires exposure caps, confidence bands and a daily sampled audit before anyone will sign it. Reading everything is the cheapest capability in the pack to govern and the one with the largest untouched population underneath it.
The delivery method around it is sequenced the same defensive way. Control mapping happens before the pilot launches rather than before production, so the pilot runs inside the controls it will eventually need. Pilots run in shadow mode alongside the existing process. Live experiments carry an automatic abort armed on two thresholds, an error rate above 0.5 percent or latency above two seconds. Tuning happens weekly at working-group level, and scaling decisions happen at a separate quarterly gate with exactly two criteria: value above cost, and risk inside appetite. Our Agentic OS work is built on that same split, because a loop that reads a whole loan book daily is only as defensible as the record of what it read and under whose identity it read it.
The outcome
What exists is a design. The three figures sit on the slide as bare tiles with nothing marking them measured or intended, and nothing has been run, so every one of them is what the design commits to rather than what it achieved.
They differ in how much weight they can carry. Full daily coverage against a top-twenty-percent sample is the most defensible of the three, because it is a property of the architecture rather than a prediction about behaviour: a pipeline either ingests the whole book every day or it does not. Two months of earlier detection is a projection against the quarterly cadence, and the page carries no measured detection lag to project from. The fifteen percent reduction in non-performing loans through pre-emptive restructuring is the weakest, and honestly so: no baseline non-performing loan rate is printed beside it, and the deliverable's benefits tracking carries non-performing loans only as an unlabelled movement from a baseline bar to a target bar, with neither rate written down. That makes the fifteen percent a ratio against an unstated base. Read the slide cold and you cannot tell whether that fifteen percent was estimated from the bank's own loss history or chosen as an ambition. It is ambiguous between the two, and the ambiguity is worth stating rather than resolving in the flattering direction.
So the first job of any build here is not the agent. It is two baselines. Detection lag, measured as the interval between the first observable deterioration signal in the data and the first internal note about it, has to be instrumented on the existing process before anything new touches the book. Non-performing loan formation has to be split by whether the exposure was inside the watched fifth or outside it. With those two numbers in hand, shadow mode turns a projection into a comparison, because the agent is reading the same book the officers are and the two records can be laid against each other.
One number the design does not carry is the one a build will meet first. Daily monitoring across the whole portfolio will generate alerts on four fifths of a book that previously generated almost none, and every flagged case still lands in front of a risk officer under the three-way decision. If the expected flag rate is not forecast and bounded before launch, the capacity constraint that produced the twenty percent in the first place reappears at the review desk within a month, wearing a different name. The measurement design in the pack is good enough to catch that early. Our Consult work starts there, on the arithmetic of what an expanded queue costs, because that is the cheapest place to find out whether coverage pays.
The reason coverage beats cleverness is not sentiment about small borrowers. On the top fifth, a better model competes with attention that already exists: a named officer already reads that file. On the other eighty percent, it competes with nothing. A mediocre signal on an exposure nobody was watching is worth more than an excellent signal on one somebody already reads every quarter, and no amount of model work changes which of those two populations is larger.
