Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsFinance

Copilots Need a Place to Put Their Answer

RealAISep 10, 20247 min read
FinanceRisk & ComplianceData Strategy

Ask a European bank this autumn what it is doing with generative AI and you will hear about a copilot. Something that sits beside an advisor, reads the client file, drafts the recommendation, pulls the right disclosure. The demos are good. Retrieval augmented generation over the product documentation works well enough that the room goes quiet, and the evaluation set says hallucination is down to a level somebody will defend upstairs.

Then somebody asks the question that ends the meeting. The model has produced a portfolio recommendation. Where does it go, and by whose authority is it allowed to say that?

Most programmes have no answer, because the copilot was built as a model problem. The recommendation lives in a chat window with no relationship to what this institution permits, which securities this advisor is qualified to propose, or how far a proposal may drift from the house allocation. The copilot has an opinion and nowhere to put it.

The layer nobody demos

Some years ago I spent time on a programme review for a large European retail banking group. Branch advisors across its local entities used an investment advisory tool. Behind it sat a second, much less glamorous application whose only job was to configure the first. Its specification runs to fifty-one pages of screen-by-screen description, dated a decade ago, written more than three years before MiFID II took effect and long before anyone in that building had said the word transformer out loud.

Read it now and it is startling how much of it is the thing the copilot conversation is missing. Not the model. The place the model's answer has to land, expressed as data a business person maintains and the advisory application reads at runtime. It is the layer we now spend most of our time on inside RealAI Platform engagements, and the one almost nobody puts in a slide.

A boundary ran through that layer. Twelve general settings, each with its own page, held what the group would not let local entities own, among them sectors, regions, security groups, duration base rates, risk types, risk brochures, the disclaimer, the regulatory text modules and the justification of a recommendation. Everything else, the product shelves and advisor roles and allocations and the local voice in the client report, belonged to the entity itself.

Almost no copilot programme can draw that line for itself. Which prompts, thresholds, reference tables and text blocks may a local team change, and which are group property nobody may fork? The bank answered on one page.

Deny by default, in one sentence

The best line in the document is seven words long. For each security group, a local entity may define how far a recommendation is allowed to differ from the optimum of the target allocation. Then the caveat: if not set, no deviation is allowed.

An unconfigured tolerance could have meant unlimited, which is what most systems quietly mean. It could have meant use a sensible default, which is worse because it looks deliberate. This spec chose zero. Nothing configured means nothing permitted.

Guardrail registries are always incomplete, because people enumerate the cases they thought of. If the unenumerated case resolves to permitted, the system is at its most free in exactly the space nobody reasoned about. A copilot proposing a portfolio shift should be bounded by a tolerance band held as data, per asset class, and should refuse where no band exists.

Note where those bands lived. Not in engine code, but in a business-maintained store with authorisation and an activation step around it, so risk appetite could move at business speed and still leave a trail. That is the mechanism that makes an assistant's behaviour changeable without a release.

Some documents may only be added

Document handling in that system encodes a control most content platforms still lack.

Documents for a given security assemble in three steps rather than resolving by override: everything from the base system, plus everything from the group entity, plus everything from the local entity whose configuration this one borrows. Union, not precedence. The spec uses precedence elsewhere, for security allocation data, and the asymmetry is deliberate. A missing disclosure is a regulatory failure. A duplicate one is clutter.

Then the rule that matters. Legally required documents cannot be changed, cannot be deleted, and cannot be marked as not used. They can only be added when missing, and they occur once per type. One permitted operation on that class of artefact, and it is insertion.

Write that on the wall of any team pointing a language model at client-facing material. A model asked to tidy up the disclosures, or to summarise the risk brochure so the client actually reads it, will do exactly that, fluently, and nobody will notice which sentence left. The prohibition belongs in the data layer, because a prompt is a request and a schema is a rule.

The last piece I would steal outright. Every document carried a category of usage: it can or must be used, it is handed over to the client, it is offered to the client, or it is not used at all. That is not metadata about what a document says but a machine-readable instruction about what the institution owes the client. An assistant drafting a client pack need not reason about disclosure obligations if the disposition is a field.

Local entities could still intervene, for three written-down reasons: the source quality was unsatisfactory, the entity wanted its own content instead of the standard, or it did not want the document at all. Central content teams usually treat local substitution as non-compliance. This spec treated it as a requirement and gave each motive its own mechanics. You can only automate a decision somebody once bothered to write down.

Bound the taxonomy, and date the numbers

Two more habits, one to copy and one to avoid.

Copy the cap. A local entity could define up to ten filter criteria, each with a name and an identifier used by the host system that classifies the securities. Ten, not unlimited. Give a multi-entity business an open-ended classification space and within three years nothing is comparable across entities and no shared list can be mapped from one to another. The same holds for the tag space behind any retrieval layer being stood up this year: an unbounded vocabulary is an unqueryable one, which is why so many pilots return documents that are topically adjacent and operationally wrong.

Avoid the staleness. The quantitative model sat in that admin tool as ordinary configuration. A security group carried its asset class, its region, its display colour for the pie chart, an expected return, an expected risk, whether the issuer was corporate, and whether the paper was investment grade. Correlations between security groups were typed in by hand, in a range the spec gives as minus one to one. Duration base rates were kept for a fixed list of eighteen currencies.

Those parameters were at least visible, approval-gated, reconstructable as they stood on any chosen past day, and attributable to the person who last changed them and the date they did it, which is more governance than most machine-learning features get. The gap is that nothing states how often any of it should be re-examined. Security data elsewhere got an explicit re-check reminder. The correlation matrix, the most volatile numbers in the model, got none.

A model reading configuration cannot tell a considered parameter from a forgotten one, and will produce an equally confident sentence either way. Every input a generated recommendation depends on needs a freshness contract and an expiry, and the system should decline to produce the output when an input is past it. Refusal is a feature, and we still mostly measure assistants on whether they answer rather than on whether they correctly decline.

Where the model earns its inference cost

None of this argues against copilots. It argues about sequence, and about where the money goes.

That same specification contains the shape of the useful assistant. An information panel reachable from every page, with sections each local entity fills itself and a static section of forms, is a retrieval problem with the governance already solved. The client report was split into two zones with different owners: the local entity wrote its philosophy, its presentation of itself and a preamble, while the disclaimer, the regulatory text modules and the justification wording sat centrally. A generative draft can be given the narrative zone and structurally denied the regulated one. That is not a prompt instruction but a boundary already in the data model.

The economics point the same way. Inference is no longer free and per-call cost is a line item people argue about. Spending tokens to have a model re-derive what a settings table already states is an expensive route to a worse answer.

Zero
Permitted deviation where no tolerance band is configured
Add only
The single permitted operation on a legally required document
10
Filter criteria a local entity may define, deliberately capped
3
Steps a document hierarchy accumulates, union not override

Where to start

Before the next copilot pilot, produce two artefacts.

The first is the boundary page. One list of what is group property and may not be forked locally, one list of what a local team owns. The banking group put the risk parameters and the regulator's words on the central side, and the product shelves, roles and local voice on the other. Yours will differ. The point is that it exists and fits on a page.

The second is the settings inventory: every threshold, tolerance, reference table, text block and disposition field an advisory decision depends on, with an owner, an approval path, a last-changed date and a freshness expectation. That is a review exercise, not a build. It is the first thing a RealAI Consult engagement produces, and it tells you how much of your advice is governed and how much is branch habit.

Then build the assistant on top and hold it to a plain test. It may read the settings layer. It may propose changes through the same approval path a human uses. It may never quietly route around it.

The bank that wrote that specification was not thinking about language models. It was thinking about what happens when a whole branch network gives advice under one licence. Same problem.

The most valuable document I read on that programme had no model in it and a screenshot on nearly every page. It described where an institution keeps its opinions. This year's copilot conversation is largely about how to generate an answer. The banks that get further will be the ones that first decide where the answer is allowed to go.

A copilot that cannot read the institution's own settings is not advising. It is guessing fluently, in the house style.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?