Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsFinance

Security Before the Data Move

RealAIApr 22, 20258 min read
BankingRisk and ComplianceData GovernanceAgentic SystemsMLOps

Every agentic pilot I am asked to look at this spring has a governance conversation attached to it, and in almost every case that conversation is happening after the data has already moved. The retrieval index is built. The connector to the core banking warehouse is live. Somebody has a notebook that pulls a year of transactions into a working store so the agent has something to reason over. Then risk gets involved, and the question on the table is what to remove.

That is a different question from the one worth asking, and the difference is entirely a matter of what order the steps were put in.

I was reminded of this by an artefact more than ten years old. It is a short proposition brief, an offering document written by an outside firm and filed with a large European cooperative banking group's material, from a point when the argument was still about whether an internal analytics capability was worth standing up at all. It claims no client results, because it is a sales document rather than a delivery record. It sets its own bar at a benefits case worth five times the investment, and it promises to run an insight project from hypothesis through to measured benefit inside eight weeks. Both of those are offers, not outcomes, and they should be read as offers.

What survives is not the pitch. It is one instruction sitting in the step where the data gets sourced.

Two ways to write the same rule

Take a single requirement, say that customer identifiers may not leave the system of record in a form that can be re-joined to a person, and write it at two different moments.

Written before collection, it changes the extract. The engineer who is asked for the data returns a view with the identifier already replaced by a surrogate key, because that is the only view that was ever requested. The pipeline that gets built has no other version to lose. The lineage is trivially clear, since there was only ever one path. The conversation with the system owner is short, because nothing is being taken away from anybody.

Written after collection, the same requirement becomes a remediation. There is now a working store, and a notebook that reads from it, and a model that was trained on what is in it, and a dashboard somebody in the business has started to rely on. The requirement has to be satisfied by deleting or masking, which means somebody has to establish what was copied, where each copy went, which downstream artefacts were derived from it, and whether the derived artefacts still carry the thing that was supposed to be removed. Every one of those is a discovery exercise before it is an engineering task.

I cannot give you a multiple for that difference from this source, because the source does not measure it. The brief states the ordering and moves on. What I can say from having watched both versions run is the direction, and the direction is not close. The cost of the second path is dominated by discovery, not by the technical work, and discovery cost rises with how long the pipeline has been useful to people.

There is also a property of data that makes the second path structurally worse. Data that has moved has been copied, and copies do not un-copy. A masking rule applied late is applied to the copies you can find. The security question after the fact is never fully answerable, only asymptotically so, and everyone in the room knows it.

What the brief did not know, and what it did

I want to be careful about what this document is being credited with. It is a general-purpose analytics offering from an era when the downstream consumer named in it was process mining, and the enabling technology question was about tooling and infrastructure rather than about models. It anticipates nothing about agentic systems. Its capability page does ask how a permanent analytics function should adapt to external developments including changes in law and privacy, which tells you the authors treated the external rule set as a design input rather than an obstacle, and that page carries a later edit, so I would not date that thinking to the original.

What the document has is a sequencing instinct, expressed once, in a place where nobody would put it for effect. It appears in the working part of the method, next to instructions about understanding the data model and working pragmatically with the engineers who hold the systems. That is where you put something you have learned by being burned, not where you put something you want a prospective client to admire.

The same rule under an agent loop

The reason this matters now more than it did then is that the sourcing step has changed shape.

In the method the brief describes, data gathering is an event. It happens once, in a window, performed by named people, and a rule placed before it is placed before it for good. Retrieval-augmented generation already erodes this, because the index build is a data move that most teams do not classify as one. Somebody points an embedding job at a document store and the security review, if there is one, happens on the assistant rather than on the corpus that now sits duplicated in a vector database with none of the original access controls attached to it.

Agent harnesses erode it further. Give a loop tool access to a warehouse, a ticketing system and a document store, and the gathering step is no longer an event at all. It is performed by the system, at runtime, in whatever combination the current task suggests, including combinations nobody scoped. A rule that lives in a project plan cannot reach that. It has to live in the thing that mediates the call.

In practice that means the constraint is expressed as a policy in front of the tool rather than as a review after the run. The harness holds the classification of every source it can reach. A retrieval call that cannot state the purpose it serves does not execute. Graded autonomy is set per source rather than per agent, so the same loop that may read a product catalogue on its own initiative needs a human in the path to touch anything that resolves to a person. That is exactly where we put the boundary in the RealAI Agentic OS, and it is the old instruction moved from a slide into an enforcement point.

The evaluation side follows from the same ordering. If you know before collection which fields a system may see, your evaluation sets can be built from that same declaration, and model risk management has something concrete to review: not a model in the abstract, but a stated data perimeter with a test that fails when the perimeter is crossed. Teams who collect first end up writing evaluations against whatever the pipeline happens to contain, which is a description of history rather than a control.

A safeguard written before the gathering decides what may be collected. The same safeguard written after decides what to strip from a copy that already exists, and nobody ever strips enough.

Why the order is cheaper under the AI Act

The AI Act is in force, and its heavier obligations arrive on a published schedule rather than all at once. What it attaches to a system, it attaches by purpose. A team that declared its purpose before it collected can answer that in an afternoon, from documents it already wrote for its own engineers. A team that collected first has to reconstruct purpose from a pipeline, retroactively, for a regulator, and the reconstruction is both expensive and unconvincing, because the pipeline was genuinely built without one.

None of this requires new technology and none of it is hard. It requires the security step to be written down above the gathering step rather than below it, in whatever document your organisation actually uses to plan work. The old brief did that quietly, in one line, inside a step about something else, and a decade of tooling change has not touched the logic of it.

Drawn from a proposition and capability brief written by an outside advisory firm and filed with a large European cooperative banking group's material: its delivery method, its page of project roles and its operating-model page. That document offered a method and set its own commercial bar. It records no engagement results, and nothing here should be read as one. Reading its step ordering as a lesson for agentic data access is ours.

A safeguard written before the gathering decides what may be collected. The same safeguard written after decides what to strip from a copy that already exists, and nobody ever strips enough.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?