Every platform decision I have been asked about this year has the same shape. A group has picked a stack for its first retrieval-augmented copilot, and somebody senior wants to know whether the stack will hold. Not whether the demonstration works, which it does. Whether this is still the right foundation once one pilot becomes twenty use cases and a handful of indexed documents becomes an estate.
Nobody can settle that yet, and everybody in the room knows it. What follows is where the mistake happens. Two responses go on the table and get treated as rival proposals: assess the platform properly, or design the route you take if it fails. One gets funded. The other gets described as a want of conviction.
I want to argue for funding both, and the argument comes from a document with both remedies sitting in it, one page apart, in a section never meant to be read as an argument at all.
The register
Nearly a decade ago I led an independent review of a failed offshore software build for a European banking group IT services subsidiary. The engagement was over by the time I arrived. My job was to establish what had happened across the delivery approach, the project management and the software, and to write recommendations that would survive the next sourcing decision.
Reviews of that kind live or die on their evidence register, and this one listed fifteen classes of artefact obtained and read: the statement of work and its addendums, the master agreement, the requirements, the design documents, the architecture diagram, the project plan and milestones, the weekly reports, the risk log, the issue log, the test plan, the change log, the defect log, the code fix log, the monthly bills for supplier headcount, and the communications plan naming the committees and the approved decision makers.
Two of those rows are the subject of this piece.
One page, two remedies
Take the design row first. The reviewer drew the system, then obtained the supplier's design documents and used them to confirm the drawing. That sequence is what independence means operationally, and it is cheap. Read the vendor's architecture document first and everything you check afterwards is a check of your comprehension against their claim. Draw the system yourself first, from the parts you can observe directly, and their claim gets checked against something that exists outside it.
The same discipline shows in how the review scoped itself. Its technical component was declared, in the deliverable definition, as a sample review of code rather than an exhaustive audit. That declaration costs a consultant something. It is also the only reason a later reader can tell what the report's silences mean.
Now the plan row. Two variants existed. Somebody had imagined the bad case and written a plan for it before the project started, at the one moment when that is cheap to do. The project then stayed on the expected-case variant for its whole life and ended badly.
That is not a story about a missing hedge. The hedge was there. Nobody could reach it.
Nobody could settle it, and the draft said so
The review hung on ten questions in two groups. One group covered the delivery process: whether the chosen approach fitted the expected outcome, whether project management was optimal at every level, whether development followed the usual procedures, and whether those procedures met common international practice and quality criteria. The other covered the supplier and the artefact: whether the contracted output was delivered, whether the design and the implementation met the contractual obligations, whether software quality met them, whether the specifications met common international practice, whether the software met its specification, and whether it was maintainable.
That split applies two tests to one deliverable: does this meet the contract, and does this meet the standard of the trade. A thing can pass one and fail the other, and most disappointing software passes the first.
Beside those questions sat a status column. On the draft it holds eleven marks against ten questions, laid out graphically, so which mark belongs to which question is not recoverable from the file and I am not going to guess. The distribution is safe to read, and the distribution is the finding: three plain affirmatives, two plain negatives, and six marks carrying a question mark.
That review had the contract, the artefacts, the interviews and the code, and it was looking backwards at something already finished. Six of eleven verdicts were still open at draft.
If a post mortem with the body on the table cannot settle six of eleven, a steering committee will not settle the platform question before the build. What matters is not whether you can be certain. It is what you do while you cannot.
What independent means in practice
Three procedural moves carried that review, and all three transfer.
Build your own model of the thing first. For a retrieval copilot, that means writing your evaluation set before the vendor demonstration rather than after, on your documents and the questions your people actually ask. A vendor benchmark tells you how the vendor scores itself. An evaluation set you wrote tells you whether the answers are usable here, and it is the only artefact in the exercise that stays comparable while everything underneath it moves.
Read the money. The register row that noticed a missing architect was reading invoices, not org charts. If nobody bills time against the person who would own the retrieval design, the document lineage and the evaluation regime, that role does not exist, whatever the programme structure says. RealAI's Consult team reads the staffing ledger of an AI programme for exactly that signature, because it is the cheapest evidence in the room.
Declare your coverage. Say which use cases the evaluation set covers and which it does not, in the same document that reports the score. An assessment that does not state its boundary gets read as though it had none.
- 15
- Artefact classes obtained and read in the review's evidence register
- 2
- Plan variants prepared for the project, one of them never entered
- 6 of 11
- Status marks still carrying a question mark at draft
- 0 of 15
- Register rows carrying an artefact identifier, in a column provided for one
The route that was written and never entered
Go back to the plan variants. The adverse-case plan was a real document. What it lacked was a condition. Nobody had written the sentence that would move the project from one plan onto the other, so the switch was never anybody's decision to make. Every month it was slightly too early to call, and then it was much too late.
The same shape appears twice more in the register. Its own reference column is unpopulated: fifteen rows, an identifier in none of them, and the only thing ever written in that column is a query about where the references were meant to come from. A citation discipline had been specified and could not be operated. And in the requirements row the reviewer flags a non-functional requirement stating that mobile capability need not be a concern, noting it could have been misunderstood if read in isolation. Written down, unwired, then read by somebody who had not been in the room.
Three instruments, all bought, none connected to anything.
A second plan is not a hedge until somebody has written the sentence that moves you onto it. Until then it is a document, and a project will stay on the plan it started with right up to the point where neither plan is available.
Both remedies, priced
The two remedies were not bought alike. The alternative route cost a second variant of a plan that already existed, and it was paid for before anything went wrong. The independent assessment was bought afterwards, when the engagement was over and the money was already spent: a drawing of the system, fifteen artefact classes read, a set of interviews. Neither is an expensive instrument. What was never paid for at all was the trigger, and the trigger was free.
Translate that to a platform bet now. The assessment is an evaluation set, one carefully scoped pilot, and a data readiness pass on the documents the retrieval layer will index, because most of what stops a copilot being useful sits upstream of the model, which is where a RealAI Platform engagement starts rather than on the stack. The alternative route is a short list of portability decisions taken at the start rather than at the crisis: where the embeddings live and whether you can regenerate them, whether the retrieval layer is bound to one vendor's vector store or sits behind an interface of your own, whether the prompts and evaluation sets are your assets or the platform's, and whether every answer carries lineage back to a document you own.
None of those is expensive at one use case. All are expensive at twenty. That asymmetry is the argument, and it is the same one that made a second plan cheap to write at the start of a project and impossible to adopt in its ninth month.
The governance case runs the same way. The European AI Act is agreed, and its obligations arrive on a published schedule. Systems scoped this spring are the ones that will have to be described when they do, and a platform you cannot leave is a platform whose documentation you do not control. Everything built in the meantime, from copilots to the first carefully scoped autonomous experiments, inherits that.
Recommending both remedies at once reads, in a steering meeting, as a refusal to commit. It is the opposite. Committing to a platform is what makes the second remedy necessary, because a commitment that cannot be reversed is not a decision, it is an exposure. The assessment tells you what you bought. The alternative route tells you what you do when the assessment comes back wrong. Holding both costs a known amount at the start; finding out late costs the platform and everything standing on it.
Then write the trigger. That is what was missing from the register, and it is what is missing from most of the AI platform decisions I read.
Drawn from an independent review I led of a failed offshore software build for a European banking group IT services subsidiary: its evidence register, its question set and the status column on the draft I worked from. That review produced findings and recommendations after the engagement ended, not delivered results. Reading its register as a pair of remedies is mine.
“A second plan is not a hedge until somebody has written the sentence that moves you onto it. Until then it is a document, and a project will stay on the plan it started with right up to the point where neither plan is available.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
