Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

Need to Check This

RealAIJun 26, 20247 min read
Delivery AssuranceBankingSourcingAI StrategyEvidence

Every review produces one sentence somebody will read out loud. In a delivery post-mortem I worked on for a European banking group's IT services subsidiary, examining an offshore build that had gone badly enough to be brought back in house, that sentence was short and sat inside quotation marks. The supplier "did not deliver quality code".

It is the line an executive quotes, and the line that travels to a steering committee without the slides around it. On the working draft it is also the only line on its page that its own author marked as not yet established. At the end of the same bullet, after the attribution and inside a parenthesis, four words: need to check this.

I have gone back to that page because the habit it records is getting rarer, not commoner, at exactly the moment the tooling makes it matter more.

What is actually on the slide

The delivery-approach page is a working page, organised under two headings, one per side. Against the supplier: did not have sufficiently qualified staff, did not write requirements, did not take issues seriously enough, over confident of the client relationship. Against the client, under its own heading and in the same plain register: already nervous about delivery, under pressure, tight deadline; potentially rushed into the contract; could have changed the way it approached testing; should have sent a technical architect over to run the project if it did not believe the supplier capable.

A reviewer who wanted a quiet life writes the first list and stops. This one wrote both, which is the first sign the page is honest.

The second sign is smaller and better. Of all those bullets, exactly one is in quotation marks, exactly one is attributed to somebody other than the reviewer, and exactly one carries a warning. They are the same bullet.

What the four words mean

Read quickly, need to check this looks like uncertainty about whether the code was bad. It is not. The technical section of the same review carries its own observations under that heading, specific and unflattering: a team new to the front-end framework with little experience of it, requirements never written, an inexperienced group, no technical architect in place throughout. Those were written in the reviewer's own voice, credited to nobody else and hedged nowhere. The question was never whether there was a problem.

The note is about a different property of the sentence: where it came from. Everything else on that page is first-person observation. That line is a report of what one contracting party said about the other, quoted, mid-dispute, with commercial consequences attached to the answer. It is admissible as evidence that the client believes it. It is not yet evidence that it is true, and the two are separated by a piece of work nobody had done yet.

The distinction survives into how each version would be evidenced. To establish that the client believes it you interview the client, which had been done. To establish that the code was bad you read the code against criteria written down before you looked, which is a different workstream with a different artefact at the end. That is a boundary, and the parenthesis is where it is drawn.

The grammar of what we still owe

Once you notice the annotation you find the whole family of them, because the notes follow a consistent shorthand and were left in the draft rather than tidied away.

The code-quality page opens with the client's concern about the quality of the supplier's code and then, on that same line, where the substantiating detail belongs, a run of placeholder characters and a bracketed note: what did we find. Below it sits the client architects' view that the platform architecture allows too little code reuse. Under recommendations there are three more lines that are nothing but a repeated letter. The page is not claiming anything. It is holding a shape open until evidence arrives.

Under possible reasons, the first explanation offered, that the architecture design was not in line because architecture had never been a key priority, carries good question to ask the architects. The second, that the supplier could be putting relatively inexperienced developers onto projects, carries need to find this out. On the estimation page it is not the perception that is flagged but the account of how the numbers were built, by assigning roughly the number of days a piece of work is likely to take on gut feel: this is key aspect to ask both sides. And on the divider for the delivery-approach section, a drafting instruction to a colleague: make it more structure and show evidence such as artefact 1 read and this done.

None of that was written for a reader. All of it is the residue of a person tracking, claim by claim, which sentences had earned their place. The findings pages show what earning it was going to look like: each is a register with the same set of columns, the aspect, the two views, the finding, and a reference to the artefact it came from. Fill that last column or the row does not ship. In this draft none of those registers has a row in it yet, which is the part worth noticing. The column for the evidence was built before there was any evidence to put in it, and the working pages carry parentheses precisely because it is still empty.

need to check this
Beside the review's most quotable sentence
what did we find
Where the code-quality evidence belongs
good question to ask the architects
Beside the first of two possible reasons
need to find this out
Beside the second

Why this is what makes a review survive

A post-mortem of a failed outsourced build is read by two organisations with money and reputation in the outcome, and every sentence in it is a sentence somebody is motivated to break. The ones that break are the ones where the reviewer's own confidence outran the reviewer's own evidence, because the counterparty only has to find one and the rest of the report inherits the doubt.

The estimation section shows the alternative. The client believed the supplier's estimates ran two to three times high, and one stakeholder put it at five times what was first quoted. The review does not convert that into a finding about the supplier. It records the perception as a perception, then files under possible reasons the thing that makes it untestable: the client had no defined method for estimating effort, so its own figures were inconsistent across projects and the supplier had nothing stable to learn from either. What follows is not a verdict but a method, carrying an unusual instruction: the client's own team should learn the chosen approach and apply it to earlier releases, comparing its results against what the supplier had quoted, before holding anybody to it.

The same instinct shows up in the recommendations at the back, where a set of proposed status definitions for future assignments is offered with an explicit disclaimer that this does not imply the client's past ratings were wrong. It is a template, so that on the next engagement there can be no disagreement about how good or bad things are. Agree the meaning before you need it to settle an argument.

The habit that generated analysis does not have

I keep this page open whenever we design, with the RealAI Platform team, anything that drafts analysis for a human to sign.

A retrieval-augmented system built over a document estate will happily produce the sentence about quality code. It will produce it fluently, in the same register and with the same confidence as the sentence about the missing technical architect, and nothing in the output will tell the reader that one of them was read off a code review and the other was overheard in an interview with a party in dispute. Citation does not fix this. A chunk retrieved from a vector store tells you which document a sentence came from, and that is provenance, not verification. The document may be a transcript of somebody's accusation.

So the property to engineer for is not fluency, and not even accuracy in the narrow sense. It is that the output separates what it has evidenced from what it has been told, and marks the second visibly enough that a reader in a hurry cannot miss it. That means evaluation sets scoring more than whether the answer is right: whether the claimed confidence matches the support, and whether the assistant refuses to promote a reported claim into a finding. It means lineage carried as a first-class field, so the status of a statement travels with the statement. And in copilots drafting for regulated functions it means the interface has somewhere to put the parenthesis, because if there is nowhere to write need to check this, nobody writes it.

Banking has a further reason to care. The European rules on AI have reached political agreement, and the shape of what supervisors will ask is visible already: not what the system said, but what the claim rested on and who checked. An estate that cannot answer that about its own generated text has a documentation problem well before it has a model problem. The same holds for the early, carefully scoped autonomous experiments worth running at all: the ones that can be widened later are the ones that mark their own boundary.

The reviewer had no tooling for any of this. There was a bracket, a lower-case note, and the discipline to leave it in the draft where colleagues would see it. That is the cheapest instrument in the business, and the most reliable thing on the page.

Drawn from the working draft of an independent delivery review of an outsourced software build, commissioned by a European banking group's IT services subsidiary after the work was brought back in house. The quoted annotations are the drafting notes as written. The review produced findings and recommendations, not results, and this piece takes no position on whether the flagged claim was ultimately established.

The most quotable sentence in the review is the one its author declined to stand behind yet. Four words in a parenthesis are doing more work than the five words they qualify.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?