The deliverable page is the most-read page in any proposal and usually the least useful one. It names things that will arrive. Arrival is easy to verify and tells you almost nothing, because a document arriving on the agreed week is a fact about a calendar, not a fact about your operation.
I have kept one deliverable page for years because it does something I still rarely see. It was written for an international health insurance group being offered a design engagement across two strands of a wider change portfolio, and it runs in two columns. The left column names an output and describes what it physically is. The right column, set in plain second person, says what the group will be entitled to be sure of once that output lands.
Two columns, and only one of them is about us
The engagement was described as twelve numbered outputs rather than as a run of activities. Six of the twelve were then specified properly, and the specification has two halves.
The first half says what the thing is. Not its title, its form. One is a document setting scope, objectives, the participation required from stakeholders, the delivery plan and the method. One is a pack of operating-model options, each carrying an impact assessment, the capabilities the model would need, and the functional responsibilities that go with them. One is an analysis pack agreeing the future capability set across people, process and technology and measuring the gap against the model in use today. One is a slide deck holding the end-state design with regional and country views, not a single global picture. One is a spreadsheet model of costs and benefits, comparing the cost of the proposed organisation against the cost of the current one and modelling the automation benefit separately. One is a single picture backed by a phased plan in which only the first phase carries named resources.
The second half is a sentence about the reader. Knowing what is being delivered, by when and how, so that delivery can be measured. Having a clear and detailed analysis of the options for a model that fits the purpose. Understanding the distance between the model being run now and the one agreed for later. Holding a picture of the operating model supported by descriptions of the people, the processes and the systems underneath it. Having a view of the cost and savings impact of the change. Having clarity over what has to be delivered once implementation starts.
Read the two halves back to back and the grammar does the work. The left column's subject is the team producing the output. The right column's subject is the organisation paying for it. Nothing on the left can be checked without the seller present, because only the seller knows whether the pack contains what the pack was meant to contain. Everything on the right can be checked by the buyer alone, long after the team has gone, by trying to say the sentence out loud and noticing whether it is true.
The pairings mostly announce themselves. A document that sets scope, method and dates is the only one of the six that can answer an assurance about knowing what is delivered, by when and how. A capability gap analysis is the only one that can answer an assurance about the scale of change between today and the agreed future state. The ledger is not being clever. It is being explicit about something most proposals leave the client to infer, which is why each output was worth buying.
What the second column is actually for
An assurance sentence is an acceptance criterion, written before anybody knows whether it will be comfortable to meet.
That timing is the whole value. Criteria written at handover are negotiated between two parties who both now know how the work went, which means they are written to be passed. Criteria written into the proposal are written by people who cannot yet see the result. They are also short, and a sentence a chief operating officer can hold in their head survives a change of programme director where a requirements annex nobody reopens does not.
The column is also a filter. Writing it exposes the outputs with no assurance behind them. If the only sentence you can manage restates the deliverable's name, the deliverable is decoration. If two outputs generate the same assurance, one is a duplicate with a different cover.
The column that machine learning work never writes
Now put an artificial intelligence programme through the same page, because this is where the conflation lives.
A current engagement list looks something like this. A data readiness assessment. A lineage record for the fields that feed a decision. An evaluation set built from real cases and approved by the business. A vector store over a document estate. A retrieval-augmented assistant for a service team. A copilot inside a claims workflow. Perhaps one early and carefully scoped autonomous experiment in a queue where being wrong is cheap.
Every item on that list is a thing that arrives. Not one of them, stated as a name, says what the organisation can then be sure of. So write the second column and watch what happens.
The evaluation set entitles you to say that any future change to the assistant can be scored against a fixed set of cases your own people wrote and signed off, so improvement becomes a measurement instead of an impression. The lineage record entitles you to say, for any figure the model consumes, which system produced it, when, and under whose definition. The retrieval assistant entitles you to say what proportion of its answers cited a document a named person can open and check, measured on that approved set rather than on the demo. The copilot entitles you to say what share of a specific queue now clears without a person touching it, and what happens to the share that does not.
Notice that every one of those sentences is about the operation, and none of them is about the model. That is the point of the exercise. "We delivered a model" and "we changed the operation" are two different claims, and in a deliverable list with only one column they arrive fused, because the artefact is persuasive before it is assured. A model demonstration looks finished in a way that a capability gap analysis never does. It runs, it answers, it impresses a steering group, and it can do all of that while the queue behind it is untouched, the threshold for acting on its output is undecided, and nobody has been given authority to act differently because of what it said.
That is why the assurance column matters more in this work than it did in operating-model design. The deliverable is more convincing and the change is less visible.
A model that scores a case changes nothing until a queue, a threshold and somebody's authority to act change with it. The assurance sentence is where you find out whether any of those three were in scope.
The half of the ledger that had nothing
Six of the twelve outputs carried this treatment. The other six carried a number, a name and a week.
I read that as instructive rather than sloppy. Assurance prose is expensive, and an organisation writing it for every output would produce a page nobody reads. The six that got it are the ones a decision hangs on: what we are buying, what the options are, how far the change reaches, what it looks like at the end, what it costs, what happens next. The stakeholder plan and the refined journeys did not need a sentence, because nobody was ever going to argue about whether they had arrived.
The same proposal is honest in one more way worth copying. Its phased plan resources the first phase in detail and leaves the later phases at outline level. Assurance narrows with distance, and saying so on the page is better than pretending a plan far out is anything other than an intention. Any governance regime drawn up around this technology will be easier to live with in an organisation that already distinguishes what it can promise from what it can only intend.
Five lines to steal
Take the next statement of work that involves a model, a copilot or a retrieval system.
For every named output, write one sentence beginning "you will be able to say". If the sentence needs the deliverable's own name to make sense, the output is not yet specified. If two outputs produce the same sentence, delete one. Name the person who has to be able to say it, and make it somebody in the operation rather than somebody in the programme. Then check that at least one sentence is about a queue, a cycle time or a decision rather than about an artefact, because a list of assurances that are all about artefacts describes a very good pilot and no change at all.
That check is the opening move of a RealAI Platform engagement, and it is why we ask what you want to be able to say before we ask what you want us to build.
The ledger described here comes from a proposal to an international health insurance group covering two strands of a wider change portfolio, much of it the proposing firm describing what it can do. The twelve outputs, the six detailed specifications and every paired assurance were offered. Nothing was delivered, no assurance was ever tested, and the source does not say whether the engagement was bought. Reading the second column as the reusable part is ours.
“A deliverable list is written in our grammar. An assurance list is written in yours. Only one of those two sentences can be tested by the person paying for the work.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
