Almost every AI portfolio I am asked to look at is ordered by something other than comparison. A proposal arrives and is argued on ground its sponsor picked. The retrieval-augmented search pilot is defended on how much of the document estate it reaches, the sales copilot on minutes saved per seller, the forecasting rebuild on lineage and data readiness. Each case is competently made. No two are commensurable.
So the queue gets ordered by whatever is left: which sponsor sits highest, which team demonstrated best, which number sounded largest read out loud. That is not dishonesty. It is what happens when nobody wrote down what the organisation is scoring for, so every proposal nominates its own scoreboard on the way in.
The cleanest counter-example I have worked on sits in a digital blueprint written for a global staffing and HR services group. It ran two tracks: a short-cycle track that took a business wish list down to a handful of high-impact ideas, and a longer track arguing for a single technology layer above the country back offices. The interesting thing is not either conclusion. It is that the second track published its scoring frame in its opening pages and never allowed a decision to be argued off it.
The rubric came before the options
The order of the document matters more than its contents. The five principles sit in the section establishing business requirements, upstream of every architecture page. They are stated plainly enough that nobody needs a workshop to interpret them, and generally enough to apply to a platform choice, an operating-model choice and a capability model without rewriting.
That breadth is usually treated as a weakness. Principles this general, the objection runs, cannot discriminate between serious options. The opposite happened. Because the criteria were broad they survived contact with three very different decisions, and because they survived, those decisions became comparable to each other. A criterion narrow enough to settle one argument cleanly gets retired before the next arrives.
The group also ran two business lines with different centres of gravity: throughput work, where the win is straight-through processing and efficient technology, and professional placement, where the win is a consistent experience across channels. They want different things badly enough that the obvious answer is to build two of everything. Five shared principles made the single-platform argument reviewable rather than merely asserted, because both lines had agreed what good looked like before either was asked to give anything up.
What a shared frame does to a decision
The operating-model page is the most rigorous artefact in the document, and it is worth being precise about what it is. Four options for running digital capability across a multi-country group, laid against the five principles. Twenty cells, each a sentence of judgement about how that option performs against that criterion. No scoring, no weighting, no total.
I have come to prefer that to a scored matrix. A score compresses a judgement into a number that can be argued about arithmetically, which moves the discussion onto weights and away from the sentence underneath. A written cell cannot be re-weighted. It can only be agreed with or contradicted, and contradicting it means saying something specific.
Read across the rows, the summary is that centralised options lose. Read down the columns, which is the reading the document itself invites, something more useful appears. Two of the four options carry substantially the same judgement on four of the five criteria. Both are judged to engage local markets enough that the solution can be tailored. Both are credited with using skills already present in those markets, with pairing central oversight and local execution to hold consistency, and with sharing practice because the central unit has local representatives sitting in it. On paper they are the same model.
They diverge in one cell. One option has an accountable leader for delivery and coordination. The other has no single point of accountability and depends on goodwill. That cell decided a multi-country operating model.
Nobody reaches that conclusion arguing options one at a time. Argued separately, both would have been defended as federated, locally engaged and pragmatic, and both defences would have been true. The frame made the one real difference visible, in a form a reader can check.
Losing legibly
The other half of the value shows up on the losing side. Every rejected option has its failure mode written out in ordinary language. Solutions that often miss local needs. Quick to mobilise, slow to market. Hard to implement where the organisation is not centralised. Poor consistency because local adoption is hard to win.
None of those sentences is cruel and none is vague. A sponsor whose preferred model lost can read the exact criterion it lost on and the sentence that did the losing. Two things are then possible: accept it, or come back with evidence that the sentence is wrong. Both are healthy.
A per-proposal argument produces the third outcome instead. The sponsor concludes the decision was political, waits a quarter, and returns with the same idea attached to a different justification. Much of the AI programme rework I see is that loop running quietly.
A shared frame does not make any single score correct. It makes the losing argument legible, and a losing argument that is legible does not come back next quarter wearing a different number.
The same discipline at two other scales
The habit repeats at two other scales in the same programme, which is what convinces me it was a habit and not a single good page.
At the level of individual initiatives, the short-cycle track scored candidates on a two-axis grid: value of the opportunity against ease of implementation, three bands on each axis. Every idea in the business's unfiltered ambition went through the same two questions, and the document puts the whole run, wish list to completed playbook, at as little as six to eight weeks. A short clock keeps a prioritisation grid honest, because a frame with no deadline gets re-cut until the favoured project lands in the favoured quadrant.
At the level of country units, a capability review compared operating companies against each other on five dimensions covering data management, people, technology, organisation and users, in four steps: participants fill it in themselves, an internal and optionally external benchmark follows, the result becomes a value and capability roadmap, and the outcome is shared in a workshop. Its stated purpose was to promote reuse and sharing of best practice. That is a comparison instrument aimed at peer pressure rather than audit, and it works for the same reason the option grid works. One country unit adopts what another built when it can see, on a frame it filled in itself, exactly where it sits.
- 5
- Design principles published once, then reused as the evaluation frame twice
- 4
- Operating-model options put through that same frame
- 20
- Cells of written judgement in the comparison grid, no scores
- 1
- Criterion that separated the two options still standing
Why this is more urgent now than it was then
Nothing above needed machine learning, which is why it is worth raising this quarter rather than filing it.
The number of AI proposals inside a large enterprise has multiplied since generative models put a working demo within reach of any team, and the marginal cost of producing one is close to nothing. Anyone can stand a retrieval-augmented prototype up over a vector store in an afternoon and demonstrate it convincingly on Friday. The same goes for copilot pilots, and even the carefully scoped autonomous experiments, the ones with a person approving every consequential step, are cheap enough that a business unit can run one without asking. The supply of plausible things to fund has outrun any leadership team's capacity to compare them.
The cost of a wrong ordering has risen at the same time. Data readiness and lineage work is slow and unglamorous and gets deferred every time a demo outranks it. Europe's AI Act has been through political agreement, and the obligations it will bring are foreseeable enough to belong in a scoring frame rather than in a scramble later. An organisation that cannot order its own queue keeps funding what demonstrates well and starving what everything else depends on.
An evaluation set is the same argument one layer down, and it works for the same reason: written before the candidates, applied unchanged to all of them, readable by whoever lost. Teams that accept that discipline for a model and refuse it for a portfolio have the logic exactly half applied.
Writing one
Four rules, none of which requires a data scientist. Publish the frame before the options, and date it: a rubric produced after the proposals arrive is a rationalisation with a table around it. Keep the criteria few and broad enough to survive three different decisions, because a criterion retired after one use never made anything comparable. Write cells, not scores, so disagreeing with a judgement means saying something specific rather than adjusting a weight. Put every rejected option's failure mode in writing, in language its sponsor would recognise, because a decision nobody can lose gracefully gets relitigated.
That is the opening move in a RealAI Consult engagement, and it is why we ask what you are scoring for before we ask what you want to build. Answer it and you have built nothing. You have a queue that can be put in order, and a record of why it is in that order.
Drawn from a digital blueprint written for a global staffing and HR services group: its published design principles, its capability model and its operating-model comparison. That document produced findings and a recommendation, not delivered results. Reading it as a decision-hygiene lesson is ours.
“A shared frame does not make any single score correct. It makes the losing argument legible, and a losing argument that is legible does not come back next quarter wearing a different number.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
