A scoring matrix exists to stop a decision from being an opinion. You write down what you care about before you look at the options, you mark every option against every criterion, and the answer falls out of the arithmetic instead of out of the room. A matrix is also the instrument most often overruled on the very next page, and the overruling is usually more interesting than the score.
A large European public-sector organisation was deciding how to source its IT service desk. Four options sat on the table: keep the incumbent arrangement and renegotiate it, bring the desk back in-house, retender the desk roughly as it stood, or retender a wider scope that folded the desk together with the rest of end-user support, hardware provisioning included. Each option was marked against design principles the organisation had agreed before the options existed, every one of them a statement about what it wanted the service itself to be.
The marks put one option in front. The recommendation went to a different one.
The challenge
Start with what the page can and cannot say, because that boundary matters. The marks on that page are shaded glyphs rather than printed figures, so what it makes legible is an ordering, not an audited total. No total is reproduced here, none is reconstructed, and none is needed. The ordering alone carries the finding: the option the rubric put in front was not the option recommended.
That is not a scandal, and reading it as one would be the lazy conclusion. The stated reason for going wide is written plainly in the option narrative itself. A retender of the service desk alone risks a limited market response because the scope is small. Widen the scope to include the hardware estate and the deal becomes worth winning, which brings in a wider field. The failure mode being avoided is familiar: you tighten the specification until the deal looks clean, two bidders show up, one is the incumbent, and you spend the next several years negotiating from a position you designed for yourself.
So the argument is right. The problem is that the argument is not in the rubric. Whether a scope will attract a strong field is nowhere among the principles being scored. Every row is about what the organisation wants from the service once it exists. Not one row is about whether anyone good will bid to provide it. The matrix was asked how attractive each option is to us, and the decision turned on how attractive each option is to them.
The recommended option is the one that carries the most future scope and therefore the least present evidence. Read the grounds given for it and every one of them is a potential rather than a measurement: better service integration, a wider field of bidders, a better price later. None of that is dishonest. It is what happens when you keep a scoring instrument in the deck after the decisive criterion has moved outside it. The instrument stops being load-bearing and becomes decoration, at precisely the moment a reader is relying on it to show the work.
There is a matching asymmetry in the numbers that were published. The two options that were rejected are the two that carry a euro figure, one footnoted with where the number came from, the other with a disclosed list of what the estimate leaves out. The two retender options carry no figure at all. That asymmetry is structural rather than sloppy, because you can cost a contract you already hold and a payroll you could build, and you genuinely cannot cost a tender nobody has run yet. But the effect on a reader is that measured things are being compared with hoped-for things, and the hoped-for thing wins.
The approach
There are two honest repairs, and both are cheap.
The first is to add the missing row and mark again. Attractiveness to the supplier market is not a soft criterion. It has observable proxies: how many credible providers serve a scope of this shape in this region, how many bid the last time a comparable scope went out, what the minimum viable contract value looks like for a provider with a delivery organisation nearby. Score that alongside the rest and the wide retender probably wins on the arithmetic too, which would make the recommendation follow the page instead of contradicting it.
The second is to keep the rubric as it stands and say out loud that the recommendation departs from it, naming the criterion that overrode the score. That is a stronger document than a quietly inconsistent one. A reader who is told what the rubric captured, and told that the decision turned on something outside it, can argue with the thing outside it. A reader left to notice the mismatch alone stops trusting the rest.
The general rule underneath both repairs is the one worth taking away. Before you score anything, write down the sentence that will actually end the argument in the room. If that sentence is not a row in your instrument, either put it in or put the instrument away.
The outcome
What this engagement delivered is a recommendation and a sequence, not a result. The advice was to agree the outsourced scope first, then stand up a project to develop the technical conditions and the contract, then plan the exit from the current arrangement in parallel so that the transition material exists whether or not the incumbent stays. Nothing had been tendered. No saving had been banked. The value of the piece is the reasoning, and the reasoning has a hole in it that is visible only because someone drew a matrix and then argued past it.
The same shape shows up now in every technology sourcing decision we are asked to sit in on, and the current one is machine learning tooling. Organisations build a feature grid for platform selection, mark the candidates on model deployment, monitoring, pipeline orchestration and cost, and then choose on something that never appears in the grid at all. Sometimes it is which vendor the data engineering team already knows. Sometimes it is which one the security review will clear this quarter. Sometimes it is the honest and unwritten judgement that one supplier will still exist in three years and another will not.
Our Consult work starts by asking a client to name that unwritten criterion before the grid gets built, because it is far cheaper to argue about it at the start than to discover it as a footnote on the recommendation slide. The two rows most often missing are the same two every time. The first is data readiness, meaning whether the feature store, the lineage and the access controls exist well enough that a deployed model can be traced when someone asks why it decided what it decided. The second is operability under real volume, which process mining over the existing workflow answers in a fortnight and a demonstration answers not at all. A platform that wins a feature grid and loses on those two rows will be replaced, at your cost, by whoever is in the chair after you.
The early language-model pilots being run in service desks right now sharpen the point rather than changing it. They are being scored on how good the answer looks in a demonstration. They will be lived with on the basis of what happens to the answer afterwards: whether it lands in a system of record, whether the resolution can be audited, whether a person has to retype it into a ticket. That last question decides everything and it is on almost nobody's scorecard. It is the row Hominis was built around, because the hard part of putting a machine-generated answer in front of a person is not producing the answer, it is what happens to it next.
A rubric that omits the thing you will really decide on is not neutral. It is a promise of rigour that the next page breaks.
