Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesIT service management

Case study
IT service managementA large European public-sector organisation

One option led outright on five of the six principles, and the only zero in the grid sits below the scale the grid prints above itself

A large European public-sector organisation had to decide what to do about the outsourcing contract behind its first-line IT service desk. The review that fed the decision drew all four options against one identical list of thirteen end-user services components, so the only thing moving between the four pictures was the contract boundary, then scored each option on the organisation's own six design principles. Consolidating and outsourcing the whole end-user estate led outright on five of the six and tied on the sixth. Two things in the grid deserve reading carefully. No weights are declared anywhere, so an equal weighting is assumed rather than argued, and the column totals are arithmetic performed on the table rather than figures the review printed. And the only zero in twenty-four cells, in-sourcing on value for money, sits below the one-to-five scale printed above the grid, on an estimate the review's own footnote says is incomplete. What was delivered was a review, four scored options, a recommended target model and a phased transition, not a let contract or a moved desk.

24 cellsScored judgements behind one sourcing recommendation
Client
A large European public-sector organisation
Duration
Sourcing options review, findings and recommendations
AI · RIDGE E18.2 N50ρmax 1.00
5 of 6Principles the recommended option led outright
13Scope components held constant across all four options
0In-sourcing's value-for-money score, on a one-to-five scale

A sourcing recommendation arrives in one of two conditions. Either the reader sees the frame it was judged against before they see the answer, or they do not, in which case the recommendation is an announcement and everything after it is a conversation about whose judgement to trust.

A large European public-sector organisation ran its first-line IT service desk under a long-standing outsourcing contract and had reached the point of deciding what to do about it. The review feeding that decision put four sourcing options across the top of one page and six design principles down the side, scored every intersection against a scale the page prints as one to five, and printed the grid. Twenty-four cells. The answer is legible before the argument for it begins.

The challenge

Two findings set up the decision, and neither was about how hard the supplier was trying.

The commercial one first. The contract paid by volume, which the review recorded plainly as leaving no motivation to exceed service levels and no flexibility in pricing or resourcing. Cost per incident ran between EUR 38 and EUR 49 across the five months of data the review had, standing at EUR 47 in the last of them, above the market comparators the review carried.

Then governance. The meeting machinery existed and ran at a sensible level. What sat underneath it did not. Reporting requirements were never clearly defined in the contract, and neither were the meetings and reviews themselves, so a structure that met regularly had no written obligation to produce anything. Two roles were absent or left undefined: someone accountable for whether the supplier was actually performing, and someone accountable for whether the knowledge base was correct. A structure pointed at nothing in particular.

Neither is fixed by trying harder, and neither is the supplier's to fix alone. Each is a consequence of how the service was bought. That makes this a sourcing question rather than a performance question, and it is why the frame had to be published before the answer.

The approach

The first move is the one most options papers skip. All four were drawn against one identical list of end-user services components, thirteen of them, repeated unchanged under each option: the first-line desk and the back office behind it, desk-side support, client hardware, file and print, unified communications, identity management, desktop engineering, desktop virtualisation and support for the organisation's own applications. The only thing that moves between the four pictures is where the contract boundary is drawn. That makes the options structurally comparable rather than rhetorically comparable: a reader can see what each buys and what each surrenders.

The four were: stay with the incumbent supplier and renegotiate; bring the desk and desk-side support back in house; re-tender the desk on its own; or re-tender the entire end-user support estate as one consolidated contract.

Then the frame, which belonged to the organisation rather than to the reviewers. A desk flexible and agile enough to absorb continuous change in the IT estate. Higher productivity, efficiency and quality. A service aligned to the business and working correctly from the start. A customer-centric service. One unified desk covering every end-user issue, with a stated preference for a single provider. Value for money.

Then the scores, in the order the options were listed above. Flexibility: two, three, three, five. Productivity, efficiency and quality: two, three, three, four. Business alignment and service centricity: one, three, three, three. Customer centricity: one, two, four, five. A unified desk under a single provider: two, one, three, four. Value for money: two, zero, four, five.

Consolidating the whole estate leads outright on five principles and ties on the sixth. Nobody needs the recommendation read out.

Two things in that grid deserve saying out loud, because a frame published in good faith still has to survive close reading. No weights are declared anywhere. Six principles are scored and none is said to matter more than another, so an equal weighting has been assumed without ever being argued. Add the columns and the spread is wide, ten, twelve, twenty and twenty-six out of a possible thirty, but those totals are arithmetic we performed rather than figures the review printed. The winning option wins on breadth, not by being decisive on any single principle. And breadth is what makes the missing weights hard to exploit. Against the desk-only re-tender, consolidation leads by one point on value for money and by an average of one point across the other five, so shifting weight onto value for money moves the gap not at all. The only reweighting that erases it is loading the frame onto business alignment, the single principle on which those two tie.

Then the zero. It is the only zero in twenty-four cells, and it lands on in-sourcing under value for money. Two things about it. The scale printed above the grid runs one to five, so the hardest score in the matrix is not on the scale the review declared, and nothing on the page says what a zero means. And the number underneath it is softer than the score looks. Turn back one page to where the options are defined and the headcount for in-sourcing is left as a placeholder, a rise of an unstated number of full-time equivalents plus unquantified increases in recruiting, training and technology. The cost itself does appear, earlier, in the review's own summary of the four options: EUR 4.3M for in-sourcing against EUR 2.7M for the current contracts. Its own footnote then removes travel, recruitment, technology, training, overtime and depreciation from that figure and applies internal day rates only. The zero is defensible. It is defensible on an estimate the review flags as incomplete, and incomplete in the direction that flatters the option it was used to reject.

The test transfers to any business case. Find the cell that decides the outcome, then find the page that printed the number behind it, and the footnote under that number.

The outcome

What this engagement produced was a review, four scored options, a recommended target model with a service integration layer at its centre, and a phased transition with warnings attached. No contract was let and no desk was moved. The scores are judgements about attractiveness, not results.

The consolidation case rested on a benchmark of eight organisations mapped across nine end-user services components. All eight had the service desk inside the bundle; only two had bought the service integration layer the recommended model puts at the centre. Four carried quantified outcomes, and breadth does not track them: the largest cost reduction stated as a figure, thirty-five percent alongside thirty percent fewer tickets, came from a four-component bundle, while the only quantified fall in incidents, forty-eight percent with eighteen percent off operating cost, came from the narrowest quantified bundle of the four. Those are outcomes other organisations reported to the reviewers, not results this client obtained.

The transition warnings age best. A supplier arrives with a professional transition team and the buyer arrives with a project manager. Leave the supplier holding the sequencing decisions and they get taken for the supplier's reasons, saving its cost rather than reducing the buyer's risk. The recommendation was a mirror organisation on the client side from the outset, and knowledge transfer from the outgoing supplier treated as a deliverable rather than a courtesy.

All of this is now being replayed with a different noun in the middle of the sentence. The live question is no longer only who staffs the first line, it is how much of it a retrieval-augmented copilot can hold and what that copilot may touch. The temptation is unchanged: pick the answer, then assemble the reasons. So is the discipline. Write the principles down before any candidate exists, hold the scope list constant, score every option on the same lines, publish the grid rather than the conclusion.

The value-for-money zero is the score most likely to move, because the headcount arithmetic that made in-sourcing unaffordable is what an assisted desk changes. And the frame now needs a line it did not have: whether the material the answers would be drawn from is fit to answer from. The same review found ticket text that could not be searched and a knowledge base that did not update itself. No model quality survives that. Data readiness, lineage and a written evaluation set belong on the sourcing grid beside value for money, not in a workstream that starts after signature.

Our Consult work opens on the frame rather than the answer, and the Platform work behind it starts with the substrate: what the estate can expose, and whether the record of past resolutions is worth retrieving from. A copilot specified without either is a demonstration.

A grid does not make a decision correct. It makes the decision arguable, which is the only property a sourcing choice can honestly claim in advance. Publish the frame first and the argument that follows is about principles, not about who gets believed.

NEXT STEP

Ready to make AI real?