Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesIT service management

Case study
IT service managementA large European public-sector organisation

Two lines in a long recommendation list name a quantity, and a recommendation that names a quantity can be finished

A large European public-sector organisation had its outsourced IT service desk reviewed ahead of a re-sourcing decision. The technology recommendations run to a familiar list: make the ticket tool easier to configure, build interfaces to the second-line tools, launch a usable portal, add chat, publish self-help videos. Two of them are a different kind of instruction. Automate password reset in the phone system, and identify and automate the five most frequent service requests. Those two name a quantity, which means the work has an end and a team can be told when it is over. The rest of the same review shows why the bounded version is the only one that would have survived: ten support groups named around a single outsourced first line, a contract scoped as a list of activities rather than outcomes, monthly demand swinging from 3,592 incidents to 9,658, and a unit cost between 38 and 49 euros against benchmarks the review itself quotes at roughly 22 and roughly 10. What was delivered was a diagnosis, a bounded recommendation and four sourcing options, not an automated request.

5 + 1Requests named as the automation scope: the top five, plus password reset
Client
A large European public-sector organisation
Duration
Current-state service desk review, findings and recommendations
AI · RIDGE E59.1 N37.5ρmax 1.00
10Support groups the review names as relevant to the desk's work
2.7xSwing between the quietest and busiest month of recorded demand
EUR 38 to 49Cost per incident across the months the review costed

A recommendation you can finish is worth more than a recommendation you can agree with. Most of what a service desk review produces is the second kind: define the processes clearly, establish ownership, align the reporting to the business, improve knowledge management. All of it true. None of it countable, and none of it done on any particular Friday.

Two lines in one review of an outsourced IT service desk at a large European public-sector organisation are the other kind. Automate password reset in the phone system. Identify and automate the five most frequent service requests. Both sit in the technology section, in the same block as the portal, the chat channel and the self-help videos, and they are structurally unlike everything around them. They name a quantity. Somebody can be told when the work is over.

The challenge

The estate underneath explains why the bounded version is the only one that would have held. The outsourced desk was the single first point of contact, and the review names ten support groups as relevant to its work, first, second and third line among them. The contract described the service as a list of activities rather than a list of outcomes, and the activities it excluded covered the organisation's own bespoke applications and problem resolution generally. Every ambiguity in that boundary arrived as an escalation.

Demand did not sit still either. Across the recorded window, monthly incidents ran from a low of 3,592 to a high of 9,658, a swing of about 2.7 times, and the review attributes the peaks to technology roll-outs. Unit cost sat between 38 and 49 euros per incident across the months the review costed, standing at 47 in the last of them, against the two benchmarks the review carries alongside it: roughly 22 euros for an on-premise desk and roughly 10 for a remote one.

Now imagine writing a business case that says "automate the service desk" against that. Which desk? The activities the contract lists, or those plus the excluded ones that keep arriving anyway? Which volume, when the busiest month is nearly three times the quietest? Which unit cost, when the same document reports two different totals for the same month depending on whether you count contacts or incidents? A programme scoped against the whole inherits every one of those arguments and never reaches a first release. A programme scoped as five named request types inherits none of them. It needs one ranked list.

The catch is that the ranked list did not exist. The review's own findings about the ticket tool are unsparing: free text inside incident records could not be searched, the automatic update into the knowledge base was not functional, and service level reports could not be generated at all. So the instruction is not really "automate the top five". It is "count first, then automate the top five", and the counting is the part the organisation was least equipped to do. That ordering is the whole method, and it is the step almost every automation programme skips, because counting produces a smaller answer than a strategy does.

The approach

Password reset is named separately from the top five, and the separation is the interesting part. It is not on the list because it is frequent, although it is. It is on the list because of its shape. The resolution is a single state change behind an identity check, there is no judgement between the request and the outcome, the entitlement is machine-checkable, and a failure is loud and immediate rather than quietly wrong. That is the profile of work a machine can own outright.

Note also where the review puts it. Not in the portal. In the phone system, which is where roughly four contacts in five actually arrived. The organisation had a self-service portal that took a rounding error of the month's traffic, and the reflex would have been to fix the portal and route the resets through it. Automating inside the channel people already use removes the request without asking anyone to change their behaviour first. Deflection that depends on a habit change is a project. Deflection that meets the existing habit is a feature.

That gives a rule the review does not state but its two recommendations imply together. Frequency picks the shortlist; shape decides what makes the cut. A request qualifies when the outcome is a defined state rather than an opinion, the entitlement can be checked without a human reading a policy, and a wrong answer is visible to the person who asked. Intersect the ranked list with that test and the top five is usually a genuine five, not a hopeful fifteen.

The tooling has since moved the ceiling on that test without moving the method. An agent loop with tools and a harness around it can carry a request further than a phone menu ever could, and graded autonomy lets one request type start as a drafted suggestion, move to an action a human approves, and end as an action the agent takes alone once an evaluation set says it is safe to. Every step of that progression still begins with the same counted list, which is where our Platform work starts. With the EU AI Act now in force, an identity-adjacent action taken automatically also has to leave a decision trail somebody can inspect afterwards, which is one more reason to begin with a request type whose correct outcome is unambiguous. Retrieval-augmented answers are the same story: they are only as good as the corpus behind them, and this desk was running its knowledge in two competing places, neither of them updated automatically by the ticket tool.

The outcome

This engagement produced a review. A current-state diagnosis, a market study, a bounded set of recommendations and four sourcing options drawn against one common scope so they could be compared honestly. No request was automated in this phase, no password reset flow was built, and the top five was never actually counted inside the document. The recommendation is to identify it. What exists here is a method and a scope, not a result, and the distinction matters because the interesting claim is about the shape of the instruction rather than about anything that shipped.

One commercial detail decides whether the instruction could ever have been carried out. The incumbent contract was volume-priced, with service levels driven by the number of incidents, and the review records the consequence plainly: no motivation to exceed service levels. Under that structure, every request the supplier automates is revenue the supplier gives up. The bounded backlog is not blocked by the technology in this case. It is blocked by an arithmetic in the contract, which is why the automation recommendation and the pricing recommendation had to travel together, and why any organisation buying automated support today should read its own payment terms before it reads a vendor's capability deck.

What would be done differently now is mostly that the counting stops being a one-off. The ranked list of request types, their true handling cost and their rework rate can be mined continuously from the ticket records rather than assembled by hand for a review, which turns the top five from a finding into a standing measurement that reorders itself as the estate changes. Our Consult engagements open on that count, because a scope agreed without it gets renegotiated by month four. Then the number goes up honestly: you earn the next five by finishing the first five, and the evidence that you finished is the queue getting shorter, not a slide saying the programme is on track.

The reason to prefer a stated top five over a blanket automation programme is not modesty. It is that a bounded backlog can be completed, and a completed thing changes what an organisation believes it can do next.

NEXT STEP

Ready to make AI real?