Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesIT service management

Case study
IT service managementA large European public-sector organisation

Six request types was not a technology limit, it was the point where a per-case build stopped paying for itself

A large European public-sector organisation had its outsourced IT service desk reviewed ahead of a re-sourcing decision. The automation recommendation was precise and small: automate password reset in the phone system, identify and automate the five most frequent service requests, and use scripted automation for common fixes. Six things, in a service where 844 incidents were pushed past the first line in a single sampled month, about a fifth of everything logged, and where one call in two was not closed on the call. The gap between the six and the rest was never a question of what a machine could do. It was that every additional request type carried its own build, its own integration and its own maintenance, while the volume behind each one fell away steeply. Retrieval over a maintained knowledge base changes what the next case costs to cover, which moves the boundary from what is worth building to what is worth writing down and testing. What was delivered here was a diagnosis and a set of recommendations, not an automated request.

844Incidents pushed past the first line in one sampled month
Client
A large European public-sector organisation
Duration
Current-state service desk review, findings and recommendations
AI · RIDGE E72.7 N62.5ρmax 1.00
4Request types the contract named as fulfilment scope
47%One-call resolution against a contract floor of 65 percent
5 min 45sAverage talk time against a contract ceiling of five minutes

Every automation programme has a frontier, and almost nobody writes down what sets it. Ask why a service desk automates six things rather than sixty and the answer that comes back is usually about technology maturity, or user readiness, or risk appetite. It is none of those. It is arithmetic. Each additional request type carries a build, and the volume behind each one falls away much faster than the build cost does.

A review of the outsourced IT service desk at a large European public-sector organisation states the frontier out loud, in a recommendation that names quantities. Automate password reset inside the phone system. Identify and automate the five most frequent service requests. Use scripted automation for fixes of common issues. Six things, plus a general instruction. Read as a technology list it is unremarkable. Read as an economic boundary it is the shape of what was reachable at the time, and the boundary is worth tracing precisely, because it is the thing that has moved since.

The challenge

Start with what the desk was contractually asked to fulfil. Four request types: information or advice, a standard change, access to an application, and a password reset. Four. Everything else a user needed arrived as an incident, was diagnosed by a first-line agent on the telephone, and was either closed on the call or pushed onward.

Onward was busy. In the month the review sampled, 844 incidents went to second and third line groups, roughly a fifth of everything the desk logged. One-call resolution ran at 47 percent against a contract floor of 65 and against the 75 percent the review cites as the industry figure, so slightly more than half of contacts were not finished by the person who took them. Average talk time sat at 5 minutes 45 seconds against a contract ceiling of five minutes, and the review notes the uncomfortable pairing directly: the longer conversations were not buying a higher resolution rate.

None of that is a desk answering badly. Average speed to answer was 17.42 seconds against a target of under 25 and an industry figure of 30. The desk picked up quickly, talked for a while, and then handed a large share of the work to somebody else. That distribution is the point. There is a head of frequent, shaped, repeatable requests, and behind it a long tail of things that each happen rarely and each need somebody who knows the application.

Now put the recommendation back against that picture. Automating password reset and the five most frequent requests addresses the head and stops. Not because case number twelve was harder, but because case number twelve had a twelfth of the volume and the same build in front of it: a script to write, an integration into the ticket tool, a path through the phone system, a test, an owner, and a maintenance liability every time the underlying application changes. Divide a fixed build by a falling volume and the frontier draws itself. Six was where the line fell.

The review says as much in its own technology section without framing it that way. It asks for the ticket tool to be configured so that knowledge base and incident model updates become simpler and easier. That is a request to reduce the cost of adding the next case. It was the right instinct and it was aimed at the wrong constraint, because making a hand build cheaper by a third still leaves you with a hand build.

The approach

What actually moves a frontier like this is a change in the cost of covering case number twelve, and that is where the current generation of language technology earns its place rather than where the demonstrations usually put it.

The shape is unglamorous. Retrieval-augmented generation over a single maintained knowledge base and the accumulated ticket history means an additional request type is covered by writing the article and assembling an evaluation set for it, not by commissioning a build. A copilot sitting beside the first-line agent does not need a separate integration per request type, because the retrieval layer and the vector store are shared across all of them. The marginal case becomes a content and evaluation exercise. That is a different curve, and a much flatter one.

It also relocates the bottleneck rather than removing it, and the review is unusually good evidence for where the new bottleneck lands. This organisation was running two knowledge bases side by side. The supplier maintained one and the review found it outdated and infrequently updated. The client maintained the other through a separate support team, kept it in a document-collaboration site, and had not integrated it with the ticket tool, so agents worked across two windows. Only simple reports on telephony and email could be produced, and those by hand.

Read that list again with the new economics in mind. Every one of those findings is now a direct cost. Retrieval over a stale knowledge base returns stale answers with a confident tone. Ticket history that cannot be queried cannot tell you what the tail actually contains, which means you cannot rank it, which means you cannot decide what to cover next. The data readiness work that looked like hygiene under the old model is the capital expenditure under the new one, and lineage from an answer back to the article it came from is what makes any of it defensible to an auditor. That is the part of the Platform work we do that clients are least prepared for and that costs the most.

The narrow cases still deserve their own treatment. Password reset was named separately from the top five for a reason of shape rather than frequency: a single state change behind an identity check, no judgement between the request and the outcome, an entitlement a machine can verify, and a failure that is loud and immediate rather than quietly wrong. That profile is where an early and carefully scoped autonomous experiment belongs. The rest of the tail belongs, for now, beside an agent rather than instead of one, and the honest reason is that a wrong answer to a rare question is expensive precisely because nobody on the floor has the experience to catch it. Anyone selling the reverse ordering has not looked at an escalation log.

There is also a governance clock running. The European rules on artificial intelligence are in force, with the obligations landing in stages, and a public-sector body automating decisions that touch its own staff will be asked to show its documentation, its evaluation evidence and its human oversight arrangements. Building the evaluation sets now because they make the system better is considerably cheaper than assembling them later because somebody asked.

The outcome

This engagement produced findings, recommendations and sourcing options. No automation was built here, no request type was retired, and the automation lines sat inside a long recommendation list in a review whose centre of gravity was a re-sourcing decision. Six named targets against a desk that sent 844 incidents past its own first line in the sampled month is not a plan to remove the tail. It was never presented as one.

What is worth carrying forward is the reason the number was six. The frontier was set by a per-case build cost, and any organisation that still funds automation case by case is drawing the same line today with better tools behind it. Our Consult work now starts by ranking the tail from the ticket history and pricing the marginal case twice, once as a build and once as an article with an evaluation set beside it. The two numbers are rarely within an order of magnitude of each other, and the difference between them is the whole argument.

NEXT STEP

Ready to make AI real?