Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesIT service management

Case study
IT service managementA large European public-sector organisation

844 incidents escaped first line in a single month, and they escaped to the same few places

A large European public-sector organisation commissioned a current-state review of its outsourced IT service desk. In the month examined in detail, 844 incidents were escalated to second and third level, about a fifth of everything the desk logged. Speed to answer, service level and abandonment rate all sat inside contract targets. One call resolution sat at 47 per cent against a contract target above 65. The escalations concentrated on the organisation's own applications, the exact scope a dedicated second-level expert arrangement was supposed to cover, and on monitoring-generated tickets that were forwarded wholesale. What the engagement produced was a baseline, a process assessment and recommendations for the next contract, not a reduction in escalations.

844Incidents escalated past first line in a single month
Client
A large European public-sector organisation
Duration
Current-state review, findings and recommendations
AI · RIDGE E45.5 N37.5ρmax 1.00
~20%Share of logged incidents that escaped first line
47%One call resolution, against a contract target above 65%
94.9%Service level achieved, against a target above 78%

A service desk can hit nearly every number in its contract and still fail at the only thing it exists to do. Speed to answer, abandonment rate, service level: each of those measures how quickly a call is picked up. None of them measures whether the person who picked it up can finish the job.

A large European public-sector organisation commissioned a current-state review of its outsourced IT service desk. In the month examined in detail, 844 incidents were escalated to second and third level. That is roughly one in five of everything the desk logged. The reviewers then did the thing that turns a volume figure into a decision: they cut the escalations two ways, by the group that received them and by the category of incident, and looked at where they piled up.

They piled up. The single largest destination was the second-level team supporting the organisation's own applications, which was the precise scope a dedicated expert-support arrangement had been contracted to cover. The review's judgement on that arrangement was blunt: it had not been effective.

The challenge

The front-of-house numbers gave nobody a reason to look. Average speed to answer over the three-month window ran at 17.42 seconds against a contract target under 25 seconds. Service level, the share of calls answered inside 30 seconds, came in at 94.9 per cent against a target above 78. Abandonment rate was 2.06 per cent against a ceiling of 8. Two measures sat slightly outside: average talk time at 5 minutes 45 seconds against a target under 5 minutes, and abandonment time at 1 minute 32 seconds against a target under a minute.

One number was not close. One call resolution measured 47 per cent against a contract target above 65. The review also recorded confusion about how one call resolution was being measured, which is a finding in its own right: a headline outcome metric whose definition is unsettled is not a metric, it is a standing argument.

The priority field told a similar story. Prioritisation was largely automatic, suggested by the service-management tool from the recorded impact, and close to 95 per cent of incidents came out low priority. A field where almost every row carries the same value is not data. It is a default that nobody has had a reason to override.

So the desk was fast, available, and unable to finish. The interesting question is not why one call resolution was low in aggregate. It is where the failures to finish were concentrated, because that is the only version of the question with an action attached to it. Aggregate first-contact resolution tells you to train everyone about everything. An escalation map tells you which five queues to fix.

Two clusters carried most of the story. The first was the organisation's own applications, where the escalations that were meant to be absorbed by a specialist tier were instead arriving there in volume, because first line could not hold them. The second was monitoring. Monitoring-generated incidents were being passed to the monitoring team almost as a rule, and the review declined to read that as legitimate routing. It read it as an absence of proactive communication between the monitoring function and the desk: the desk was acting as a relay because nobody had told it what the alerts meant.

The approach

The analysis was deliberately unglamorous, and its discipline is worth copying. Escalations were counted by receiving group, which answers who absorbs the failure, and by incident category, which answers what kind of question the desk could not answer. The charts published their own noise thresholds: receiving groups below ten escalations were bucketed together, categories below five were left off. Publishing the threshold is what makes a chart like this auditable rather than persuasive, and it is the step most internal reporting skips.

Alongside the numbers, the process review found the mechanism. There were two knowledge bases in play. The provider maintained its own, which the review found was not being updated frequently and had gone out of date. The organisation maintained a separate one, kept current by a different support team, held on a document-sharing site and not integrated with the service-management tool. An agent on a call therefore had no single window: the current knowledge lived in one system and the ticket lived in another, and the review recommended integrating them.

The organisation had already tried the direct approach. It ran quizzes to test and raise agent knowledge of its own applications. The questions were repeated across sittings and communicated in advance, and it still took close to six months for scores to show a minor increase. The review offered two explanations, both about people: motivation on the floor, and local management capability.

Both may be true. The escalation data supports a third that costs less to act on. If the current knowledge is held outside the tool the agent works in, repetition is the wrong instrument. You are asking people to memorise what the system should be handing them. The desk's incident close-out process points at the same gap from the other side: until shortly before the review, the desk did not retain ownership of a ticket after hand-off to second line, and there were no follow-up calls back to the user to confirm the fix had actually held. So the resolutions that second line produced never travelled back to first line in any form anybody could search. Every escalation was a lesson that left the building.

The outcome

This engagement delivered a review. It produced a performance baseline against the contract, a process and governance assessment, an escalation analysis, and a set of recommendations for the next service desk contract, including making the knowledge base integration and the reporting automation contractual rather than aspirational. No reduction in escalations was delivered in this phase, and the numbers above describe a month of operations, not an improvement.

Three things would be done differently if the same review ran now, and none of them are exotic.

Escalation concentration should be a standing measure, not a monthly slide. Process mining over ticket events reconstructs the path each ticket actually took, first-line attempt, escalation, hand-back, closure, and ranks destinations by both escalated volume and time spent after escalation. That second dimension matters, because the queue with the most escalations is not always the queue costing the most. The Platform work we do starts at exactly this point, with the event log rather than with a model.

Knowledge base integration should be specified for lineage, not for convenience. If a resolution records which article the agent used, every article acquires a hit rate and a decay curve, and the ones that are quietly wrong become visible. Without that link, a knowledge base is a folder that people are told to like.

And the data has to be readied before any of it can carry a model. A pipeline trained on a priority field where 95 per cent of rows share one value will learn nothing worth deploying. Closure categories were already intended as the basis for root cause analysis, which makes them the field to clean first. The current wave of cautious language-model pilots in support desks changes the arithmetic less than the vendors suggest: retrieval into the agent's window over an outdated knowledge base returns outdated answers faster. Our Consult engagements start by testing whether the content is current, because that is the cheapest test available and it decides everything downstream.

The finding here is not that a fifth of tickets escaped. It is that they escaped to the same few places, in plain sight, in data the organisation already owned. Where escalations cluster tells you where the knowledge is missing.

NEXT STEP

Ready to make AI real?