Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesIT service management

Case study
IT service managementA large European public-sector organisation

A desk resolving 47 percent at first contact does not lose the rest, it creates them again somewhere else

A current-state assessment of an outsourced IT service desk at a large European public-sector organisation put one quantified line among nine listed symptoms: first call resolution at 47 percent, against a contract target above 65 percent and a cited industry figure of 75 percent. The other eight symptoms read as complaints and work as mechanisms, each one a route by which a single user need becomes two or more recorded contacts, including duplicate incidents created when a user rang to chase the incident they already had. The retained second-line groups reported absorbing routine work the desk could have closed. What the review never recorded is a contact volume, a ticket count or a cost per contact, so the multiplier is visible and unsizeable in the same document. The deliverable is a diagnosis, a market study and four sourcing options, not a result.

47%First call resolution, against a contract target above 65%
Client
A large European public-sector organisation
Duration
Current-state assessment, findings and sourcing options
AI · RIDGE E45.5 N25ρmax 1.00
75%Industry figure it was measured against, cited with no source
18+ ptsBelow the desk's own contractual floor
0Contact volume figures anywhere in the summary

A desk that resolves 47 calls in every 100 at first contact has not lost the other 53. It has created them again. They come back as a call to a second-line engineer, a ticket reopened, a colleague asked in a corridor, or a fresh incident logged by the same person chasing the one they already raised. That figure is not a quality score. It is a multiplier on the volume of everything downstream of the phone.

The number sat on a single line of a current-state assessment of an outsourced IT service desk run for a large European public-sector organisation: first call resolution measured at 47 percent, against a contract target above 65 percent and a cited industry figure of 75 percent. The desk was the first point of contact for roughly 8,000 users across several of the organisation's sites.

The challenge

That line was the only quantified item on the page. Beside it sat eight further symptoms, all of them qualitative: an average or patchy user experience, incidents not categorised properly, inadequate and inconsistent information in escalations to other groups, misrouted tickets passed around and across teams, incidents resolved with unsatisfactory or inappropriate resolutions, critical incidents not detected quickly enough and escalated slowly, ping-pong between the desk and other support teams, and duplicate incidents with no means of checking them.

Read as a list of complaints, those eight are a service-quality anecdote. Read as mechanisms, seven of them name a distinct route by which a single user need turns into two or more recorded contacts, and the eighth, the patchy experience, is the judgement the other seven add up to. A misrouted ticket is handled twice. Ping-pong is handled repeatedly. An inappropriate resolution closes a ticket and produces a phone call. The most literal example is the duplicate one: the review recorded that a user contacting the desk to check an update on their own incident was logged as another incident rather than as a diary entry against the existing one.

Where that work went was visible on the previous page. The organisation's retained second-line groups reported being increasingly tasked with routine desktop and application issues the desk could plausibly have solved, and the pressure on technical and application support was described as reaching the lines of business. The eighteen or more points by which the desk sat below its own contractual floor did not evaporate. What the review has in place of a measurement is the retained groups' own account of where the work landed, absorbed by engineers retained for something else, arriving without a ticket saying where it had come from.

Then the part that should stop anyone reaching for a calculator. Nowhere in the summary is there a contact volume, a ticket count, an agent headcount or a cost per contact. The incumbent arrangement across two vendor contracts was costed at EUR 2.7 million, roughly 8,000 users are named, and no denominator joins them. The multiplier is real and it cannot be sized from this document. The same page also explains why any volume figure would be untrustworthy if it existed: when chasing your own ticket creates a second ticket, a count of tickets is partly a count of the failure to resolve them.

The approach

The diagnosis was structured with unusual restraint. Nine observed symptoms sat in one undifferentiated list. Against them, fifteen root causes grouped under people, process, technology, and governance and reporting. Not one line was drawn between any symptom and any cause.

That refusal is the craft. Ticket ping-pong in a desk this size comes from undefined handoffs, unclear accountability and a stale knowledge base at once, in proportions nobody can honestly state. An arrow from one symptom to one cause is a precision claim, and a precision claim is what gets budgeted against.

The causes are the useful inventory. The contract defined no minimum skill set for an agent before that agent could take live calls. The known error knowledge base was updated irregularly, so the same issues escalated again and again. Text inside incidents was not searchable, and the knowledge base auto-update did not work. The service management tool was hard to integrate with the tools the other support groups ran, so every escalation crossed a data boundary. Operating level agreements with those groups were not defined. Reporting requirements and review cadence were not written into the contract at all.

One of those has to be read against the headline number. The same set of causes states that reporting functionality was poor and that service level reports could not be generated. The 47 percent therefore came from somewhere other than the system of record, and the document does not say where. The 75 percent industry figure carries no source anywhere in it. The gap is almost certainly real and its direction is not in doubt, but publish it with its provenance attached, and notice that the first recommendation on measurement was to reset targets to industry standards. That is a new reading proposed on an instrument the same review had already declared broken.

The recommendations followed the causes rather than the symptoms. Ownership of a ticket through to closure, written into the contract as the desk's responsibility. A contractual minimum qualification for agents, assessed by the client, with remediation of anyone falling short at the supplier's cost. Operating level agreements at each handoff. Pricing made symmetric, with reward for exceeding targets where the client gains and penalty for missing them. Users tiered by business impact, with service levels and penalties differentiated per tier. And on automation, a short list: automated password reset in the voice menu, identify and automate the most frequent request types, script the fixes for common issues.

That last list is where a reader today should slow down rather than speed up. Retrieval-augmented generation over a knowledge base whose auto-update never worked, sitting on incident text that is not searchable, retrieves nothing worth retrieving. A copilot drafting agent responses needs an evaluation set, and the measure it would be scored on cannot be produced by the tool it would run inside. Data readiness and lineage are not preliminaries to a programme like this. They are the programme, and the vector store is the cheap part of it. Our Platform work opens at exactly that point, auditing what the ticket history actually contains, because a model pointed at a corpus that records the dysfunction learns the dysfunction. The autonomous experiments worth running here are narrow and carefully scoped by design: one request type, one measured outcome, one way to hand back to a human.

The outcome

What this engagement produced is a decision document, not a result. A current-state diagnosis, design principles for the future service, a market study, four sourcing options and a recommended way forward. No desk was rebuilt inside it, no service level moved, and no resolution figure was measured afterwards. Everything above is a finding or a recommendation.

The costing carried the same asymmetry the multiplier suffers from. The two options the review argued against carried numbers. The incumbent arrangement was priced at EUR 2.7 million across the two vendor contracts, with no period stated for it. Bringing the desk in house was estimated at EUR 4.3 million on client-supplied day rates, with travel, recruitment, technology, training, bonus, overtime, depreciation and amortisation all excluded, which makes it a generous estimate that loses anyway. The two retender options carried no figure at all, and one of them was the recommendation. A business case that prices the things it rejects and describes the thing it proposes will always read as though the proposal won on the numbers.

The honest close is that the most consequential number in this review is one nobody took. Not the 47 percent, which was contested and unsourced. The count of second contacts it generated: the reopens, the escalations a first-line agent could have closed, the calls to a colleague that never became a record, and the duplicate incidents the desk manufactured out of its own follow-ups. Our Consult engagements now open by measuring that, before anyone argues about a target, because it is the cheapest way to turn a service score into a number a finance director recognises.

NEXT STEP

Ready to make AI real?