Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesPublic sector

Case study
Public sectorA large European public-sector organisation

Four in five agreed the service was about what they expected, and the same people split three ways on the ratio the whole operation runs on

A large European public-sector organisation reviewed the outsourced desk its own staff called for IT support, ahead of a decision about how to source it next. On the summary question, 80 percent of respondents said the service was about what they expected and the other 20 percent said slightly worse, with nobody rating it better. Asked what share of incidents the desk resolved at first level, the same respondents split into exact thirds: some, about half, most. The review traced that split to conditions rather than opinion. Data was mostly captured but comparison across it was difficult, departments used different tools with no way to bring the information together, tickets were misrouted between teams several times a month, and incidents were closed as assumed resolved without the user confirming. What was delivered was a current-state picture, a set of design principles, a market view of how such desks are bought, and recommendations. No desk was rebuilt and no ratio was recomputed inside the engagement.

80%Said the service was about what they expected; the rest said slightly worse
Client
A large European public-sector organisation
Duration
Current-state review and market study, findings and recommendations
AI · AUDITED €115MFY · APPROPRIATIONS
33 / 33 / 33Three-way split on the share of incidents resolved at first level
0Respondents who rated the overall experience better than expected
4 bandsUsed on perceived knowledge, against two on overall experience

Ask people how a service is doing and they answer against what they have learned to expect from it. Ask them what share of its work it finishes on first contact and they ought to give you the same number, because that number is a fact about a queue rather than a feeling about one. In a review of an outsourced internal IT service desk, the first question produced near-unanimity and the second produced a three-way tie.

A large European public-sector organisation had contracted out the desk its own staff called when something stopped working. Before deciding how to source that service next, it commissioned a review: performance data, process, governance and reporting, a structured round of interviews with the managers whose teams absorbed the escalations, and a market study of how desks like it were being bought elsewhere. The interviews carried a short survey on the service itself.

On the summary question, the one asking about the overall experience, 80 percent said the service was about the same as what they expected. The other 20 percent said slightly worse. Nobody said better. Read alone, that is a settled group describing an unremarkable service that mostly works.

Then the same people were asked what share of incidents the desk resolved at first level. The answers split into exact thirds. A third said most of them. A third said about half. A third said some of them.

The challenge

First-level resolution is the ratio the operation is built around. It sets how many agents you need and at what skill, what the second and third lines absorb, what a ticket costs, and whether a self-help channel pays. A third of an informed audience believing the desk clears most tickets itself, while another third believes it clears only some, means two groups are describing organisations with different cost structures and different remedies, using the same service name.

No question underneath the summary one produced anything like the same agreement. Eagerness to help drew just two answers, moderately eager and slightly eager, splitting 63 to 37. Perceived knowledge spread across all four bands, from very knowledgeable to not knowledgeable at all. Asked how often users were kept informed while an incident was open, two thirds answered sometimes; asked how often incidents came back as escalations, half answered sometimes. Sometimes is the answer a frequency question gets when nobody is counting.

The interesting question is what made that arithmetic unavailable, and the answer sits in the current-state work rather than in the survey.

Reporting was the first condition. Most data was being captured, but reporting off it was difficult and comparison for analysis was difficult, which is a precise way of saying the numbers existed and did not reconcile. Technology was the second. Different departments ran different tools with no way to bring the information together, the ticket system resisted integration, its reporting function was weak, ticket text was hard to search, and priority and severity could not be set inside it at all.

Process was the third and the most corrosive. Tickets were misrouted and passed between teams several times a month by each team. Incidents were closed as assumed resolved rather than confirmed by the user. Knowledge sat with a small number of agents for particular applications, so a single absence produced a backlog that grew until that person returned, and the same incident was escalated repeatedly because nothing was written back into a knowledge base after it was solved.

Every one of those bends the ratio in a different direction. A ticket that bounces twice and is settled by a second-line team is a first-level failure in one accounting and a closed ticket in another, and one assumed closed counts as resolved for whoever closed it. Ask a manager who sees the bounces and one who sees the closures, and both answer honestly, and they describe different desks. Our Consult engagements now open by writing that definition down, because a review that puts the ratio to a survey before defining it measures the disagreement rather than the desk.

The approach

The review did not try to settle the ratio by asking harder. It went after the conditions that made it unanswerable, writing current state and target state side by side, in words.

Ownership came first. There was a single owner of the desk, but the roles and responsibilities under it were not documented in a way anyone could point at, and the structure had grown a lot of layers. Continuous improvement was happening, but as separate initiatives launched by the organisation and by the provider in parallel. The governance rhythm itself was working: the meetings were sensibly structured and the people in them were senior enough to decide.

The design principles from the interviews are unglamorous and worth listing, because almost none of them is a tooling decision. Define the desk's scope. Raise first-contact resolution by helping the user directly. Streamline the incident and problem processes, which had grown complex enough that categorisation itself was a source of error. Install accountability for the service. Give agents easy access to information and a knowledge base worth searching. Tell agents from the onset what is expected of them. Then, last in the list and last in the logic, a service management tool that supports all of the above.

The market study is where this stops being a service desk story. Fixed-price and unit-based contracts were the two predominant ways these desks were bought, with outcome-based arrangements moving among buyers looking for value beyond cost reduction. The example given for an outcome worth paying against was a sustained increase in first-contact resolution above target. An organisation that cannot produce that ratio to a definition its own managers share cannot buy on it, cannot verify it and cannot dispute it. The commercial model available to you is limited by the instrumentation you have.

The outcome

This engagement produced findings, design principles, a market view and recommendations. No desk was rebuilt inside it, no ratio was recomputed, and no sourcing decision was taken within its scope. The strongest thing it produced was not in the performance pack. It was the discovery that a group of well-informed managers, all watching the same service, agreed about how it felt and could not agree about how it ran.

What we would do differently now is not ask. The ratio does not have to be surveyed, because every ticket already records its own route: opened, categorised, assigned, reassigned, resolved, reopened. Process mining over that event log answers the question the survey could not, against a definition written once, applied to every ticket and rerun whenever anyone wants. The precondition is data readiness rather than analytics: consistent capture, one identifier for an incident across every tool that touches it, and lineage from the number on the slide back to the events that produced it. Our Platform work starts there, and this survey is the reason it starts there.

Once the ratio is instrumented, the automation candidates name themselves. The reassignment loops, the categories misrouted every time, the population closed as assumed resolved: each becomes a measured volume with a measured cost, which is what makes a case for straight-through handling rather than a hope for it. And the early language-model pilots now being run over ticket text, drafting a first response or suggesting a knowledge article, need the same foundation. A pilot measured against a contested definition will report a win by moving tickets into a category nobody counts.

Satisfaction was never the problem. Four in five said the service was about what they expected, and expectation is a moving anchor that settles wherever a service has been for long enough. The problem was that the organisation had a mood it could describe and a ratio it could not.

NEXT STEP

Ready to make AI real?