Almost every assessment of a support function starts by looking outward. Someone pulls a cost per contact from an analyst house, a first-contact resolution figure from a vendor's published material, a staffing ratio from a peer institution in another country. The comparison is then presented to an executive committee, which asks the only sensible question available to it: are those organisations anything like us?
They usually are not, and everybody in the room knows it, so the benchmark carries less weight than the effort that produced it.
I keep coming back to one current-state review of an outsourced IT service desk at a large European public-sector organisation, because the research design did something I have rarely seen since. It interviewed the desks in the same building.
The cohort nobody usually interviews
Set the stakeholder list beside a standard benchmarking exercise and the difference is structural rather than cosmetic.
Three of the four cohorts are the ones anyone would pick. The application groups across the business areas, who raise tickets and live with the outcome. The heads of the departments behind the desk, running infrastructure, platforms, operational security, service integration, asset and configuration management. The senior management who own the contract and answer for it.
The fourth cohort is the interesting one: directors of other service desks in the same organisation, sitting outside the IT organisation. Service functions with their own users, their own queues, their own contracted staff, their own definitions of an acceptable wait.
They have no view on ticket categorisation and were not asked for one. What they have is a running answer to the question the review actually needed settled: what a well-run desk looks like inside this institution, under its constraints.
What the internal comparator bought
The clearest illustration of what an internal comparator is worth sits in the assessment of agent knowledge and motivation.
The organisation had been running its own knowledge checks on the outsourced agents, covering the applications specific to its work. The results were poor and stayed poor. The same questions were reused, four rounds in a row, with the agents told in advance that the questions would repeat. Correct answers barely moved. It took close to six months for even a minor improvement to show.
That finding invites a training recommendation, and a training recommendation is what an external benchmark would have produced. Industry practice includes satisfaction and motivation indicators in service desk contracts, so add them. Perfectly true, and perfectly easy to defer, because nothing about it says the institution can do it.
The review went the other way. It recorded that a similar contract, one carrying performance-related bonuses, already existed at another service desk in the same organisation, and recommended evaluating that contract for the IT desk. Not a target, not an outcome, not a saving. A recommendation with a live local precedent attached.
The difference between those two recommendations is entirely about the room they are read in. The first asks an executive to believe a practice observed elsewhere transfers here. The second asks an executive to accept that something already signed, already administered and already surviving inside their own institution could be signed again a corridor away. One is a matter of judgment. The other is paperwork.
The evidence an average cannot carry
The rest of the stakeholder material shows why the internal view was worth the interview time. Almost none of it would survive being compressed into a benchmark.
On overall experience the interviewed stakeholders were settled and unremarkable: four in five described the service as about what they expected, the remainder slightly worse, nobody better. That is a service nobody is escalating and nobody is defending, and exactly the profile an external comparison renders as adequate and moves past.
Underneath it, the qualitative material is specific and awkward. Incidents closed on assumption rather than on confirmation from the user, which one stakeholder objected to in plain terms. Tickets passed between teams two or three times a month by each team. Escalated tickets arriving at second and third level with too little information to act on, so the receiving group starts again. Knowledge concentrated so narrowly that a family of the institution's own applications was handled by two agents, and a backlog formed when one of them took leave. The same incident escalated repeatedly because nothing was written back into a knowledge base after it was solved. No quality checks or audits on incidents, and no accountability attached to one.
Meanwhile close to four in five contacts arrived by telephone, self-service accounted for a fraction of a percent, and roughly one logged incident in five left the desk for a specialist group.
Asked how many incidents were resolved at first contact, the interviewed stakeholders split evenly into thirds: most of them, about half of them, some of them. On eagerness to help, the answers divided between moderately eager and slightly eager, with nobody at all reaching for the top of the scale.
None of that is a number you can benchmark. All of it is a description of where the work actually goes, and the only people who can give it to you are people inside the institution who are close enough to the service to have opinions and far enough from it to say them.
- Four in five
- Stakeholders describing the service as about what they expected, none better
- Thirds
- Even split among interviewed stakeholders on how many incidents resolve at first contact
- Roughly 1 in 5
- Logged incidents escalated out of the desk to specialist groups
- No information available
- Agent utilisation, never written into the contract as an indicator
Where this lands for retrieval and copilots
The reason I am writing this up now rather than filing it as an operations lesson is that the same design question is in front of every organisation currently deciding where to put its first assisted-support capability.
A service desk is the obvious candidate. Retrieval-augmented generation over an incident history and a knowledge base, a copilot beside the agent rather than in front of the user, surfacing the closest prior resolutions. The technology is not the hard part any more.
The hard part is the substrate, and the review describes it precisely. Two knowledge bases ran in parallel: the supplier's own, which was outdated and not being refreshed, and the organisation's, maintained separately and not integrated with the ticketing tool at all. Categorisation was applied by whoever took the call, with no audit behind it. Reporting ran on two tracks with no declared source of record, and different departments used different tools with nothing pulling them together.
Point a vector store at that estate and it will index the stale corpus as confidently as the maintained one. Build an evaluation set from resolved tickets and it inherits a category field nobody governed. The lineage question, which knowledge article a suggested answer came from and whether anyone still stands behind it, has no answer in that estate, and it is the question that decides whether an agent trusts the suggestion twice.
So the deployment question is not which function has the most tickets. It is which function has a maintained corpus, a definition of resolved that two people would apply the same way, and a manager who already measures something. That is a question about the inside of the building, and the review shows you how to answer it: go and interview the desks that are not yours.
An industry average tells you something is possible somewhere. A desk two floors away tells you it is possible here, under your procurement rules, your staff regulations and your budget cycle, and it has already survived every objection your executives are about to raise.
How to run the internal pass
The mechanics are unglamorous, which is why they get skipped.
List every function in the organisation that answers unscheduled requests from a queue. Payroll, records, procurement support, an internal helpline. Most institutions have more of these than they think, and they almost never compare notes because they report through different management lines.
For each one, ask four things. What do you count, and who wrote the definition down. What is in your contract that you would sign again. What do your people actually search when they do not know the answer. What would break first if volume doubled next month.
Then read the answers as a shortlist. The function with a maintained corpus and a settled definition of done is where an assisted-support capability shows a result inside a quarter. The function without either is where the same capability produces a convincing answer to the wrong question, and the pilot is quietly stood down without anyone learning why.
None of this needs a model, a vector store or a procurement exercise, and all of it has to land before any of those is worth funding. A RealAI Platform engagement opens on exactly this pass, because the answer decides whether there is anything underneath worth building on. It also fits the direction European rules are travelling: the EU AI Act's obligations have not landed yet, and organisations that can say where a suggested answer came from and who maintains it will find that conversation considerably shorter than those that cannot.
The instinct to look outward is not wrong. It is just expensive, slow and easy to dismiss. The comparator that changes a decision is usually already on the payroll, running a queue nobody in IT has ever asked about.
Drawn from the stakeholder-interview and current-state sections of a review of an outsourced IT service desk at a large European public-sector organisation. That review produced findings and recommendations ahead of a re-sourcing decision, not delivered results; the internal comparator contract it points at existed, and the proposal to evaluate it for the IT desk was a recommendation. The market-study and vendor sections of the same pack are third-party reference material and are deliberately not drawn on here, and no score or target from the review's maturity assessment appears anywhere above. Reading the interview design as a template for deployment assessment is ours.
“An industry average tells you something is possible somewhere. A desk two floors away tells you it is possible here, under your procurement rules, your staff regulations and your budget cycle, and it has already survived every objection your executives are about to raise.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
