Reviews get commissioned to confirm things. Nobody says so out loud, but the shape of the request gives it away. The service is under contract, the reports come in green, somebody senior wants an independent view before the next procurement, and what they hope to hear is that the arrangement is sound and the paperwork can be renewed with light edits.
One such review, at a large European public-sector organisation running an outsourced IT service desk, carried a line that was not that. Sitting among the recommendations on the contract's service levels is a sentence saying third-party contractors should not be used to staff the desk, precisely so that motivation and performance management have somewhere to live. It is an uncomfortable thing to write into a deliverable. It touches procurement, it touches an incumbent relationship, and it was never the review's headline: the headline was a consolidated re-tender of the wider end-user estate, and the staffing sentence rides along underneath it.
It is also the sentence with the most evidence behind it and the least support from the rest of the document, because the desk the same review holds up as the internal model worth copying turns out to be more than half contracted. Both of those are worth having in front of you at once. That is why I keep returning to this work.
What the green numbers were counting
Start with the defence of the incumbent. Speed to answer beat both the contract target and the industry comparator. Abandonment rate came in well inside a ceiling the review itself described as set too high. Service level ran above a floor the review noted had been set below the industry average. Two figures missed, both narrowly, both on the clock rather than on the outcome. On the whole, a telephone service performing close to the terms it had signed.
Now look at what those figures have in common. Each is a by-product of routing a telephone call. The gap between a call arriving and a headset lifting exists as a number before anyone decides it is a metric, and it means the same thing every month because a machine settled the definition rather than a person. Those metrics report themselves. They would report themselves on a desk with no management at all.
The rest of the scoreboard is the tell. One call resolution, the one telephony-side figure requiring somebody to have written down what counts as resolved and what counts as one call, sat at 47 percent against a target above 65 and an industry average of 75. It appears as a single month's figure rather than a run, because the supplier held no history for it, and the review flagged confusion about how it was being measured at all. Incident acknowledgement had no target in the agreement outside two email routes and no measured average anywhere. Incident resolution had targets at every priority level and not one figure to place against them. Priority itself had quietly collapsed: close to 95 percent of incidents were logged as low.
The service had a fully instrumented front door and an unmeasured interior. The pack showed the front door.
The staffing evidence, and why it never reached the pack
The maturity work behind the review is where the recommendation comes from, and it reads bluntly. Against performance management, the recorded current state was that performance is not managed. Against objectives for desk staff, the note is that this was not in the buying organisation's scope, alongside a perception that no objectives were linked to performance and a supplier statement that they existed. Against motivation: no formal training, recruitment or career progression plans, addressed reactively and only when a need arose. Against qualifications: some staff held a recognised service-management certification because they were key or because they were keen. Against career path: the supplier stated paths existed and were managed locally, and the reviewers judged that plausible for senior people only.
Read those as one item rather than five and a structure appears. The organisation paying for the service had no standing to manage individual performance, because the staff were not theirs. The supplier had the standing but no obligation, because minimum qualification levels and training requirements had never been written into the agreement. Development therefore happened when the client asked for it and escalated until it happened, which is what the record shows: customer-handling seminars and training courses both began only after repeated requests and escalations.
That is not a story about the people who worked on that desk, and the review is careful not to make it one. It is a story about an arrangement in which nobody was both able and required to build capability. When the organisation ran technical knowledge assessments of its own, results came back far below what it expected even on basic questions, and running the same test four times over, having announced that the questions would repeat, barely moved the score.
The comparator that contradicts the recommendation
Elsewhere in the same review sits a short study of another service desk inside the same organisation, a non-IT one, offered as evidence of what good looks like in-house. Its staffing is recorded plainly: 10 permanent staff alongside 11 from a contract supplier. More than half the heads on the desk held up as the model are contracted, on a desk the review recommends copying, in a document that recommends not staffing this way.
I do not think that makes the recommendation wrong. I think it locates it correctly. What the review actually credits the comparator desk with is not its employment mix. It is a commercial construction: a service-credit regime where points are deducted per error, and a service-level bonus at review that is split between the supplier company and the individual agents rather than stopping at the corporate boundary. Objectives reach the person. The reward reaches the person. Employment status never had to.
So the defensible version of the recommendation is narrower and harder than the sentence on the page. Do not accept a staffing arrangement in which no party is both able and obliged to develop and manage the individual. Contract staffing is one way to end up there, and it is the common way, which is why the shorthand exists. It is not the only way, and the same document contains the counter-example.
The users had already said it, quietly
The stakeholder work is the part I would put in front of a sceptical executive, because it shows how a service can be compliant and unloved at once.
Asked about their overall experience, four in five said the service was about what they expected and the remaining fifth said slightly worse. Those two answers account for the entire sample. On perceived knowledge, responses landed on all four answers offered, from very knowledgeable to not knowledgeable at all. Asked how many incidents were resolved at first contact, the accountable group split into even thirds: most of them, about half of them, some of them.
That last split is the most useful line on the page. When the people accountable for a service disagree evenly about its central operating ratio, that ratio is not published anywhere they can check. A distribution with nothing above "as expected" in it is the signature of a service that satisfies its contract and nothing beyond it.
Underneath, the specifics were about capability and ownership rather than speed: incidents categorised inconsistently, escalations arriving without enough information to work with, tickets moving between groups two or three times a month per team, incidents closed on assumption rather than confirmation, knowledge concentrated in two individuals so a single absence created a backlog, and no quality checks or audits on incidents at all. Not one of those shows up on a telephony scoreboard.
Why this argument matters more now than it did then
We are all putting copilots on service desks. Retrieval over the knowledge base, drafted responses, suggested categorisation, an agent loop that takes a low-risk request end to end under graded autonomy with a human confirming anything above the line. I think that is right, and RealAI's Agentic OS work is largely about making those loops governable rather than merely demonstrable.
Automation does not repair the structure it lands on. It inherits it.
An agent loop needs the same things the human tier needed and did not get. Someone whose objectives include its performance. An evaluation set built from cases a named person defined, which means somebody has to settle what counts as resolved before a single case can be scored. Quality checks and audits on its output, on a desk where none were carried out on human output. A categorisation field that means something, on a desk where 95 percent of incidents defaulted to low priority. And under the EU AI Act now in force, that accountability chain written down rather than assumed, with the human review step resourced by people equipped and motivated to perform it.
Put an agent loop on a desk where performance is not managed and you get an unmanaged loop. Its front-door metrics will be excellent, for the same reason the telephony metrics were: deflection and handling times are by-products of the machinery. Its resolution quality will be unmeasured, because measuring it requires a definitional decision nobody owns. You will have bought a faster version of the same blind spot, and the pack will be greener than ever.
A scoreboard assembled from what a telephone switch counts on its own answers the question a telephone switch can ask. The staffing question never reaches it, because no instrument on the page is pointed at it.
What the review recommended, and what to carry forward
Four sourcing options were assessed against six design principles the organisation had already written down: renegotiate in place, insource, re-tender the desk alone, or re-tender a consolidated end-user support scope. The two re-tender options scored at or above the other two on every principle, and consolidation led outright on five of the six. Insourcing, which is the option that would have removed third-party staffing from the desk entirely, scored lowest overall and took the grid's only zero, on value for money.
That is worth sitting with, because it is the grid arguing against the staffing sentence rather than for it. Pricing pulled the same way. The contract ran on volume-based pricing with nothing in it that rewarded exceeding a target, at a cost per incident ranging between 38 and 49 euros over the recent window, sitting near the top of that range in the latest month and higher than the market. Separately, the service levels carried exceptions and volume thresholds that made achievement easier, and incidents raised by email or self-service portal fell outside them altogether.
The changes worth carrying forward have nothing to do with any one supplier. Write minimum qualification levels and training obligations into the agreement rather than requesting them later. Make training outcomes measurable through a formal skills assessment. Define acknowledgement and resolution targets by severity and measure them, so the interior of the service has numbers at all. Avoid exclusions and volume thresholds that make a target easier to hit than the service is to deliver. Reward exceeding an outcome and penalise missing one, and push that reward far enough down that it reaches the individual, which is the mechanism the internal comparator desk actually ran on. And treat any staffing arrangement where nobody can develop or manage the individual as the defect, whatever the contract calls it.
None of that requires a model, a pipeline or a platform, and all of it has to land before any of those are worth buying. Reading a review against itself, which is what the comparator desk forces here, is where our Consult work on a sourcing decision starts.
Figures and findings are as recorded in a current-state review of an outsourced IT service desk at a large European public-sector organisation: its contract scoreboard, a maturity assessment, stakeholder interviews, an internal comparator study and a sourcing options analysis. That review produced findings and recommendations ahead of a re-sourcing decision, not delivered results, and the option scores are judgements of attractiveness rather than outcomes. Reading it as a governance problem for automated service tiers, and reading its staffing recommendation against its own comparator, are ours.
“A scoreboard assembled from what a telephone switch counts on its own answers the question a telephone switch can ask. The staffing question never reaches it, because no instrument on the page is pointed at it.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
