A performance conversation needs a number both sides will accept. The review of an outsourced IT service desk at a large European public-sector organisation recorded that performance was not being managed. Not managed poorly. Not managed. No figure anywhere in the reporting belonged to a person, so an objective had nothing to attach to, a bonus had nothing to trigger on, and a training plan had no way of knowing who was weak at what.
Read as a people finding, that invites a culture programme. Read against the organisation's own contract scoreboard, it stops being about people at all. The scoreboard shows exactly where measurement stopped, on a line nobody decided to draw.
The challenge
The agreement carried two kinds of commitment. One kind falls out of the telephone system as a by-product of connecting the call. The other needs a human being to have written down what the words mean.
Every commitment of the first kind reported a figure. Speed to answer, 17.42 seconds against a contract target under 25 and an industry average of 30. Service level, the share of calls answered inside 30 seconds, 94.9 percent against a target above 78. Abandonment rate, 2.06 percent against a ceiling of 8. Two of the five ran the wrong way and still reported cleanly: talk time at 5 minutes 45 seconds against a target under 5 minutes, abandonment time at 1 minute 32 seconds against a target under 1 minute. Three comfortable, two missed, five present.
The counting went past the targets. In one month the switch logged 10,029 calls, and the review could say answer times rose that month because the desk was short of a couple of agents. Volume, staffing and answer speed lined up because all three had numbers.
Now the other kind.
One-call resolution carried an industry average of 75 percent and a contract target above 65. The figure in its cell was 47 percent, tagged April, while the row it sat in was headed as the three-month average for January to March. The review noted in its own words that there was confusion around measuring one-call resolution. The one telephone metric that depends on a human judgement about whether the caller's problem went away is the one that came back contested and carrying the wrong month.
Incident acknowledgement, across all five priority levels: no contract target on any level, no measured actual on any. The panel reads NO INFORMATION AVAILABLE.
Incident resolution is worse, because here the contract did its half of the job. A target sits on every level: one hour at the top, then four hours, twenty-four hours, one week, two weeks. Every measured actual against those five came back as nothing, and the panel reads NO INFORMATION AVAILABLE again.
Then agent utilisation, the measure closest to an individual's productivity. It was in neither the service levels nor the KPIs, its panel returned NO INFORMATION AVAILABLE too, and the recorded consequence was that staffing could not be flexed against demand.
Put the halves side by side. This estate had not stopped capturing data. It was capturing continuously, on the phone side, in a system nobody had to maintain a definition for. Everything that needed somebody to write down what counted as acknowledged, as resolved, or as an hour of productive work was blank. The difference is definitional, not technical.
One individual measure did exist, and the buyer had to build it outside the contract. The organisation ran technical knowledge tests on the desk's agents. Results came back far below expectation on questions the review calls very basic, and when the same test was repeated four times in a row, with advance notice that the questions would be reused, the number of correct answers barely moved. Two readings were recorded: the agents were not motivated to improve, and local management was not able to drive performance. Both are fair, and neither could be acted on, because the result fed no objective, no bonus and no training plan.
The approach
The review put the contract scoreboard next to stakeholder interviews and a user survey, and the three converged. The interviews returned behaviour rather than technology. Incidents closed on the desk's own assumption rather than on the user's confirmation, which one interviewee named as the thing they hated most. Agents more eager to close a ticket than to fully resolve the issue. Knowledge silos, with two agents between them holding an application area, so a backlog built whenever one took leave. No quality checks, no audits, no accountability attached to an incident. Four in five survey respondents called the experience about the same as they expected, and the rest called it slightly worse.
None of that is a character defect. It is the rational response to an operation where the only visible unit of work is a ticket that opens and closes. When closure is the only thing counted, closure is what improves.
There was a commercial reason the gap had persisted. The agreement bought a service, not the management of the people delivering it, and not one line of the scoreboard resolves to a person. So whether the desk's agents worked to objectives of their own was a matter of assertion rather than record. The buyer's people believed there were none; the provider said there were, at every level; and nothing either side could open would settle it. That is itself the finding: an obligation nobody can observe is an obligation nobody can enforce.
So the recommendations went to the contract rather than to the culture. Define acknowledgement and resolution explicitly in the agreement. Put agent utilisation in as a named KPI, because it is the honest measure of productivity and it moves inversely to cost per incident. Motivation surveys and performance-related pay are ordinary contents of a service desk contract, and a comparable arrangement already existed on another desk in the same organisation. Rebuild the accountability chart on the future-state processes, with dedicated incident and problem management roles facing the provider's desk manager, a knowledge manager, and an audit team checking service level adherence.
The stakeholders had named the same thing unprompted. Their design principles for a future desk included accountability for the service provided, agents knowing what is expected of them from the outset, and a working environment with motivation, incentives and development prospects. All three need a measure nobody had.
The outcome
What was delivered was a current-state picture and recommendations ahead of a re-sourcing decision. Nothing was rebuilt in this phase, and no number above is a result. The value is in the ordering: almost every improvement anyone proposed for this desk depends on a measure that did not exist, so the measure is the first build rather than the last.
That build is ordinary engineering, and it is the work RealAI's Platform is for. One incident record, one set of definitions for acknowledgement, resolution and closure, and the second and third level groups reporting into it rather than alongside it. The monthly reporting then becomes a pipeline: resolution time by priority, reopen rate, rework by team. A management metric is a feature like any other, and a feature needs a definition, an owner and lineage back to the event that produced it. That is the discipline a feature store exists to impose, and it is why data readiness here is a definition question before it is a volume question.
Once those numbers exist, the people work becomes possible rather than rhetorical. Objectives can be set against something both parties can observe, training stops being ad hoc when the record shows which categories generate the most rework and which agents carry them, and a knowledge test tied to the queues an agent actually handles becomes an assessment somebody can be appraised against instead of a quiz nobody acts on.
A management layer cannot appraise anyone on rows that return no information available. What this review recorded is not a people problem awaiting a culture programme. It is a missing instrument, and instruments are something you build.
