A review of an outsourced IT service desk, reported three months after its measurement window closed, set three numbers beside each other for every service level in the contract: industry average, contractual target, and what the desk actually did across January to March. On how fast a human picked up the phone, the desk was excellent. On whether the caller's problem went away, the same table reported 47% one-call resolution against a 75% industry average.
The challenge
The desk was the front door for staff across several sites and a network of affiliated partner organisations: intake, logging, phone and remote triage, desk-side visits, escalation, knowledge-database upkeep, access provisioning. It ran 07:30 to 18:30 Monday to Thursday and 07:30 to 17:00 on Friday, with no weekend coverage.
Demand arrived almost entirely by voice. In the final month of the window telephone carried 3,662 contacts from internal users and 4 from external ones, 79.4% of the month; email 857 and 5, or 18.7%; on-site visits 23, or 0.5%; self-service 15 contacts, 0.3%. Internal users were 98.9% of the total. The reason given for the self-service figure was usability: the portal "is not user friendly".
The pickup metrics were comfortable. Average speed to answer was 17.42 seconds against a target of under 25 and industry at 30. Service level, the share of calls answered within 30 seconds, was 94.9% against a target above 78% and industry at 85%. Abandonment rate was 2.06% against a target under 8% and industry at 3.5%. The last two carried targets written softer than the market, though on these three months the desk cleared the market number as well. Both targets should be reset, the review said, abandonment rate "to the industry average or maybe best in industry".
Three service levels missed. Average talk time was 5 minutes 45 seconds, over both the sub-5-minute target and the 5 minute 30 second industry average, and the longer calls bought nothing: talk time was not reflected in one-call or first-call resolution. Abandonment time was 1 minute 32 seconds against a target under one minute and industry at 25 seconds, so not every clock was green. And one-call resolution was 47% against a contractual floor above 65% and industry at 75%. That last figure is the one cell in the table carrying a date of its own, April, rather than the January to March average printed in its row label.
Then there were the service levels that did not exist. Acknowledgement carried industry reference times across five priority bands, 15 minutes at P1 up to one business day at P5, and no contractual target at any of them. Resolution targets existed and were loose: one hour at P1 against a 15-minute industry norm, two weeks at P5 against one business day. Every measured-performance cell in both tables read as unavailable. Email acknowledgement existed for two sources only: one hour for internal users, 30 minutes for partner organisations.
Those two tables and the six service levels above them describe one object: the life of a single call, measured at every point where somebody thought to put a clock.
The two widest hoops stand on bands that clear their targets. The hoop that decides whether the caller has to ring back narrows to 47%, its band hangs below the plane, and the stages at either end of the descent have no band at all.
The cause was mechanical. The service management tool could not represent the contract's five priority levels, so the resolution service level went unreported by any other means. Priority collapsed to low, medium and high on impact, and close to 95% of incidents were set to low.
Everything the telephony platform emitted carried a number, answer speed down to hundredths of a second. The two things a user actually feels, acknowledgement and resolution, had no clock at all.
Downstream, 844 incidents went to second and third level in that same final month, roughly 20% of those logged, the largest share landing on the team covering the organisation's own applications, the area a dedicated expert contract was meant to cover.
The approach
The current state was assessed through four lenses: performance, processes, governance and reporting, and maturity plus target operating model. One discipline in the performance work is worth copying. Every service level was read against the market, the contract, and measured reality, so the target itself could be audited, not only the achievement against it.
The process review walked the documented lifecycle from logging through workaround, escalation and close-out. Users were never told expected response or resolution times, and the provider worked from its own knowledge base rather than the client's, contrary to the contract. Quizzes on the organisation's own applications took agents almost six months to move the score, with the questions repeated in advance.
Cost was read together with productivity. Cost per incident had always been high and rose sharply from the month before the measurement window opened, the stated cause being a second supplier providing extended user support: floor-walkers who headed off incidents before they were logged. Volumes stabilised from the window's first month for the same reason. Prevention worked, and it made the unit-cost metric look worse, by taking tickets out of the denominator. Agent utilisation appeared in no service level and no KPI, which the review tied to inflexible staffing.
What came out was a design and a set of recommendations, not a delivered service: make agent utilisation a key KPI in the next contract; define acknowledgement and resolution explicitly, and set priority levels against what the tool can hold; reset the soft targets toward industry; and weigh a mixed on-premise and off-site model, since roll-outs are continuous and each is a demand shock.
The outcome
The appendix was an evidence base for a decision, not a report card: current-state analysis, stakeholder views, a market study of outsourcing trends, vendor profiles, a transition framework. A review ending in vendor profiles and a transition plan is feeding a re-procurement.
- 17.42 sec
- Average speed to answer
- 47%
- One-call resolution
- 844
- Escalations to L2 and L3 in one month
Read with a decade of hindsight, the useful part is not that an outsourced desk underperformed. It is that the failure was legible in the measurement regime before it showed up in the service, and every AI service-desk programme now being scoped inherits that regime, starting with the inversion in the title. An agentic tier makes answer speed close to free, and 17.42 seconds already beat target and market, so nothing is bought there. The only scarce quantity was the 47%. Point an autonomous tier at a desk tuned for pickup speed and you get an instant answer that resolves nothing, reported green. The review also recorded "some confusion around measuring one call resolution", and the number itself came stamped with a different month from every other cell in its table. An unstable definition makes both the human baseline and any machine improvement claim unfalsifiable, which is a problem before a single model is chosen. Behind it sit the empty cells: no clock on acknowledgement or resolution, and no model supplies one. The first honest deliverable of most AI programmes in service management is a timestamp, not an agent.
Triage is the highest-value target for classification here, and the 95% figure shows where the trap sits. Priority collapsed because the contract spoke in five levels, the tool held three, and people defaulted to the bottom one. An agent writing into that same field reproduces the collapse, faster. The schema has to hold the answer before the classifier is worth building.
That makes this a harness problem before it is a model problem. The harness is everything an agent reaches when it acts: the ticket schema it writes into, the knowledge base it reads, the provisioning system it changes, the escalation queues it hands work to, and the clocks that record what happened. On this desk every one of those was already bent. The tool could not represent the contract's five priority levels. The provider worked from its own knowledge base rather than the client's, contrary to the contract. Acknowledgement and resolution emitted nothing at all, so no measurement existed to tell a solved incident from a quiet one. Put the strongest available model inside that and it inherits each defect and runs it at speed. Rebuild the harness and the same model is working with a full set of instruments for the first time.
The loop is the other half, and it is the half this desk never had. A desk run as a loop does not stop at the answer. It reads the contact, proposes a priority against a schema wide enough to hold it, attempts the fix, checks whether the user's problem actually went away, and feeds that check back as the signal that grades the next attempt. One-call resolution is precisely the number a loop moves, because it measures the loop closing; average speed to answer measures the loop starting and nothing after it. Loop and harness engineering is what turns the finding in this review into a workstream rather than a re-procurement. Both are engineering with a known shape, done against definitions the client and the provider agree on first, and neither waits on a contract cycle to show whether it worked.
The channel data points at the prize. Self-service took 15 contacts in a month against 3,662 internal calls by telephone. Portals failed because they made the user do the triage: pick the category, pick the priority, describe the fault in someone else's taxonomy. A conversational agent inverts that, and the honest measure of success is the 79.4% phone share, not portal visits. It also takes in what a phone line cannot carry. The error dialog a caller would otherwise read down the phone, an access request that arrives as a scanned form: computer vision reads each of those into a filled field and a proposed priority without making the user translate anything into someone else's taxonomy first. Email was the second channel by volume on this desk, and pictures and documents travel by email. The same reading works at the other end of the loop, where an agent that can see the screen after a change can check the fault is actually gone rather than asking the caller to vouch for it, which is the check one-call resolution was always meant to be. Coverage is the other cheap win, given a desk that ran four long weekdays and nothing at weekends. Which contacts it may close unsupervised is policy, not model capability.
Two warnings come with it. The denominator trap is the first: this client did the right thing, putting people in front of users to prevent incidents, and its cost-per-incident number got worse. Deflection takes the cheap contacts first and leaves the expensive residue, so an AI programme judged on cost per ticket looks like it is failing while it works. The boundary is the second. This contract put first-level support for the client's own applications in scope while assigning nearly every duty specific to them elsewhere, and the largest escalation destination was that seam. A human analyst quietly negotiates an ambiguous carve-out. An agent cannot, so the boundary has to be written precisely enough for a machine to read: a contracting problem, not an engineering one.
Put those pieces together and the target is larger than a bot on the front door. It is a set of pipelines connected to each other: the intake that classifies a contact, the provisioning path it fires, the roll-out calendar that anticipates the next demand shock instead of absorbing it, the escalation route into second and third level, and a reporting layer that at last has timestamps to report. Each of those already exists on this desk as a manual hand-off, and each hand-off is a place the clock went missing. Wire them together and agents carry a real share of how the estate runs from one day to the next, while people keep the judgment calls, the ambiguous carve-outs and the incidents nobody has seen before. The evidence base this review assembled is the right document to start from, because it already names every seam and says plainly which numbers were never measured at all.
