A performance scorecard measures what somebody wrote down. That is a dull sentence until money is attached to the writing down, and then it becomes the most important thing about the contract.
A large European public-sector organisation asked for a review of its outsourced IT service desk. The desk sat on a volume-based contract, and the review found what volume-based contracts produce. What made the work unusual is where the useful evidence came from: not the vendor market and not a peer organisation, but a second outsourced desk inside the same organisation, on a different commercial model. That desk was on points.
The challenge
The IT desk was priced on volume, with service levels driven by the number of incidents. The review recorded three consequences: no flexibility in pricing or resourcing, no motivation to exceed service levels, and no way to accommodate peak demand. Cost per incident ran between EUR 38 and EUR 49 across a five-month window, standing at EUR 47 in the most recent month, which the review described as high against the market.
The measurement around it was worse than the price. There were many exceptions to achieving the service levels, and volume thresholds in the contract made achievement easier rather than harder. Incidents logged by email or through the self-help portal were not covered by the service levels at all. There was no contractual provision for calling a user back once an incident was closed, so no measurement of customer satisfaction existed anywhere. And the direction of travel was away from the telephone: the review's own recommendation was to move more contacts onto the self-service portal and onto chat, to take volume off the phone. Every contact that moves off the telephone leaves the measured service and becomes invisible to it.
Then there was the culture. The review described the operating model as one where problems get fixed by whoever makes the most noise about them rather than by the process. It records an instance: a senior manager whose problem the desk could not solve had it solved instead by an internal application team. What the review captures is the workaround, not a ticket for it, and that is the shape of the exposure rather than a proven count. Work that reaches resolution outside the desk is work the desk's own data does not describe, and this route was open to exactly the people whose problems were least routine.
The approach
The comparator sat one department away. A facilities and switchboard service line, also outsourced, deliberately kept onsite because the service depends on knowing the building, and dual-sited so it survives losing one location. A contracted speed to answer of 94% of calls inside 20 seconds. Two named KPIs: accuracy of incident categorisation, and recording the incident in the system within 20 minutes or less.
The commercial regime had two mechanisms working on different clocks. The bonus was worth 5% of the contract value for the period under review, generally assessed at a half-yearly review, with an additional clause dividing that bonus equally between the supplier company and the agents themselves. The malus ran monthly: the supplier is given 50 credit points at the start of the month, a pre-defined number of points is deducted for each error in a band of 10 to 20, an extra bonus follows if the balance survives, and a penalty follows if it goes negative.
Run the arithmetic on those two numbers, which is ours rather than the review's, and the allowance is tight by design. Three errors at the top of the deduction band overrun the allowance. Five at the bottom exhaust it exactly, and the penalty is written against a negative balance. And the clause that most contracts never contain is the split: the upside reaches the person answering the call, not only the account manager who negotiated the terms.
On the strength of that comparator the review recommended a bonus and malus regime for the IT desk too, with a caution attached in the same breath. The value for money handed back to the supplier has to be measurable, and the reward has to be the right quantity. Rewarding a desk for answering more questions is not the same as rewarding it for asking more questions in order to resolve the problem properly. Both sentences appear in the review, and it does not settle the tension between them. Neither will we, because wording does not settle it.
Now look at what the points are deducted against. Categorisation accuracy, and elapsed time from contact to record. Both are properties of the incident record. The supplier writes the record. A malus levied on record quality is a malus levied on data authored by the party being penalised, assessed after the fact, by a client reading that same data.
Two things follow, and neither appears on the scorecard. Contacts can fail to become records at all. Records that exist can be shaped so they survive the scoring: categorised into the safest bucket, timestamped at the convenient moment, split or merged so a single failure counts once. The scorecard sits downstream of the record and therefore cannot see either. The IT desk was already exposed on both routes with no points regime on it at all. On the first, the bypass culture and two whole channels sitting outside the service levels. On the second, the review's own current-state findings: incidents not categorised properly, and duplicate incidents with no check on them, to the point that a user ringing back to ask about an open incident was recorded as a fresh incident rather than a note on the existing one. Under a contract paid by volume, that is the record bending toward the money. Put a tight credit-point regime on a desk with those holes still open and the likely result is a cleaner scorecard drawn from a smaller sample.
The outcome
What this engagement produced was a review. A current-state diagnosis, an internal comparator worth copying, and recommendations covering the commercial model, the service levels and the measurement gaps. The credit-point regime was already running on the facilities line before anyone arrived, and none of its performance is ours to claim. No contract was signed in this phase, no service level was renegotiated, and no cost per incident moved. The contribution was reading the two contracts against each other and pointing at the one already in force down the corridor.
If the same review ran now, three things would change.
The first is that the incentive question and the data question stop being separate conversations held by separate committees. Any straight-through processing of routine requests, any model that predicts which contacts are about to breach, any classifier that assigns a category on arrival, is trained on the incident record. The record is the artefact the regime pays for. That makes the commercial regime a data-generation regime, and almost nobody treats it as one. A feature store built on tickets from a penalised desk is a feature store built on a filtered sample, and the filter is strongest exactly where the work is hardest.
The second is that the bypass stops being an anecdote. Process mining across the ticketing tool, the telephony platform and the second-level queues shows demand that arrives and never becomes a ticket, and shows where resolution work actually lands, including in teams the service desk contract never names. Lineage on the same events shows which records were edited after creation and when, which is about as close to a direct measurement of scoring pressure as anyone gets. Both are joins over data the organisation already holds. This is where the Platform work starts for us, because a measurement perimeter is a data readiness question before it is a commercial one.
The third is the one everyone wants to talk about. Early language-model pilots on ticket summarisation and first-draft categorisation are worth running, carefully and small, and they change nothing above. A model that categorises well makes the categorisation KPI easier to hit. It has nothing to say about whether the contact was logged. Consult engagements now open on the perimeter question for that reason: the cheapest way to spoil a support analytics programme is to build it on the desk's own account of itself.
Fifty points a month with 10 to 20 off per error, and a bonus worth 5% of contract value that reaches the agent, is a better design than most support contracts contain. The lesson is not the mechanism. It is the order of operations. Decide what a record must contain and how you will know independently that it exists, and only then decide what you will deduct against it. Do it the other way round and you have bought a supplier a strong, rational, entirely deniable reason to be careful about what gets written down.
