A field that almost always holds the same value is not a field. It is a formality with a database column behind it. At a large European public-sector organisation running an outsourced IT service desk, close to 95 percent of incidents were set to low priority. Read quickly, that looks like a triage failure, the sort of thing you fix by retraining the people doing the logging. Read properly, it is not a behaviour problem at all. It is what happens when a contract measures something the tooling cannot represent.
The contract defined five priority levels and hung resolution targets off each one: one hour at the top, then four hours, twenty-four hours, one week, and two weeks at the bottom. The service management tool the desk actually worked in could express three values, low, medium and high, assigned on impact. Worse, that assignment was largely automatic. The tool carried pre-defined incident models that suggested a priority from the impact recorded at logging, and the suggestion was mostly accepted. Nobody chose to file 95 percent of the estate's incidents as unimportant. The path of least resistance chose it, several thousand times a month.
The challenge
Follow the consequence rather than the cause, because the consequence is where the money went.
If the tool cannot hold the five levels the contract is written against, then no measurement of the contract's resolution targets is possible. That is exactly what the review found. Across all five levels, the measured performance for the reporting window came back as no information available. Not missed. Not disputed. Absent. The acknowledgement side was worse: no contractual target had ever been defined for any of the five levels, and the only acknowledgement commitments in the agreement at all were response times for two categories of inbound email, one hour and thirty minutes. Every other route into the desk had no defined acknowledgement time.
Set that against what the same review could measure without effort. Average speed to answer, 17.42 seconds against a target under 25. Service level, 94.9 percent against a target above 78. Abandonment rate, 2.06 percent against a ceiling of 8. Three others ran the other way: average talk time at 5 minutes 45 seconds against a target under 5 minutes, abandonment time at 1 minute 32 seconds against a target under 1 minute, one-call resolution at 47 percent against a target above 65. Six telephony targets, three met and three missed.
The point is not that the phone half was healthy, because half of it was not. The point is that every one of those six had a number attached, some of them to two decimal places, and they had one because a telephone switch counts whether anyone asks it to or not. Three of those targets could be argued about, and the organisation was in a position to argue. The ticket half could not be argued about at all, because there was nothing to argue with. Both halves were in the same contract, signed at the same time, by the same people.
This is the shape of the problem, and it is worth naming plainly: an organisation does not measure what matters, it measures what its systems emit. Where the two disagree, the systems win, quietly, and the contract becomes a document about a service nobody is observing.
The process review run alongside the performance numbers reached the same place from a different direction. Asked how an incident actually got its classification, the answer on record was the agent's own reading of the problem plus whatever the tool proposed. No rule anyone could point to, and nothing a second agent would reliably reproduce on the same ticket. Stakeholders interviewed for the same review listed their complaints about the tool: priority and severity could not be set, incident text could not be searched, second-line groups kept their own parallel trackers, the tool was hard to integrate with anything, and reporting out of it was poor. Five complaints, one root. The system of record could not represent the work, so people kept records elsewhere, and the central record decayed further.
The approach
The recommendation the review put forward was deliberately unglamorous: define acknowledgement and resolution explicitly in the contract, and set the priority levels against what the service management tool can genuinely express. Two levers, and only two. Either the contract narrows to what the tool can hold, or the tool gains the expressiveness the contract assumed. Choosing neither is what had been happening, and choosing neither costs more than either option, because it buys an unmeasurable agreement at the price of a measured one.
That recommendation was where this engagement stopped. Nothing here was rebuilt in this phase, and the numbers above are a diagnosis feeding a re-sourcing decision, not an improvement anyone has yet booked.
What we would add today sits one layer under the recommendation. A priority scheme is a data model, and a data model that has collapsed to one value can usually be reconstructed from evidence the organisation is already keeping. Roughly a fifth of logged incidents were escalated to second and third line. The process required a closure category on every ticket and named it as the basis for root cause analysis and trend reporting. Time to close, reopen behaviour, channel of arrival and escalation path all exist in the workflow event log whether or not the priority field is honest. Process mining over that log recovers the real distribution of urgency directionally, and it does so from behaviour rather than from a dropdown. The review's own evidence supports the direction rather than a figure, and we would state it that way to a client: the true spread of urgency is knowable from the event trail, and the number it produces has to be measured, not assumed.
That reconstruction is also the precondition for anything further. Automated classification is the first place machine learning tends to earn its keep in service management, and the first place it quietly fails, because a supervised model trained on a label that is 95 percent one class learns to emit that class with high accuracy and no information. Model deployment on a corrupted label is worse than no model, because it launders the corruption into a prediction that carries an air of objectivity. The Platform work we do therefore starts with the label, its lineage, and whether the field it lives in has enough levels to hold a signal at all. Data readiness in service management is not a volume question. This organisation had years of tickets. It had almost no variance.
The outcome
What was delivered was a current-state picture and a set of recommendations. The scoreboard was documented rather than repaired: eleven contractual targets, of which only six had any measured actual, three met and three missed, and the five acknowledgement rows never defined in the first place. That is the honest account, and it is more useful than a claimed improvement, because it tells the next negotiator exactly which clauses will be enforceable and which will not.
The lesson generalises past this desk. Every contract carries an assumption about what the tooling can say, and nobody writes that assumption down. When the tool is replaced or upgraded and the contract is not reopened, or the contract is renegotiated without anyone opening the tool, the two drift apart and the reporting continues without interruption, which is the dangerous part. A green dashboard sitting on top of a field with no variance in it looks exactly like a green dashboard sitting on top of a well-run service.
The current wave of cautious language-model pilots in service desks does not change this arithmetic, and in one respect makes it sharper. A model that reads an incident description well and drafts a plausible resolution still has to write its output into the same record. If that record cannot express urgency, the pilot produces faster tickets at the same undifferentiated priority, and the queue behind it is unchanged. Our Consult engagements start by checking the destination fields before anyone scores the model, because the cheapest way to tell a demonstration from a programme is to ask where the answer lands.
Ninety-five percent at one value is not a triage problem. It is a design problem with a triage symptom, and no amount of retraining the people doing the logging will move it while the field itself has nothing to say.
