Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsData Strategy

Logged Perfectly, Resolved Badly

RealAIApr 21, 20238 min read
Data StrategyIT Service ManagementPublic SectorData QualityMLOpsProcess Mining

There are two disciplines running inside every operations function, and because they are performed by the same people using the same tool at the same moment, almost nobody separates them.

The first is recording discipline: making sure that a piece of work which happened leaves a row behind it. Somebody rang, a ticket exists. The call arrived by email, the row says email. It came in at 09:12, the row says 09:12. That discipline is largely mechanical, easy to contract for and audit, and enterprises are genuinely good at it.

The second has no settled name, so call it semantic discipline: making sure the row says what the work actually was. What kind of fault this is. How urgent it is. Which group can fix it. What eventually fixed it, described so that you would recognise the same fault arriving next quarter under a different set of words.

Recording discipline gets you a corpus. Semantic discipline gets you a corpus you can compute on. They are different capabilities, they fail independently, and I have not yet seen an organisation buy the second on purpose.

The desk wrote everything down

Read the operating description and the recording obligation is unambiguous. The desk receives, monitors and manages incoming contacts, and its stated first responsibility is that the call gets logged promptly and with the relevant details captured. Priority is assigned at logging. The reference number goes back to the user. None of that is aspiration; it is the job as bought.

And it worked. A month of contacts resolves into per-route counts, the routes carrying a handful showing up as cleanly as the one carrying thousands. Escalations out of the desk are counted for the same month by receiving group and by category. Nobody questioned any of it: when the review needed to explain a slow month it reached for the logged call count and the argument was accepted.

That is a functioning recording layer. If the question were only whether the organisation had data, the answer was yes, continuously, in one place.

What the record did not say

Now look at the same tickets as a set of statements about the world.

Priority first, because it is the field every operations model reaches for. The agreement defined five levels and hung resolution targets on each, from one hour at the sharp end out to two weeks at the bottom. The tool the desk actually used supported three, set by impact. The contractual vocabulary and the recording vocabulary were never reconciled, and the field then collapsed in use: close to 95 percent of incidents were recorded at the lowest priority. It was populated every time and told you nothing, because a value appearing on nineteen tickets in twenty is not a description of a ticket. It is a description of the form.

Then classification. Incidents were categorised on the agent's own understanding of the fault, assisted by a suggestion the tool offered. Standard resolution procedures existed for some incident types and not others, so for much of the intake there was no written answer to what this kind of fault is called. The departments on the receiving end reported the consequence plainly: incidents not categorised properly, and a categorisation process too complex to follow. Two people meeting the same fault produced two different rows, and neither row was wrong in any way the system could detect.

Then the handover. When the desk could not resolve something it escalated to a second-line queue or to a third party, and the groups receiving that work reported the information arriving with it as inadequate and inconsistent for them to carry the analysis further. The ticket was complete as a record and empty as a brief. Misrouted tickets were passed around teams twice or three times a month per team, which is the signature of a routing decision made on a field that cannot carry one.

And then retrieval, which is the one that should worry anybody planning to build on this. The text inside incidents and tickets was not searchable, so even where an agent had written a careful and accurate description of a fault, the estate could not find it again.

The contract asked for the second discipline and paid for the first

The most instructive thing in the whole review is not a gap the reviewers found. It is a clause the agreement already contained.

Root cause documentation was owed in three circumstances: for every severity one incident, for severity two incidents involving an application outage, and for any incident that had recurred three or more times in production. The first two are events: somebody notices an outage and the obligation fires. The third is not an event. It is a query, answerable only if two tickets raised months apart, by different agents, about the same underlying fault, can be recognised as the same thing. The clause was unenforceable not because anyone avoided it but because the data could not answer the question.

The closure step repeats the structure. The technician assigns a closure category on resolution, and the agreement says what that category is for: it should reflect the immediate cause, and it is the basis for root cause analysis and trend reporting. The whole analytical apparatus was specified to sit on one free-judgment field, filled at the end of a call by whoever wanted to finish the call.

Knowledge shows the same split. The agreement put the shared knowledge database in scope for the desk. In practice the supplier kept its own, which had gone out of date, while the organisation kept another, maintained by a different support team on a document site not connected to the service management tool. Two knowledge stores, neither of them where the tickets were, and the receiving groups recorded the result: the same incident escalated over and over, because nothing captured at resolution came back within reach of the next agent.

Roles carry it too. Responsibilities were documented for incident handling and not for problem, change or configuration work, and until shortly before the review the desk did not own a ticket through to closure at all.

Every one of those is a failure of meaning, not of recording. Not one of them would be repaired by logging more.

3,662
Telephone contacts logged in one month, alongside 857 by email
5 vs 3
Contracted priority levels against the levels the tool could hold
~95%
Of incidents recorded at the lowest priority
3+
Recurrences triggering a root cause obligation nothing could detect

Why this decides what you can build

Every technique currently worth applying to an operations estate consumes the semantic layer, not the recording layer.

Process mining needs three things from a log: a case identifier, a timestamp, and the name of the activity that happened. Here the first two were solid and the third was a folk vocabulary. Run process mining on that estate and you get a rendered map of disagreement, drawn precisely enough that nobody questions it.

Straight-through processing needs a class you can act on without a human reading the text. The agreement lists the candidates: requests for information or advice, standard changes, application access, password resets. Those classes exist in the contract and not in the data, so the automation has nothing to attach to.

A feature store is a promise about meaning, and lineage is the record of who made it. A priority feature drawn from this estate would be a well-governed, version-controlled, fully lineage-tracked encoding of a habit.

And the language-model pilots now arriving on service desks are, underneath the demo, retrieval problems. Retrieval over which knowledge base, the stale one or the disconnected one, and finding the prior case how, in text nobody could search. The pilot will still produce fluent answers, which is what makes it dangerous here rather than merely disappointing.

Every ticket had a number, a channel, a timestamp and a priority. The number was reliable, the channel was reliable, the timestamp was reliable, and the priority was a habit. Three of those four are things a machine writes down. The fourth needed somebody to have decided what it meant.

What we would fix, and in what order

None of this needs a data scientist, and all of it has to land before one is worth hiring.

Publish a closed list of categories, short enough that two people handling the same fault land on the same value, and audit how it lands. Reconcile the priority vocabulary in the agreement with the vocabulary the tool can store, in whichever direction is cheaper, because a level that cannot be recorded cannot be owed. Write down what a resolution is, and what makes an incident the same incident as one from last quarter, since every recurrence clause and every deduplication step depends on that definition. Define the minimum a receiving group needs in an escalation and make the ticket refuse to move without it. Put the knowledge store where the tickets are, with one owner.

That sequence is the opening block of a RealAI Platform engagement, and it is why our Consult team asks which field you want to trust before asking which model you want to run. Answer those five and you have built nothing. You have a corpus that can carry a model, which is a different thing from a corpus that is large.

Details are as recorded in a current-state review of an outsourced IT service desk at a large European public-sector organisation: the contracted service description, the operating detail, the volumes and the interviews run alongside. That review produced findings and recommendations ahead of a re-sourcing decision, not delivered results. Reading them as the difference between recording discipline and semantic discipline is ours.

Every ticket had a number, a channel, a timestamp and a priority. The number was reliable, the channel was reliable, the timestamp was reliable, and the priority was a habit. Three of those four are things a machine writes down. The fourth needed somebody to have decided what it meant.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?