Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsData Strategy

Most Data Captured, Comparison Impossible

RealAIMar 5, 20237 min read
Data StrategyPublic SectorIT Service ManagementData QualityMLOps

Almost every data readiness conversation I have had this quarter opens with a shortage story. Not enough history, not enough labels, not enough volume to justify a pipeline. Then we go and look at the estate and it is full of data. Ticketing, telephony, monitoring, time recording, all writing rows for years without pause.

The shortage is almost never in the capture. It is in the ability to put two of those rows beside each other and mean something by the comparison.

The cleanest demonstration of that I have worked on sits in a current-state review of an outsourced IT service desk at a large European public-sector organisation. The review laid the contract scoreboard out plainly: each service level the agreement named, an industry benchmark, the contract target, the measured three-month average. Read as an operations document it is a performance discussion. Read by anyone who has to stand a machine learning pipeline up on that estate it is a diagnosis, because the scoreboard splits in two and the split has nothing to do with how much data the organisation held.

One scoreboard, two different kinds of number

The telephony metrics are all there. Speed to answer averaged 17.42 seconds over the window, talk time 5 minutes 45 seconds, and abandonment time, abandonment rate and service level all reported. Nobody ran a project to produce those. Timing the gap between a call arriving and a headset lifting is a by-product of routing the call at all: the number exists before anybody decides it is a metric, and it means the same thing every month because a machine settled the definition rather than a person.

The incident metrics are not there at all. Acknowledgement had no target written into the agreement for any priority level, and no measured average against any. Resolution had targets at every level, from one hour at the sharp end out to two weeks at the bottom, and not one figure to place against them. Both panels carry the same flat phrase: no information available.

It would be comfortable to read that as a systems gap, the ticket side being less instrumented than the phone side. It was not. Logging ran continuously and the counts were trusted enough to argue from. The review quotes 10,029 calls logged in a single month without hesitation, using it to explain a staffing shortfall. A month of contacts breaks down by arrival channel to a tenth of a percent. Escalations out of the desk are counted for the same month by receiving group and by incident category. What was missing was never the rows.

The metric that needed a person

The difference between the halves is definitional, and one metric sits on the seam, which makes it the most useful thing on the page.

One-call resolution has a number. That number is 47 percent. It also has two problems no instrumentation would have caught. It sits in a row headed as a three-month average for January to March and carries the parenthetical label April: somebody reached for the most recent figure they had, put it in a column defined as something else, and shipped the page. The review states the second problem out loud, that there was confusion around how one call resolution was being measured at all.

Both have the same root. Resolving a call in one contact is not an event a switch can observe. It needs somebody to have decided in advance what counts as resolved, what counts as one call, whether a user ringing back next day about the same fault breaks the chain. Those questions have no technically correct answer. They get settled by a person and written down, or they do not, and when they do not you get a number that looks like the ones beside it and means something different each time it is produced.

That is worse than a blank, and it is the failure mode I care most about. A blank stops a pipeline honestly. A defined-by-nobody metric goes straight through it.

Why this is a machine learning problem and not a reporting problem

Take the label. Automated triage is the first thing anyone wants to build on a service desk, and the label it learns from is the category on the ticket. Here that field was filled by whoever happened to take the call, working from their own reading of the fault with a suggestion from the tool, against no published list of values. Nobody checked it afterwards either: the review records that no quality checks or audits were being carried out on incidents at all. So two tickets describing the same failure carry different categories, and two agents using the same category mean different failures. The group-by still runs, the chart still renders, and the pipeline faithfully reproduces a judgment nobody standardised. Nothing catches that, because internally there is nothing inconsistent about it.

Take lineage. Reporting ran on two tracks. The supplier produced a standard set of reports, some from the telephony system and some prepared by hand; the organisation pulled its own automated reports out of the ticketing tool. Both circulated, neither was the declared source of record, and a management dashboard was being built alongside. When the source of truth for a metric is whichever system this month's pack came from, every feature derived from it carries a coin flip nothing downstream can see. A feature store is a promise about where a value came from and what it means, and here the honest answer is that it depends who produced it.

Take joinability. Departments ran different tools with nothing pulling them together, so there was no consolidated view to report from, and the specialist groups behind the desk kept records of their own alongside the central one. Every one of those systems was capturing. What the estate lacked was a shared key anyone trusted, and one surface where two counts could sit beside each other and be about the same thing.

None of that is a volume problem, and none of it is solved by a better model.

94.9%
Calls answered inside 30 seconds, contract floor 78
2.06%
Call abandonment rate, contract ceiling 8
1 min 32 sec
Abandonment time, contract limit one minute
No information available
Measured incident acknowledgement and resolution, every level

What happens to management when nothing compares

Agent utilisation was never written into a service level or an indicator at all, so the desk had no published measure of its own productivity and staffing could not be flexed against demand. It would be easy to read what follows from that as a motivation problem, and the review does record thin training and flat morale among the agents. I read it as a downstream effect. You cannot hold anyone to numbers that will not compare across teams or months, so in the end nobody tries, and the missing comparison surfaces later as an attitude.

You can see it in what the people running the service believed. Asked how many incidents were resolved at first contact, senior stakeholders split in equal thirds between most of them, about half of them, and some of them. A handful of informed people is not a measurement, but it is a good tell: when the people accountable for a service disagree evenly about its central operating ratio, that ratio is not published anywhere they can check. On overall experience the same group was settled, four in five saying the service was about what they expected and the rest slightly worse, nobody better.

The service was not in crisis. The instrumentation was.

The switch counted whether anyone asked it to. Everything else waited on somebody having written down what it meant, and nobody had. Capture was never the constraint. Definition was.

What we would fix, and in what order

None of these repairs needs a data scientist, which is the uncomfortable part, because all of them have to land before one is worth hiring.

One definition per counted thing, with the query that produces it stored beside the definition, so a figure and its meaning travel together. One system of record per metric, named, any second path declared as derived rather than parallel. Classification turned from free judgment into a short list of values that any two people handling the same ticket would apply the same way, with somebody auditing how it lands. One named owner per number, a person rather than a standing group, who can be asked why it moved. A mapping between the records the specialist groups keep and the ones the desk keeps, so the two can be counted together.

That is the opening block of a RealAI Platform engagement, and it is why we ask which metric you want to move before which use case you want to run. Answer those five and you have built nothing. You have a sorted estate: here, telephony and volume would have come out ready, category and first-contact resolution would not, and a forecast on anything in the second group would have been fiction rendered in a clean dashboard.

The modelling end has never been cheaper. Pipelines are assembly work, deployment is close to solved in most stacks, and the first language model pilots crossing my desk stand up in days. That is why the measurement layer deserves more scrutiny than it used to get, not less. When the model was the expensive part, the inputs got examined on the way in. The model got cheap and the examination went with it.

Figures are as recorded in a current-state review of an outsourced IT service desk at a large European public-sector organisation: its contract scoreboard, its operating detail and the stakeholder interviews run alongside. That review produced findings and recommendations ahead of a re-sourcing decision, not delivered results. Reading those figures as a data readiness problem is ours.

The switch counted whether anyone asked it to. Everything else waited on somebody having written down what it meant, and nobody had. Capture was never the constraint. Definition was.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?