Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsData Strategy

The Shadow Tooling That Split the Record

RealAIJan 29, 20248 min read
Data StrategyIT Service ManagementPublic SectorData QualityLineageMLOps

Retrieval work in IT operations always starts from the same asset, the most obvious one in the building. Years of tickets, faults described by the people who had them, a resolution written at the bottom of each. Point a retrieval layer at it, put a copilot in front of the desk, and the next agent to meet that fault gets the answer the last one found.

The first question worth asking is not about volume, or which vector store. It is whether the ticket history contains the resolution at all.

In a current-state review of an outsourced IT service desk at a large European public-sector organisation, the answer was no. Roughly one incident in five left the desk at the moment it became interesting, and several of the groups it went to recorded their incidents and problems in toolsets of their own. What made those incidents worth learning from was written down in other people's tools.

The record has two halves and one owner

On paper the flow is clean, and worth reading as a data design rather than a process document. The desk logs the contact, sets a priority from the fault as diagnosed and reads a reference number back. If the agent cannot resolve it, the incident goes to a queue set up by subject matter in the second line, or out to a third-party provider, and the desk keeps ownership across that hand-off and hand-back until closure. At the end the desk technician or the resolving group assigns a closure category, and the agreement says what that field is for: it is the basis for root cause analysis and trend reports.

Read as an engineer, that is an end-to-end trace with an owner for every leg and an analysis field at the terminus. In practice the ownership across hand-off was a recent arrival, and the trace is severed at the escalation anyway, because the record then continues in a system the ticket's owner cannot open.

Where the second half went

The organisation did not have a tooling vacuum. It had a shared service-management tool, described without qualification as a common tool used across internal and external service providers. The qualification follows immediately: second-level support groups also ran their own toolsets in parallel to record incidents and problems, one of them a general-purpose issue tracker.

The interviews put it more bluntly than the assessment pages do. Different departments use different tools, so there is no way to bring all the information together. The central tool was hard to integrate with anything, which discouraged those departments from adopting it as the common view in the first place. Text search inside tickets was awkward and its reporting was poor. The knowledge base the organisation maintained sat on a separate document site with no connection to it, so the artefact designed to carry learning forward lived outside the system where the learning happened.

None of that is a failure of record keeping. Every group wrote things down diligently, in the tool that suited its own work. What the estate lacked was a shared key, and the records stopped meeting where the difficult work started.

Why the resolution clock could not be read

The contract named priority levels with resolution targets running from one hour at the sharp end out to two weeks at the bottom. Against every one of them the measured column was empty, and acknowledgement was worse: response times were defined for two arrival routes and not at all for the rest.

Some of that is a tooling limit: the tool held three impact-driven priorities against a contract that named five. Fix that tomorrow and resolution time still could not be reported, for a structural reason. A duration needs two events. Almost everything else on that scoreboard needs one: speed to answer, talk time, abandonment, call volume, where a switch observes a single moment and the number falls out. Resolution time needs a start in one place and a stop in another, and for an escalated incident the stop was recorded in whatever tool the resolving group ran. Nothing joined the two, which is the same complaint the interviews made about bringing the information together at all. So the duration was not measured badly. It was not measured, and the review's own summary of the reporting position lands where you would expect: most of the data was captured, and comparing it for analysis was difficult.

The split showed up at the user's end before any report. Two-thirds of the stakeholders surveyed said users were updated on resolution status only sometimes. Asked what an online tracker should show, they were consistent: current status, the team investigating, and above all who is working on it and when it will be fixed. Every item on that list lives in the escalated half of the record.

844
Escalations to second and third level, month sampled, as the review counted them
~20%
Share of everything the desk logged that month
No information available
Measured resolution time, every priority level
2 to 3 a month
Misrouted tickets per team, per interviews

The same fault, escalated again

The review's process complaints read like a service quality list. Read them as lineage and they are one complaint repeated. Tickets passed back and forth between teams. Escalated incidents arriving at resolver groups with inadequate and inconsistent information for those groups to analyse further. Above all, stated without hedging: a lack of learning and knowledge base updates, so the same incident was escalated over and over again. Once an incident has been resolved, one stakeholder observed, the next similar one should be handled at the desk.

That is a learning loop with its return path missing. The fix exists. It was found by a competent person and typed carefully into a system the desk does not open, and the artefact meant to carry it back sat on a document site the tool could not reach. Two knowledge bases were in play and neither closed the gap.

The contract itself depended on the missing join, and nobody had noticed. Root cause documentation was owed for any incident that had recurred three or more times in production. That is a query across the whole record, both halves, with a stable notion of what counts as the same incident. Run it against the desk's tool alone and you count arrivals rather than recurrences, and the faults that recurred most are the ones whose lives were spent elsewhere.

A duration is the only metric that needs two events. Put the start in one system and the stop in another, with nothing joining them, and the duration is not measured. It is guessed, politely, by whoever gets asked.

What a retrieval project would have inherited

Now put a copilot on that estate, which is what everybody wants this year.

Index the desk's ticket history and the retrieval layer will be excellent at the first half: how faults get described, which category people reach for, what the opening exchange looks like. It will be close to silent on what fixed anything, because for the escalated fifth the fix text sits in the resolving group's own tool. What answers it does produce come from incidents the desk closed itself, which are the easy ones, so the assistant ends up most confident where it is least needed.

An evaluation set built from that same history will not catch it. Ask the corpus what a good answer looks like and it will tell you, using the same truncated records. The scores come out respectable and the desk still escalates. That is the failure I would put ahead of invention on any operations pilot: not a model that makes things up, but a corpus that quietly excludes the cases you meant to learn from, in a way no metric computed inside it can reveal.

Traceability is becoming a governance question too. With political agreement now reached on the European AI Act, the direction of travel for systems touching operational decisions is towards showing where an answer came from, and a record scattered across unjoined tools cannot.

What we would fix, and in what order

None of this needs a data scientist, and all of it lands before one is worth hiring.

One identifier for an incident that survives every hand-off, carried into whatever tool the receiving group prefers rather than replaced by that tool's own numbering. A recorded link between the desk's ticket and the record the resolver group opens. Resolution text and closure category written back to the ticket that started it, as an obligation rather than a courtesy. One knowledge artefact reachable from inside the tool where the work happens. A start event and a stop event agreed per priority level, with the query that produces the duration stored beside the definition.

That is the opening block of a RealAI Platform engagement on any operations estate, and it is unglamorous by design. Do it and you have shipped no model. What you have is one continuous record of how work gets completed here, which is the only thing worth pointing a retrieval layer at, and the artefact root cause analysis needed anyway.

The organisation in this review was not disorganised. It had a shared tool, a written process, ownership through hand-off written into the contract, and a field whose stated purpose was analysis. The record split anyway, because tool choice at the second level was left to each group and nobody costed the join. Fragmented tooling gets discussed as an architecture preference or a licensing tidy-up. It is neither. It decides whether an organisation has a memory.

Figures and observations are as recorded in a current-state review of an outsourced IT service desk at a large European public-sector organisation: its contract scoreboard, its escalation counts for one sampled month, its process documentation and the stakeholder interviews. That review produced findings and recommendations ahead of a re-sourcing decision, not delivered results. Reading them as a lineage problem is ours.

A duration is the only metric that needs two events. Put the start in one system and the stop in another, with nothing joining them, and the duration is not measured. It is guessed, politely, by whoever gets asked.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?