Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesIT service management

Case study
IT service managementA large European public-sector organisation

Months of repeated testing on the organisation's own in-house applications, with the questions handed out beforehand, barely moved the score

A large European public-sector organisation ran repeated knowledge tests on the outsourced first-line desk that supported its own in-house applications. The questions were repeated and communicated to agents beforehand, and it still took close to six months for the average score to show a minor increase, with one month of the run skipped entirely. The review recorded that as a motivation problem and a supplier-management problem, and both readings hold. The same review also contains a second, far more expensive version of the same experiment: a dedicated second-level expert tier bought specifically for those in-house applications, which drew the largest single share of the desk's escalations and was judged not to have been effective. Two instruments aimed at holding bespoke system knowledge inside people, one cheap and one contracted, both returning close to nothing. The engagement produced a diagnosis and design principles for a future desk, not a rebuilt knowledge estate.

~6 monthsRepeated pre-announced testing before the average score showed even a minor increase
Client
A large European public-sector organisation
Duration
Current-state review of an outsourced service desk, findings and recommendations
AI · RIDGE E59.1 N62.5ρmax 1.00
2Agents holding the in-house application queue between them
1 monthSkipped in the run, with no test conducted at all

There is a version of a knowledge test that stops measuring knowledge. You write the questions, you tell the people sitting the test what the questions will be, and then you run the same test again the following month, and the month after that. Difficulty is no longer a variable. What is left is a reading on whether the material can be made to stick at all in the people you are loading it into.

A large European public-sector organisation ran that test on the outsourced desk taking first-line IT calls from its staff. The subject was the organisation's own in-house applications, the ones written for its work and running nowhere else. The stated purpose was both to verify and to improve agent knowledge. Questions were repeated across sittings and communicated to the agents ahead of time. The results came back poor, and it took close to six months of that treatment before the average score showed even a minor increase. One month in the run had no test at all.

The review recorded two causes: a lack of motivation among the agents to improve their knowledge of the subject, and a lack of management capability on the supplier's side to drive and manage performance. Both are defensible readings of the evidence, and the recommendations followed from them, which is why the report proposed motivation and satisfaction surveys as contract measures and performance-related bonuses, noting that a comparable bonus arrangement already existed on another desk inside the same organisation.

The challenge

Set the motivation question aside for a moment and look at what was being asked of the arrangement.

Knowledge of a bespoke application has no external supply. There is no vendor course for it, no certification, no prior employer an agent could have picked it up from, no public forum where the answer to last week's failure has already been written down by a stranger. Every unit of that knowledge has to be transferred internally, from the people who built or ran the thing to the people answering the phone about it. And it decays between uses, because an individual agent on a general first-line queue meets any one in-house application rarely. Add ordinary turnover in an outsourced first-line role, and the loss rate is high while the only replenishment channel is slow and manual.

Repeated testing with pre-announced questions is close to the best conditions anyone is going to arrange for that transfer. It removes difficulty, it removes surprise, it repeats the same material on a monthly cadence, and it comes with visible management attention attached. The score still barely moved. When an intervention that favourable returns almost nothing, the constraint usually sits in the mechanism rather than in the people sitting the test.

The review contains the corroboration, in the section that has nothing to do with training. When the escalations leaving first line in the month examined were cut by receiving group, the largest single destination was the second-level team dedicated to the organisation's in-house applications, a team the organisation was already paying for under an extended contract so that expert knowledge of those applications would sit somewhere reliable. The review's verdict on that arrangement was unambiguous: it had not been effective.

So the organisation had run the same experiment twice, at two price points. The cheap version put the knowledge into first-line heads by testing them repeatedly with the answers in hand. The expensive version bought a dedicated tier of specialist heads. Neither held.

There is a third data point, and it is the one that shows what the alternative looked like in practice. Where knowledge of the in-house applications had genuinely accumulated, it had accumulated in two people. Those two handled the queue between them, and when one of them took leave a backlog formed and kept growing until they came back. That is not a knowledge estate. It is a single point of failure with a holiday calendar attached.

The approach

It helps to move the noun. Knowledge of a bespoke system was being managed here as a property of staff. It behaves far better managed as an artefact, with a location, an owner and a retrieval path.

An artefact has a place it is written. The review found the desk working from the supplier's own store rather than the one the contract had named, with the consequence users described in their own words: the same incident escalated over and over to second and third line, because a resolution reached once was never written where the next agent would meet it. Every one of those repeat escalations is a piece of knowledge that existed, was used, and was then allowed to evaporate.

An artefact also has a retrieval path, and here the path stopped at the tooling. Stakeholders reported that the text of incident tickets was hard to search, and that departments ran their own tools with no way to bring what each of them held into a single view. Whatever the last identical failure on an in-house application had cost to work out, there was no dependable place to go and read it back. Memory was the only store left standing, and memory is exactly what the training programme had been trying to top up.

Once you see it that way, the test result reads differently. Six months of pre-announced quizzing is an attempt to raise the retention rate of a channel that has no business being the primary channel. Even if it had worked, it would have worked until the next roster change.

The outcome

What this engagement produced was a current-state diagnosis, a set of design principles for a future desk and recommendations feeding a re-procurement. No knowledge store was merged, no tooling was replaced, and the training and motivation recommendations are the ones that reached the report. The structural reading above is drawn from the same evidence rather than lifted from the conclusions.

If the same review ran now, the first change would be to stop testing the agents and start instrumenting the flow. Repeat-escalation rate per application, reassignment counts per ticket, the point at which an escalated ticket stops moving: all of that already sits in the incident records, and process mining over them produces it without anyone being able to prepare. None of that needs an analytics programme. It needs the records to be in a state where they can be counted, and in this estate they were not. Platform work of ours is mostly that repair: consolidating where tickets live, making a category mean the same thing twice, and keeping the trail from any published figure back to the rows behind it.

The second change is to make the resolution itself the deliverable. Write-back at closure, into one store, keyed to the application and the symptom, surfaced at the moment the next contact arrives rather than filed for someone to study later. That turns bespoke knowledge into something the organisation owns instead of something it rents, which matters most in exactly this setting: a first-line contract that assumes system knowledge will accumulate inside the supplier's people is buying an asset that walks out of the building when the term ends.

The same holds for the language-model pilots being tried on support content. Drafting an answer out of past resolutions works only where past resolutions exist as text; on a bespoke application nobody ever wrote up, a model has nothing to draw on and will say something anyway. That is why a Consult engagement with us opens on the corpus question, which is whether the resolutions were ever recorded at all, before anybody scopes an assistant to serve them back.

The limit this review exposes has nothing to do with whether a group of agents would learn. A system nobody wrote down cannot be taught fast enough to outrun the rate at which it is forgotten.

NEXT STEP

Ready to make AI real?