A test whose questions you have been given in advance is no longer a test of what you know. It is a test of whether you opened the questions. Run it once and the result is ambiguous. Run it four times in a row, announcing each time that the same questions will be used again, and the ambiguity goes away. Whatever the score does across those four sittings is a reading on preparation, not on knowledge.
At a large European public-sector organisation, the score did almost nothing. Technical knowledge tests had been introduced part way through the life of an outsourced service desk contract, and the first results came back far below expectations on questions the organisation itself described as very basic. The test was then repeated four times in a row, each time with the announcement that the same questions would be used again. The number of correct answers hardly increased.
That sentence sat in an appendix, one bullet under a heading about people. It is the most informative measurement in the review.
The challenge
The organisation was reviewing an outsourced desk that took internal IT support calls from staff across several internal departments. The review gathered stakeholder views from application management teams and department heads, scored the current state, and asked what a future desk should be designed to do.
The survey results are what you would expect of a service under review and nothing more. Perceived knowledge of the desk clustered in the middle, with the largest single share, 43 percent, at moderately knowledgeable, and the two ends of the scale carrying 14 percent each. First-level resolution split into three exactly equal thirds: some incidents, about half of them, most of them, depending on who you asked. Overall experience had two slices and only two, about the same as expected against slightly worse, split 80 to 20, with no better-than-expected slice at all. Shares reported this way come off a base you can count on two hands. They are a mood, not a measurement.
The four sittings are a different kind of evidence. They are an experiment, whether or not anyone intended one, and it has a control built into it. If knowledge were the binding constraint, pre-announcing the questions removes it, and the score climbs. The score did not climb. So the binding constraint was somewhere else, and the rest of the review says where.
Regular training was not in the scope of the arrangement and had to be explicitly requested. Customer handling seminars began only after the organisation asked for them, with repeated reminders and formal escalations before anything happened. Sending agents on courses started the same way, by request rather than by design. Knowledge sat in silos: one class of application incidents was handled by two agents, and when one of them took leave a backlog built and kept building until that person came back. There was no knowledge base worth the name, so the same incident was escalated to second and third line over and over rather than being answered once and then answered thereafter. No quality checks or audits were carried out on incidents, and no one was accountable for an individual incident once it was closed.
Then the closure behaviour, which is the part that makes the test result inevitable. Tickets were shut on the assumption that the problem had gone away rather than on confirmation from the person who raised it. Nobody called back once a ticket closed, and no spot checks were run on user feedback. Interviewees described agents who seemed more eager to close a problem than to resolve it, which is exactly the behaviour you would design for if you built the incentives on purpose.
Put those together and the four sittings stop being surprising. An agent facing a pre-announced test has to believe two things before preparing is rational: that the answers can be found somewhere, and that being right will be noticed. Neither held.
The approach
The useful move in the review was to stop treating the test as a verdict on agents and start treating it as an instrument reading on the arrangement. That inverts the question: rather than asking why the agents did not learn, it asks what would have had to be true for the score to move.
Retrievability failed first. Text search across incident tickets was difficult. Priority and severity could not be set in the tool. Different departments ran different tools, so there was no way to bring the information together into a consolidated view, and the incumbent tool was hard enough to integrate that it discouraged anyone from adopting it as a shared one. Reporting out of it was poor. An agent who wanted to look up how the last identical incident was resolved had no reliable place to look.
Ownership failed second. The governance findings describe a service with a single owner but without clearly documented and articulated roles and responsibilities. Improvement activity was being driven from both sides at once rather than sitting with the party contracted to run the service. Data was mostly captured, but comparison for analysis was difficult, so even where numbers existed nobody could put two periods side by side and argue about the difference.
Feedback failed third, and it failed at the point closest to the agent. Escalated incidents reached other groups with inadequate and inconsistent information for those groups to work from. Tickets were misrouted and passed between teams twice or three times a month per team. None of that came back to the person who had answered the call, because no audit existed to carry it back.
The outcome
What this engagement produced was a diagnosis and a set of design principles for a future desk, developed with the people who use it. Nothing was implemented in this phase, and none of the material below is a result.
The principles the stakeholders asked for map one to one onto the failures above. Easy access to information and a sound knowledge database. Higher first-call resolution. Accountability installed for the service provided. Agents knowing from the onset what is expected of them. And a service management tool that supports all of it rather than obstructing it. On channels the views were more mixed than the sponsors expected: self-help was thought workable for general issues and not much else, email was described as a source of delay and repeated contact, and satisfaction surveys after closure drew a response rate of about 10 percent in one area, with a strong preference for one question and one click over anything longer.
If the same review ran now, the test would not be repeated a fifth time. The evidence of what agents know is already in the ticket log. Process mining over the incident flow gives reassignment counts, reopen rates, the misrouting the teams described from memory, and the point at which an escalated ticket stops moving. Nobody prepares for that measurement, which is what makes it worth having. The condition is data readiness rather than analytics: one ticket store, consistent categorisation, and enough lineage to say which record produced which number. Our Platform work builds that lineage before anyone reads a number off it, because a figure nobody can trace back to a record is not evidence.
The current wave of cautious language-model pilots on support content sits in the same place. A model that drafts an answer from past tickets is a genuine improvement to retrievability, and it changes nothing about accountability. Where closure is assumed rather than confirmed and no audit follows the answer home, the effect is a faster wrong answer and a shorter queue. Our Consult work now starts by checking the closure loop before anyone scopes the assistant, because a support estate that cannot tell a resolved incident from an abandoned one cannot tell a good model from a bad one either.
A repeated measurement that does not move is not a stubborn workforce. It is a well-designed experiment nobody meant to run, reporting that the thing being measured was never the thing that needed to change.
