A public body cannot shortlist a supplier because the demo went well. Months later, in front of people who were not in the room and some of whom represent the bidders who lost, it has to show how the field was narrowed and why each supplier that fell out fell out. That obligation is not a step you add at the end of a market study. It decides what the study has to be from the first day.
A large European public-sector organisation was preparing to retender an outsourced IT service desk serving roughly 8,000 internal users across several sites. A study of the supplier market was one of the four contracted parts of the review: ten suppliers approached with one questionnaire written for the purpose, the six measures the contract already carried, and results distributed into quartiles.
The challenge
Vendor market studies fail in a quiet, familiar way. Each supplier is asked what it can do, each answers in its own vocabulary against its own definitions, and the comparison is a table of things that are not the same shape. Nothing in it can be defended later, because nothing in it was collected under the same conditions.
The incumbent arrangement supplied a second problem. First-call resolution was measured at 47 per cent, against a contracted target above 65 per cent and a cited industry average of 75 per cent. No historic series for that measure was available from the incumbent, so nobody could say whether the number was improving or degrading. A supplier that cannot produce a time series on the measure it is missing has made the miss difficult to challenge.
There is a subtlety inside those figures that a careless reading destroys. The review's own benchmark card put the market norm for first-call resolution at 65 per cent, inside a range from 40 to 95, so the contracted target sat exactly on the norm rather than beneath it. The separate finding, that on three of the six contracted measures every other supplier offered a higher level as part of a standard offer, must belong to the other measures. A cited industry average, an advisor's norm and what a supplier will put in writing are three different numbers answering three different questions, and only the third one is contractible.
The approach
Four design choices carried the defensibility, and none of them is expensive.
One questionnaire, written for this review rather than reused, issued to all ten suppliers. Same wording, same definitions. Comparability is a property of how answers are collected, not something analysis can add later.
A fixed set of six measures, and the fixed set was the one the contract already carried: average speed to answer, average talk time, first-call resolution, abandonment time, abandonment rate, and service level. The last was defined on the page as the share of calls answered within thirty seconds, which reads as pedantry until two suppliers answer it against different thresholds.
Distribution into quartiles per measure instead of an overall ranking. Offers varied across all six, so no supplier was uniformly best, and a composite score would have hidden that behind a single number.
And the limits printed on the face of the analysis rather than in a footnote nobody reads. These were supplier-stated offerings, not measured performance, and where a supplier gave a tiered answer the best value was taken, so the ranges flatter the market by construction. Both admissions make the study look weaker and make it far harder to attack.
The spread was the finding. Answering one identical question, the ten suppliers committed to first-call resolution anywhere from 40 to 85 per cent. Average speed to answer ran from 22 to 50 seconds, average talk time from 2.5 to 10 minutes, abandonment time from 9 to 37 seconds, abandonment rate from 2 to 5 per cent, and service level from 80 to 100 per cent. Anyone writing the requirement at the middle of the market would have quietly excluded the top quartile or accepted the bottom one.
The narrowing rule the study recommended names no supplier: rank the six measures by business importance first, then shortlist from the top two quartiles on the measures that rank highest. Six of the ten sat in the top two quartiles on at least three of the six. Three sat in the bottom quartile on at least three. That is a shortlist a procurement officer can defend, because the rule that produces it can be written down, published, and applied by someone who has never seen the responses.
Two further disciplines are worth naming. A detailed analysis of the tools market was declared out of scope on the page, so nobody could later claim the study had covered it. And the benchmark card carried eleven measures while the contract carried six, which is where two of the recommendations came from: define incident acknowledgement and resolution time explicitly and by priority level, which the contract left undefined for most sources, and carry agent utilisation, which the review called the best measure of productivity and which the contract did not carry at all. Benchmark wider than the contract, then let the difference set the agenda.
The outcome
What this engagement produced was a study, a quartile distribution and a recommended shortlisting band. No supplier was appointed, no service level was renegotiated, and every figure above is either an offer or a market benchmark. Reading the quartiles as performance would repeat the mistake the study printed a warning against.
Three things would be done differently now.
All six measures are things the organisation's own telephony and ticketing systems record, and the incumbent already reported five of them as three-month averages, the sixth only as a single month. A study run today would pull those distributions out of the source systems first and set the supplier-stated offers against them, so the comparison has a measured column as well as a stated one. The instrument stays; the evidence behind one side of it gets much better.
Second, suppliers no longer arrive selling only agents and shift patterns. They arrive with copilots for the desk and retrieval-augmented answering over the knowledge base, with deflection and containment percentages attached. Those percentages have precisely the property the review labelled on its own face: they are supplier-stated. The questionnaire discipline transfers without modification. What changes is the measure set. Add the ones the original six cannot carry, and require them measured rather than offered: answer accuracy against a held-out evaluation set built from the organisation's own resolved tickets, how often the assistant escalates what it should have answered and answers what it should have escalated, and the latency of the human review step. A supplier who will not answer against your evaluation set has told you something useful.
Third, data readiness belongs before the questionnaire rather than after it. A retrieval answer is only as good as the corpus behind it, and a knowledge estate split between an outsourced desk and a retained second line, with no shared store and no expiry date on articles, gives a vector store fragments to average. Test that before asking anyone what their containment rate is. The Platform work we do starts there, with what the ticket history and the knowledge base actually hold, because an accuracy figure quoted against someone else's corpus predicts nothing about yours.
One clause is worth adding while the contract is still open. Ask for lineage on every assisted answer: which article it drew on, which version, and who last changed it. This organisation could not obtain a historic series for first-call resolution from its incumbent. The same gap on a model-assisted desk is worse, because once the knowledge base has moved on there is no way to reconstruct why an answer was given. The AI Act is agreed rather than in force, and the record-keeping it points toward costs very little to design into a procurement and a great deal to retrofit into a running service.
The finding the study was built to survive is not that one supplier turned out better than another. It is that ten suppliers answering one question gave answers 45 percentage points apart, and the only reason anyone can act on that spread is that the rule for narrowing it can be stated without naming a single one of them.
