A contract can transfer an activity. It cannot transfer the consequence of that activity going badly. The gap between those two things is cheap to sign and expensive to live in, and it does not show up in the service level report. It shows up the day somebody asks why the service behaved as it did and nobody on the buying side can say.
A large European public-sector organisation had contracted out the internal IT service desk its own staff called when something stopped working. Ahead of a decision about how to buy that service next, it commissioned a current-state review: performance data, process, technology, the people running the desk, governance, interviews with the managers whose teams absorbed the escalations, and a study of how comparable desks are bought.
The finding I keep returning to is not in the performance data. It sits in the contract's own split of responsibilities, on the row covering how the desk was staffed. Against that row the supplier's desk manager was both accountable and responsible. The client's contract manager was informed. Not consulted. Informed.
The challenge
Follow that one row into the review and it accounts for most of what the people section was unable to say.
Every time the review reached something that happens inside the supplier's own management, the honest entry was that it sat outside the scope of the contract. Whether desk staff carried service management objectives linked to their performance was outside scope. Whether they had a career path was outside scope. What could be recorded instead came in two kinds: what the client observed from outside, and what the supplier asserted about itself.
The supplier's assertions were reasonable and unverifiable. It stated that it had specific, measurable objectives at every level, and that career paths were in place and managed locally. The reviewers' own read was that a career path probably existed for some senior staff. Nobody on the client side could settle it, because settling it was not a right the client had bought.
The observable side was thinner. Performance was not being managed in any form the client could see. Training and development had no particular focus and were addressed as needs came up rather than to a plan. There were no formal recruitment or progression plans. Roles and responsibilities had been documented, but nothing kept the documents current, and people were not clear what they meant in practice. Some agents held the standard service management foundation certification, and the reason recorded was that they were key people or personally keen on it, which describes individual initiative rather than a training regime.
Then there is the one occasion on which the client tried to look inside directly. It ran its own knowledge tests on the agents, covering the applications written for its own work, with the questions repeated across sittings and communicated beforehand. It still took close to six months before the average score showed even a minor increase. The review drew two conclusions, both in the deliverable: the agents were not motivated to improve their knowledge of the subject, and local supplier management was not driving and managing performance.
Read those two conclusions next to the responsibility row. Both point at management the client had contracted away. It could measure the symptom, monthly, at its own cost. It could not touch the cause. Outcomes visible, mechanism invisible, accountability sitting with the party that owns neither.
The governance section shows the same asymmetry from the other side. Most data was captured, but reporting off it was difficult and comparison across it harder, so even the visible outcomes arrived in a form that resisted analysis. There was a single named owner of the desk, and a widely held sense that the roles behind that ownership were not documented clearly enough to act on.
The approach
The recommendations that came out of this were governance ones, and deliberately dull.
Write a detailed responsibility split into the next contract, drawn on the processes as they should work rather than on how they had grown, so that clarity of roles is a contractual artefact rather than a shared assumption. Add named roles on the client side to face the supplier's desk manager day to day: a dedicated incident manager and a problem manager, so that the contract manager's attention goes to improvement rather than to individual failures. Add a knowledge owner. Add an audit function that checks process and service level adherence on a cycle, rather than discovering the state of things at renewal. And put agent motivation and satisfaction into the measured terms, which was by then becoming ordinary practice in service desk contracts.
None of that is clever. What it does is convert an assertion into an obligation. "Career paths are in place and managed locally" is an assertion, and a sincere one. "You will report this, we will audit it on this cycle, and this is what happens if the audit fails" is a term. The distance between them is the distance between a client that can diagnose its own service and a client that can only rate it.
The outcome
What this engagement produced was a current-state picture, design principles, a market view of how such desks are bought, and recommendations for the next contract. No desk was restaffed and no responsibility chart was signed. The recommendations were still recommendations when it closed.
The reason I am writing it up now is that the same trade is being signed again, quickly, in a different aisle.
Buy a copilot as a managed service and you own the outcome: an answer reaches a member of your staff or of the public with your name on it. The vendor owns the model, the evaluation set and the prompts. Retrieval-augmented generation sharpens the asymmetry rather than softening it, because the documents feel like yours. The vector store may sit in your own tenancy while the chunking, the retrieval and the ranking that decide what the model sees sit in somebody else's code. When answers get worse, you have the outcome and not the mechanism. You are on the informed row again, running your own quiz.
The clock is faster than it was on the service desk. A managed model can change underneath a service with no ticket in your queue and no line in your report. So the questions worth asking before signing are the ones this review had to ask afterwards: whose evaluation set is it, can you add cases to it, is the retrieval trace kept, can your own auditor read the lineage from a bad answer back to the passage and the version that produced it, and on what cycle does somebody check. Those are contract questions rather than model questions, and they are cheap now and unbuyable later.
The regulatory direction of travel points the same way. The European rules on AI have reached political agreement and will place documentation and oversight duties on organisations that deploy these systems, whether or not they built them. Nothing is in force yet. The practical version arrives through procurement first, because the party answering for the outcome is the party that needs the trace.
The same discipline holds for the early, carefully scoped autonomous experiments people are running at the edges of this. The narrower the scope, the more of the mechanism you should be able to see, not less. Our Consult engagements open on that ground, and the Platform work that follows is mostly about moving the evaluation set, the retrieval trace and the lineage onto the client's side of the line, where the accountability already sits. Data readiness and the plumbing habits of MLOps matter for an unglamorous reason: they are what makes an audit trail exist.
Outsourcing an activity you cannot see is a decision to be surprised on somebody else's schedule. The service desk version took years to become visible. The model version will take months.
