Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsLeadership

The Integration Layer Nobody Budgeted For

RealAIAug 29, 202610 min read
LeadershipIT Service ManagementSourcingAgentic AI

A large European public-sector organisation commissioned an outside review of its IT service desk; the executive summary is dated June 2015. The future desk was to be the first point of contact for internal users across all of the organisation's sites. The brief was a sourcing decision, and the review answered it properly: a current-state assessment, a market study, design principles, four sourcing options, one recommendation.

Read it now and something else stands out. The options were costed and scored against the systems and contracts anyone could point at. Most of what the diagnosis recorded happened somewhere else, in the handovers between those systems. That layer had no owner, no service level and no line in anyone's budget.

The complaints were all about handovers

The diagnostic page is the sharpest thing in the document. It lists nine key issues in one column and fifteen root causes sorted into four lenses: people, process, technology, and governance and reporting. What it never draws is a line from any symptom to any cause. The mapping was left for the reader.

Three of the nine are outcome measures: the first call resolution number, an average or patchy user experience, incidents resolved with unsatisfactory resolutions. The remaining six cluster hard. Incidents not categorised properly. "Inadequate and inconsistent information in escalated incidents to other groups for them to perform further analysis." Misrouted tickets passed around and across teams. Critical incidents detected too slowly and escalated slower still. Ping-pong of tickets between the service desk and other support teams. Duplicate incidents with no means of checking them, to the point that a user ringing to ask about their own open incident had that call logged as a second incident rather than a note on the first.

Only one of those six is a fault inside a system. The other five are faults in a handover between two of them.

The causes agree. Under technology the review records that the ticketing tool was "difficult to integrate with the other tools used by other support groups", and that service level reports could not be generated. Under process, in six words carrying more weight than anything else on the page: operational level agreements with other support groups were not defined. The desk had a contract. The second and third line groups had jobs. Between them, nothing written down.

The headline number sits in the same place. First call resolution measured 47 percent against a contract target above 65 percent, with 75 percent cited as the industry average. Twenty-eight points of gap, in an organisation whose own tool could not produce the report that would have tracked it.

What the contract could see, and what held it up

The sourcing options were built on things a procurement document can name. Three sat inside the incumbent supplier scope: the service desk itself, desktop hardware provision, desk-side support. Six sat in the retained organisation: the ticketing tool, second line support, in-house application support, and three other help functions the review interviewed by name.

Grouping them that way is our reading, not a count the review publishes. Nine parties make 36 pairs, and the brief was a desk that could exchange work with all of them. Eighteen of those pairs cross the line between the supplier contract and the retained organisation. That is the arithmetic of a join surface: systems rise in a line, joins rise as a square, and a contract priced per system will always underbuy the second one.

Loading exhibit
Exhibit 1Nine systems above the line. Thirty-six joins below it.Hover a platform. Above the floor sit the nine parties the review names around the future desk, three inside the contracted supplier scope and six in the retained organisation, each drawn wide in proportion to the boundary crossings it carries. Below the floor hangs one strut per pair that would have to exchange work: 36 in all, the 18 amber ones crossing the contract line. The review records operational level agreements with other support groups as not defined. Grouping the nine and squaring them is our arithmetic; depth is drawn, not measured.

Above the floor, the two incumbent contracts are stated at 2.7 million euro, sourced from the organisation's own service desk manager as at May 2015, with no period given on the page. The insourcing alternative was estimated at 4.3 million, and the review is honest that the figure excludes travel, recruitment, technology, training, bonus, overtime, depreciation and amortisation, and applies only internal employee day rates. Insourcing loses on a deliberately flattering number, which is the right way to kill an option.

Below the floor there is no number at all. Not through carelessness, but because nobody prices a layer they have not named. Those struts are what the design principles were really asking for when they called for higher first call resolution and fewer misrouted tickets reaching second and third line.

The market had a name for it and almost nobody had bought it

The review ran a market scan across seven capability dimensions and sorted what it found into four adoption bands, from the least adopted through to the practices vendors had stopped offering. Twenty-eight cells.

In the band reserved for what almost nobody had bought yet, under governance, sits a single entry: the service integration layer, outsourced to the service desk provider. The operating model cell in that same band is marked not applicable, which is its own small verdict: in 2015 there was no cutting-edge way to arrange the boxes, only a cutting-edge way to govern what ran between them, and only a handful of buyers were trying it.

In the band of practices on their way out sit two entries that describe the situation the diagnosis had just finished describing: a service management tool unable to integrate with other systems, and support groups working in silos with no coordination between them. The scan told this buyer, in its own shorthand, that it was living in the obsolete row and that the fix sat in the row nobody had bought.

The rubric scored six things and the decision rested on three

The four options were scored against six of the design principles on a five-point scale, drawn as filled circles rather than digits. Reconstructing them means reading the fill geometry of each mark, so treat the cell values as recovered rather than published, with more fill meaning more attractive as the reading convention. On that reconstruction the totals out of thirty run 10 for staying with the incumbents, 13 for insourcing, 19 for the narrow retender and 26 for the wider one. The wider retender is the option recommended, and the matrix agrees with the recommendation: it takes the highest mark on five of the six rows and ties on the sixth.

So the instrument and the answer line up. Then read the grounds the recommendation actually gives, and the alignment thins out. Three are stated: higher service integration, greater attractiveness to the market, and potentially higher value for money. Only the last of the three is a row in the matrix.

Service integration is the reason given first. It is also the thing the entire diagnosis had been about, the thing the market scan filed in the least-adopted band, and the thing that never became a scored criterion in the instrument built to choose between the options. Attractiveness to the bidder field is in the same position: the narrow retender's own narrative warns that "Market response might be limited due to the small scope". That is an argument about who will bid, not about any of the six things scored.

A rubric that leaves out the axis you will actually decide on stops carrying weight at the moment it should be carrying most. Here the omission was harmless because the scores pointed the same way anyway. That is luck, not method. The same gap shows up in AI vendor scorecards in 2026: pages of feature coverage, and no row for whether the join between the vendor's agent and your estate has an owner.

What a loop and a harness do to this same work

Everything in that review was made by hand. Stakeholder interviews across several divisions. A market study assembled from a sample of client organisations, the supplier community and a research team. Nine symptoms and fifteen causes typed into a grid, with the mapping between them left undrawn because drawing it by hand for a point-in-time study would have cost more than it returned. One measurement of first call resolution, from a tool that could not report.

Put a loop on the same work and the shape changes. A loop is an agent that observes the estate, acts, checks its own result, and escalates when the check fails. A harness is what makes that trustworthy: the rubric with the deciding axis in it, the ledger of evidence behind every claim, the counted denominator, the audit trail per action. The loop does the work. The harness is why you can believe it.

Four things become concrete.

The evidence pass stops being a sampling exercise. The scoring that carried the sourcing decision was never written as text. It was drawn as filled circles, and recovering the four totals meant reading the fill geometry of forty-one shapes across twenty-four cells. That is a computer vision problem, and a machine now reads that page more consistently than a human working through it late in a review cycle. The same capability aimed at the desk changes the ticket: a screenshot or a fifteen-second screen capture read by a vision model arrives as structured state on the incident, which is the direct answer to "inadequate and inconsistent information in escalated incidents". The user stops having to describe their screen to someone who cannot see it.

Triage becomes an agent that owns the join. Misroute, ping-pong, duplicate, miscategorised: four of the nine symptoms are routing decisions made with partial information. An agent holding the whole join surface routes on the state of the incident rather than the wording of the caller, and recognises someone asking about their own open incident instead of opening a second one. That fix removes a symptom the review had no way to count, because the document carries no ticket volume anywhere.

The 2015 automation ambition was the right instinct on the wrong mechanism. The review asked for password resets automated in the phone menu, the top five request types automated, and robotic automation of common fixes. Those are scripts: they run, and when reality differs from the script they fail quietly. A loop observes the result of its own action and hands the case up with the state attached rather than dropping it.

The eighteen crossings become executable. An operational level agreement written as a document is a promise nobody can test. Written as a pipeline it is a contract that runs: an agent each side of the boundary, a defined handover payload, a check on arrival, a clock, a log entry per crossing. The layer with no owner in 2015 becomes the most observable layer in the estate, because it is the layer the agents run through.

Speed is the point. The symptom-to-cause mapping that was too expensive to draw once by hand becomes a weekly artefact when a harness holds the evidence and re-runs the derivation. First call resolution measured once by interview becomes measured continuously against one definition. A sourcing decision that took a study becomes a standing question with a live answer.

Cost the layer under the floor

The lesson generalises past service desks. When you buy a capability you price what you can point at, and the failures show up in what connects it to everything else. That layer is missing from the business case because it has no natural owner on either side of the contract.

The test is small and uncomfortable. Count the systems your new capability must exchange work with. Square that count, subtract the count, halve what is left. That is your join surface. Now say who owns each one, what it promises, and where it appears in the price.

9
Parties the review names around the future service desk
36
Pairs among those nine (our arithmetic, not the review's)
18
Of those pairs crossing the contract boundary
47%
First call resolution against a contract target above 65%

Nobody budgets for a layer they have not named. Name it, give it an owner, and it stops being the reason everything else underperforms.


Sourced from the executive summary of a 2015 IT service desk sourcing review at a large European public-sector organisation. Figures, symptoms and recommendations are as recorded there. That document is a pre-decision deliverable: a diagnosis and a recommendation, no delivered outcome. The option scores are reconstructed from the scoring marks on the page rather than published as numbers.

The contract could see nine systems. The complaints landed in the thirty-six joins between them, and not one of those joins had an owner, a service level or a price.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?