Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsLeadership

The Vendor Benchmark That Measured the Wrong Thing

RealAIApr 16, 20248 min read
LeadershipAI StrategyData Strategy

Nearly ten years ago I spent a week inside a vendor benchmark commissioned by a European banking group's IT services subsidiary. The section I was handed covered five mid-tier offshore providers, one comparison pack each, a shortlist at the end of it. Every provider got a facts page and a page of recent news; some also got a strengths and weaknesses view or a table of recent deals. Nobody had cut corners.

It still could not answer the question the sourcing team actually had, which was whether any of these five could run a specific slice of application support for a regulated bank without the programme going sideways in year two. That gap between a well-made benchmark and a useful one is worth revisiting now, because the same instinct is being applied to a new category of supplier, and it is producing the same shape of document.

The sheet filled every cell it could reach

Read the pack cold and it looks thorough. One provider: revenue of GBP 189.2 million, 8,485 staff, 40 relationship offices across 30 countries, sector mix of 47 percent treasury and capital markets, 22 percent corporate banking, 16 percent insurance and other, 15 percent retail banking. Another: GBP 257.9 million, 11,341 staff, 31 global offices. A third: 3,200 staff and a service footprint described as reaching more than 140 countries. A fourth: USD 164.8 million, 3,352 employees of whom 2,535 sat offshore. A fifth: USD 379.6 million, 8,494 staff, over 230 clients across 18 countries, revenue split 44 percent Americas, 36 percent EMEA, 20 percent rest of world.

Precise numbers, every one of them, and every one of them public. Revenue, headcount, office counts, partner logos, award listings, board appointments, a chronology of announcements going back two years. The pack was thorough in exactly the dimension where thoroughness is cheap.

The columns that stayed empty

The instructive part of that benchmark was not the data. It was the pattern of holes.

The pack cited an independent European outsourcing survey as a comparison source. For two of the five providers the entry read that the survey did not cover them. For the other three it read as not available. Zero of five carried the one piece of evidence that came from European buyers describing what the supplier had actually been like to work with. The field for the named relationship contact was blank for four of the five.

Then the table of recent deals, which is where delivery evidence should live. For one provider it listed four engagements. Contract value was unavailable in all four. Contract start date was present for one. Contract length was present for two, at seven months and twelve months, both of them small implementation jobs in industries that had nothing to do with banking. Across those five exactly one deal row was complete: a government client in EMEA, contract starting in February, running 36 months, worth USD 13.8 million, covering application development, support and testing. One row, across five suppliers, described a relationship of the size and shape the bank was contemplating.

The benchmark did not fail because anyone was careless. It failed because every column in it could be filled from a press release, and none of them could be filled from our own work.

Two instruments, opposite verdicts

Once you notice the shape, the contradictions start showing up.

The trajectory field is a good example. One provider was recorded as growing, at 2.9 percent revenue growth year on year. In the same breath the pack noted that net profit had dropped 50.5 percent. Both facts sat in the same cell. A single word, growing, was carrying a company whose top line crept up while half its profit disappeared. Any sourcing manager scanning the summary row would have read the word and moved on.

Concentration was the second. One provider drew 55.3 percent of revenue from the United States, up from 46.6 percent the year before. Another drew 64.6 percent from the Americas, with only 7.1 percent from Asia Pacific and emerging markets. For a European bank that matters, because the accounts that get the strong delivery managers are the accounts in the growth region. The benchmark recorded the percentages faithfully and drew nothing from them.

The third contradiction was the sharpest. The provider with the highest growth in the set, 13.0 percent year on year, with 37.5 percent of its revenue from banking and financial services, was placed in the bottom band by three separate independent analyst grids, and ranked last within one of those bands. Meanwhile the provider that analysts placed highest for financial services specialism was declining at 21.9 percent year on year, and partway through the pack came the disclosure that a majority stake in it was being acquired for around USD 270 million, producing a combined entity of roughly 18,000 people and USD 826 million of revenue. On the pack's own disclosure, the company being scored was months away from being folded into a different one.

Two instruments, growth and analyst standing, pointing in opposite directions. Neither of them was about whether the supplier could run our application estate.

What the sheet should have asked

The useful benchmark starts from the scope, not the supplier. It states a defined slice of work with its volumes, regulatory perimeter, data residency constraint and integration points, then asks what evidence exists that this supplier has done that, and what happens when they do it for us.

Three columns would have been worth more than the whole pack. Named references at comparable regulatory scope, contacted directly, not logos. Attrition and tenure on the specific delivery centre that would staff us, because a provider with 2,535 offshore staff is irrelevant if the handful assigned to us turns over twice a year. And a paid trial: the same slice of real work given to the shortlisted suppliers, scored on output rather than on presentation.

That last one is the whole argument. Fitness is a property of the pairing between supplier and workload. It cannot be read off a company profile, however accurate the profile is.

The same sheet, now with model names in the first column

I am watching that pack get rebuilt. The first column now holds model names, and the fields have changed, but the logic has not. Parameter counts. Context window sizes. Public leaderboard positions. Funding rounds. A chronology of launch announcements. A demo recorded on the vendor's own data.

Public benchmark scores are this cycle's analyst grid. They measure something real, and that something is not your workload. A model that answers general knowledge questions well tells you little about whether retrieval over your own policy documents will surface the right clause, and less about what happens when it does not. Two suppliers can sit within a point of each other on a leaderboard and behave completely differently on a corpus of your credit memos, your incident tickets, the local-language product terms your legal team actually signs off.

The corrections are the same three, translated.

Build the evaluation set from your own work before you take the first demo. A few hundred real examples with agreed correct answers, including the hard and ambiguous ones, is worth more than every score published this quarter. It is also the only artefact that survives a change of supplier.

Write down the failures you cannot tolerate, and test for those specifically. A confidently wrong answer inside a customer-facing copilot is a different category of problem from a slow one. Hallucination control and guardrails belong in the acceptance criteria, not the marketing pack, and they need to be observable in production rather than demonstrated once.

Price the thing you will actually run. Cost per thousand tokens on a slide is not cost per resolved case at your volume, once retrieval, retries, longer prompts and human review are counted. Inference economics are becoming the constraint that decides whether a promising pilot ever reaches the whole department, and almost no vendor comparison sheet has a column for it.

What goes in the columns this time

The instrument matters more than the shortlist. A comparison built from public attributes is easier to produce, and it will always describe suppliers rather than predict programmes. One built from your own tasks, failure list and unit economics takes a few weeks, and is the only kind that has ever changed a decision I sat in.

That is why our own product surfaces are built the way they are. Hominis holds the evaluation sets, the review queues and the guardrails as first-class objects, so the criteria you selected a supplier against remain the criteria you keep measuring in production. Synapsa exists because the people writing those evaluation sets need to know what a good one looks like, and that is not a skill a procurement function has yet.

The five-provider pack still sits in my files, a reminder that a document can be accurate in every cell and still measure the wrong thing. The tell is always the same: if every column could have been filled without ever mentioning your work, it was never a benchmark of fitness.

5
Providers in the comparison pack
1
Deal rows with a contract value
0
With independent European buyer coverage
50.5%
Profit fall at a provider scored as growing

Fitness is a property of the pairing between supplier and workload. It cannot be read off a company profile, however accurate the profile is.

The benchmark did not fail because the analysts were careless. It failed because every column in it could be filled from a press release, and none of them could be filled from our own work.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?