Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesBanking IT sourcing

Case study
Banking IT sourcingA European banking group's IT services subsidiary

Five criteria, three points of spread, and a shared core worth 9 of 15

A European banking group's IT services subsidiary benchmarked candidate offshore providers for application development and support. The benchmarking pack rated five vendors on five criteria on a three-point scale and named the leaders in order in its headline. It never printed a total. Add each row and the totals are 14, 13, 12, 11 and 6 out of 15, and they reproduce that headline order exactly. Every rating sits next to the sentence that justifies it, which is the pack's real strength and also what makes its rubric reverse-engineerable: all five axes turn out to be functions of the cell printed beside them, and the deck contradicts itself on one of the cells that put the bottom vendor last. What was delivered was a rating table with its working shown, not an award.

9 of 15Score every leading vendor already shared
Client
A European banking group's IT services subsidiary
Duration
Vendor snapshot and benchmarking, pre-award
AI · RIDGE E68.2 N18.8ρmax 1.00
3 pointsSpread across the top four, out of fifteen
73 to 79%Profile overlap, subject with each leader
5 of 5Rows where the capability score was just an item count

A benchmarking pack ran to 41 slides and ended in an order. Five offshore providers, five criteria, a rating of 1, 2 or 3 on each, and a headline sentence naming the leaders in sequence. The pack never printed a total. Add each row and you get 14, 13, 12, 11 and 6 out of a possible 15, and those totals reproduce the headline sequence exactly, which is how you know a sum was the operative rule even though nobody wrote one down. The provider the group had brought the exercise to test came third at 12. The order looked decisive. It was three points wide across the four vendors anyone was choosing between, and every one of those points came from a judgement made by reading a website.

The challenge

The buyer was an IT services subsidiary inside a European banking group, deciding who would build and run its applications. The method decomposed into six modules: a vendor snapshot, a peer comparison, product and service capability scoring, salary benchmarks, a financial services deal history, and a country and location analysis. Each module was assembled by hand from a different body of evidence: vendor websites for capability, a compensation database for pay, a contracts database for deals, licensed analyst panels for client satisfaction.

The vendor under evaluation was the smallest firm in the set whose headcount the pack disclosed: 2,500 people, 16 offices across 10 countries, a claimed catalogue of twelve solution lines across twelve industries. Breadth was not the anomaly. Of the four rivals rated beside it, the pack carries a service catalogue for three, and those run to eleven lines, ten and nine; for the fourth it carries no profile page at all. The staffing behind the claim was the anomaly.

One row in the comparison was binary rather than comparative. Three of the five vendors offered tiered L1, L2 and L3 support; two offered basic support only. A bank buying run-and-maintain cannot take the second group at any price, so that line worked as a filter before any score mattered. It was scored anyway, as a 3 or a 2, on an axis labelled customer experience that carried no customer evidence at all.

The deeper problem is that five ratings were never a number. They were a five-element vector, and a vector is a shape.

Loading exhibit
Exhibit 1Five ratings are a region, not a sum.Each vendor is one solid. Spoke length is that vendor's 1-to-3 rating on one criterion, zero-based from a shared origin, in the deck's own column order: product and service capability, customer experience, offering strategy, verticals served, geographic presence. The solid is the convex hull of the five spoke tips, three axes on the equator and two at the poles, so a whole capability profile reads as a body rather than a total. Hull colour and edge brightness both carry the row total, dimmest at 6 and brightest at 14, which paints a confident ordering onto shapes that interpenetrate. The amber solid at the centre is the intersection of the four leading vendors, the per-criterion floor of 1, 2, 2, 2, 2 that every one of them contains, worth 9 of the 15 available points. All five share one origin, so they overlap in place instead of sitting side by side. Hover or tap a legend row to isolate one vendor and read its total and its overlap with the subject of the pack. The 1-to-3 ratings are the deck's; the totals, the shared core and the overlaps are arithmetic on them, and overlap is measured criterion by criterion as shared points over combined points rather than as enclosed volume.

Read as volumes, the four leading profiles enclose a common core worth 9 of the 15 points available. Sixty percent of the scale, and nearly two thirds of the leader's own 14, is ground every serious candidate already stood on. The subject shares 79 percent of its combined envelope with the second-placed vendor, 77 percent with the fourth and 73 percent with the first. The vendor the sum wrote off at 6 sits entirely inside all four of the others: a smaller copy, not a different kind of firm. Those overlaps are arithmetic on the deck's own ratings, not figures it printed, and there was no way to see them in a table of numbers.

The approach

The pack's real strength is that every score sits beside the sentence that produced it. Sixteen offices in ten countries. L1, L2, L3 support. Basic application development services. A challenged number could be argued on its evidence instead of on the number.

That same adjacency makes the rubric reverse-engineerable, and once reverse-engineered, the ratings stop looking like judgements. Product and service capability is a count: the score equals the number of capability items listed in the paired cell, three items scoring 3, two scoring 2, one scoring 1, five rows out of five with no exceptions. Customer experience takes exactly two values across the whole table, tiered support scoring 3 and basic support scoring 2. Offering strategy is a relabel, the words High, Medium and Low mapped straight onto 3, 2 and 1. Verticals and geography are bucketed counts: fifteen and twelve verticals both score 3, six scores 2, four scores 1, and the office counts fall into the same shape.

So every one of the five axes is consistent with a rule applied to the cell of evidence printed next to it, and not one of those rules appears anywhere in the pack. With five rows the bucket boundaries cannot be pinned down, but the direction can, and it means a vendor that had listed one more phrase on its website would have moved a point.

That matters most on geography, where the rule and the evidence pull apart. The axis does not track the country count sitting beside it: a vendor with 27 offices in 15 countries scored 2, one with 31 offices in 10 countries scored 3. Rank the rows by office count instead and the scores fall into order. The axis reads as office-weighted, and is never stated to be, so a reader checking it against the number their eye lands on will conclude the table is wrong.

The pack also disagrees with itself on one of the cells that put the bottom vendor last. The scorecard records that vendor as having offices in four countries and rates the row 1, one of the four 1s that make up its composite of 6. The deck's own profile page for the same vendor names offices in seven countries. A second vendor is recorded at six verticals on the scorecard and eight or nine on its profile. Nobody ran the internal consistency check, because a slide deck has no mechanism for one.

Meanwhile the facts that actually separated the candidates lived in the modules the scorecard did not touch. On pay, the subject's median developer salary was INR 713,000 against a peer range of INR 407,896 to 450,767, between 58 and 75 percent above every comparator. The smallest firm in the set was also the most expensive per head, so whatever the case for it was, it was not labour arbitrage. On track record, it held 1 financial services contract against 6, 4 and 4 for three peers. Neither fact moved the five-axis rating, because neither fact was one of the five axes.

The outcome

What was handed over was a rating table with its working shown, pointing at the spreadsheet detail behind capability, salary and deals. Not an award. The pack records no decision, and none should be read into it.

The evidence underneath it stretched from contract starts several years old, through satisfaction panels from the two most recent years, to a vendor news item dated ten days before the cover. An artifact whose oldest fact is nearly four years older than its newest is stale the day it prints, and this one could not be otherwise: six modules and 41 slides, each module assembled by hand from a different body of evidence, take as long as they take, and a vendor universe does not hold still meanwhile.

A ranking is only as fresh as the slowest source underneath it.

The version of this engagement you would run now is not a faster deck. It is two pieces. A harness is the rubric written down as code: five axes with their counting rules printed next to them, so that "three capability items scores 3" and "offices, not countries, drive geography" are executable lines rather than habits a reader has to infer; the support-tier gate running as a hard filter before any rating; no axis permitted to emit a number without a retrievable citation; and assertions that fail the run when the same vendor is described two ways in one deliverable. Four countries on the scorecard and seven on the profile page is an equality check. Six verticals against the eight or nine its profile lists is another. Every finding recovered by hand above becomes a test that runs in seconds. A loop is an agent driving that harness against a live vendor universe, on a schedule and on events: a demerger, a leadership change, a contract award, an office opening. All four are already sitting in this pack as news items on the profile pages, each one silently ageing a cell in the rating table that nobody went back to change. The loop re-rates and reports only what moved, so the output stops being an order and becomes a change to an order, with the sentence that changed attached to it.

Most of the evidence such a harness needs arrives as pictures of text. Filings, contract PDFs, portal screenshots, and prior packs exactly like this one. The extraction that reopened this deck flattened a five-column rating table into a single stream of names and digits, which is exactly how a 3 gets separated from the column it belonged to and the reverse-engineering above becomes hand work. Vision models read the grid instead of the stream, keep the cell, and hold each citation as a rectangle on a page, so a rating traces back to the region it came from. That is the difference between a machine-assembled scorecard that is defensible and one that is merely plausible.

Then wire it forward. The axes that scored candidates become the schedule of the contract; the contract's obligations become the checks that delivery telemetry is read against; incident tiering, the thing this pack could only test by reading a support page, becomes a measured fact once work is flowing. Agents sit at each hop with a scope, a named owner and a log. The scorecard turns into an instrument that keeps running rather than a document that dies at handover.

What the loop must not decide is the weights. Nothing in the arithmetic knows that a bank buying run-and-maintain cannot accept basic support at any price, or that a pay premium of that size is fine when the skills are scarce and indefensible when they are not. Those are the buyer's calls, and the point of the machinery is to hand a person fewer of them, better posed.

The method in this engagement was right and its cadence was wrong. Writing the five counting rules down once is now a day of work, and after that the loop reruns them every week for the life of the relationship. Judged that way, the scorecard here was not late because the analysis was slow. It was late the moment it printed, because the analysis could only happen once and the vendor set kept moving after it stopped.

NEXT STEP

Ready to make AI real?