Outcome pricing rarely dies in the negotiation. Both sides usually want it there. It dies later, in the room where somebody has to sign, because the person signing is not being asked whether the deal is clever. They are being asked for a maximum.
Everything written about paying for results argues about the shape of the incentive: how much to expose, against which target, released on what event. Almost nothing argues about the ceiling, which is the term that decides whether the contract reaches a signature at all.
I was reminded of that reading back through a post-mortem review I led of a failed offshore software build at a European banking group's IT services subsidiary. The subsidiary had contracted an offshore supplier to build a regulator-driven customer-facing application, escalated on code quality, granted a grace period, terminated the statement of work before the system reached user acceptance testing, and took the build back in-house. Of the ten client questions the review answered, nine came back negative or mixed. The single clean yes was that the requirement specifications met common international practice.
One of its recommendations was to attach a future supplier's incentive to the outcome rather than to the volume of effort. Its appendix then set out, marked illustrative throughout, what such an arrangement could contain. Nobody signed those terms and nobody operated them, which is why they are worth reading: they are what a careful reviewer writes down when no deal is on the table.
What the approval committee is actually asking
A procurement committee is not a party to the delivery argument. It cannot judge whether four percent is generous for a fifteen percent discount, and it is not trying to. It has one structural question, which is what this contract can cost at its worst, and one procedural one, which is whether that figure can be provisioned and reported against.
A cap answers both in a line. Total rewards capped at 120 percent of the contracted price gives the committee a maximum expressed as a multiple of a number already in its budget. The committee can write that down, hold it against a mandate, and move on. That half the supplier's fees sit at risk on the other side of the ledger is upside; it is not what the approval turns on.
Remove the cap and the sentence cannot be written. A share of a realised benefit has no maximum expressible in advance, because the benefit has none. That is not an objection about a supplier earning too much. The form of the commitment does not fit the form of the approval, the committee is being asked to authorise a number it cannot state, and its only lawful answer is no.
Which produces the result everyone in sourcing has watched at least once. The most sensible construct on the table, the one where the supplier eats the consequence of failing, is refused, and the deal reverts to day rates that transfer nothing. The mechanism was not rejected on its merits. It arrived without a bound.
The same defect that made cost-plus objectionable
There is a symmetry in the appendix that I think is the whole lesson.
The arrangement it criticises is a margin on spend, where the supplier is also asked to negotiate the client's discounts down. The objection is not that three percent is high. It is that the supplier's earnings rise without limit as the spend rises, so the party best placed to reduce the spend is paid for not reducing it. Unbounded, pointing the wrong way.
Paying a share of a realised benefit is unbounded pointing the right way. The incentive is aimed correctly, and a buyer who gets more benefit is happy to pay more for it. The governance defect is identical: neither construct lets anyone state the total in advance, and a buyer fresh out of a failed build knows what it costs to sign a commercial shape whose total depends on how events unfold.
The cap converts the second into something approvable while keeping its direction. Half the fees exposed to the downside, rewards bounded by a stated multiple of the contracted price and no further. The supplier keeps a real reason to deliver. The buyer keeps a number.
No baseline, no reward, no matter how well drafted
The cap makes the deal approvable. The baseline makes it payable, and the appendix is careful about this in a way most commercial drafting is not.
An outcome-linked element is tied there to three things the buyer supplies before signature: how much control over delivery the supplier actually has, whether a baseline exists to measure improvement against, and whether both sides have agreed a scorecard. The formula for calculating, timing and paying the result-based fee is settled at commencement, not discovered later, and the payment profile follows completion of stages, stated as percentages of the contracted price.
Read those as a filter. When a review has to record that the delivered software did not meet the requirement specification, that the development process was not compliant with usual procedures and that the output was not delivered as specified, the instrumentation needed to compute a benefit share fairly is plainly not in the building. The trap is familiar: the situation that most wants outcome-linked payment is the one least equipped to operate it.
- 3%
- Supplier margin on spend in the criticised construction, at any spend level
- 2% to 4%
- Band the margin flexes across in the simpler option described
- 50%
- Share of fees the developed option places at risk
- 120%
- Ceiling the illustrative option puts on total rewards, as a share of the contracted price
Why this sits on my desk again now
Every serious proposal for a copilot deployment crossing my desk this quarter carries a version of the same offer: pay us out of what it saves you. Handling time on a service desk. Deflection into self-service. Throughput on code review. The offer is honest and it beats a day rate, because the supplier is the one claiming a benefit and should be the one exposed to it.
It reaches a committee and fails there for the two reasons above.
There is no maximum. A share of a saving on a process with no measured pre-state is a number nobody can state before the fact, which is the unbounded shape again.
And there is no baseline. Handling time before the retrieval-augmented assistant went in, counted the same way it will be counted afterwards, is a figure most organisations discover they do not hold. The evaluation set that would tell you whether answer quality moved has to exist before the deployment, and lineage on the underlying figures has to be good enough that the supplier's own reporting is not the source of record for its own fee. Unglamorous MLOps work, the part of a RealAI Platform engagement that lands before anyone models anything, and the difference between an outcome contract and an argument.
Political agreement on the EU AI Act has already pushed documentation, record-keeping and monitoring up the agenda for systems of this kind, so the apparatus an outcome-linked fee needs will be built for other reasons anyway. Build it first and the commercial construct becomes available. Build it afterwards and you are negotiating the meaning of a number while the invoice sits on the table.
A cap is not the buyer holding money back. It is the buyer being able to say in one sentence what the worst case costs. Without that sentence there is no approval, and without approval there is no outcome pricing at all.
What I would put in the paper
Four terms, in this order, because each is unapprovable without the one before it.
A baseline, measured before anything is deployed, on an instrument both sides can run. A scorecard naming what is being paid for, agreed at commencement rather than at the first invoice. An exposure large enough that the supplier notices losing it. And a cap, stated as a multiple of the contracted price, so the worst case is a sentence and not a scenario.
Suppliers resist the fourth term more than the third, which tells you how these conversations go. A cap on the upside is the price of being allowed to play at all, and a supplier who will not accept a bound on what the arrangement can pay does not believe the baseline either. Nothing in the review says a capped construct was operated on that engagement. What it says is that this was the shape worth reaching for next time, and the ceiling is the part that would have made reaching for it possible.
Drawn from the commercial appendix of a post-mortem review of a failed offshore software build at a European banking group's IT services subsidiary, which I led. Every construct described there is marked illustrative; none was a term of that contract, none was signed and none was operated. The review produced findings and recommendations, not results. Figures are as recorded in it. Reading the reward ceiling as the term that decides approvability is ours.
“A cap is not the buyer holding money back. It is the buyer being able to say in one sentence what the worst case costs. Without that sentence there is no approval, and without approval there is no outcome pricing at all.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
