Every few years a buyer decides that paying for effort is absurd and that they would rather pay for results. The argument never changes and it is sound: paying by the day pays the supplier to consume more days. What follows is less clean, because the moment both sides agree money should be at risk, somebody has to say how much. How much is not a detail of the mechanism. It is the mechanism.
The sharpest statement of that problem I have worked with sits inside a post-mortem review of a failed outsourced software build at a European banking group's IT services subsidiary. The subsidiary had contracted an offshore supplier to build a regulator-driven front-end application, escalated twice at steering level, first on code quality, terminated the statement of work before the system reached user acceptance testing, took the build back in-house, and commissioned an independent review of delivery approach, project management and engineering practice. That review produced findings and recommendations, not results, and nothing described below was ever run.
One recommendation was to put a risk and reward mechanism in place for future work of this kind, so that the supplier's incentive would attach to the outcome rather than to the volume of effort. Under potential risks with such models, the review sets down the tradeoff as a stated principle about outcome-based deals generally: "Larger the pie, bigger the risk; smaller the pie, less motivation." Eleven words, entirely correct, and then the document moves on without sizing the pie.
Two sizings of the same idea, one appendix apart
Put the two examples side by side and the gap is the whole article.
The first is a procurement mechanism. A supplier is paid cost plus a fixed three percent while also being responsible for negotiating discounts from third parties: the larger the bill, the larger the fee. The repair is small and precise. Set a target discount at a benchmark, ten percent in the example, and hold three percent as the neutral case. If the supplier lands only five percent, the margin drops to two; if it lands fifteen, it rises to four. The at-risk amount is a third of a three percent margin, which on a one million euro spend is ten thousand euros against a fee of thirty thousand.
The second is a different animal. Half of all supplier fees at risk against the product working and the targeted benefits landing, half of an initially negotiated discount released as the reward for on-time delivery against acceptance tests, further rewards for efficiency gains, and total rewards capped at one hundred and twenty percent of the contracted price. The document is candid that this suits a long-term relationship, and sketches a further step where the supplier is paid entirely out of realised benefits.
What sets the ceiling, and what sets the floor
The band has two walls, of different material.
The ceiling is survivability. Past some point, losing the at-risk portion does not sharpen a supplier's behaviour, it replaces that behaviour with defensive scope arguments, change requests as a revenue channel, and the best engineers rotated onto accounts where the money is safe. Fifty percent of fees is not fifty percent of margin: on a delivery business running a normal services margin, half the fee is several times the entire profit on the account, so once the outcome looks doubtful the rational move is to stop spending on it. A penalty large enough to be existential converts the supplier from an aligned party into a litigant.
The floor is salience. Below some threshold the at-risk portion costs less to lose than it costs to chase, and becomes an accounting entry rather than an incentive. No delivery director restructures a team over a sum that rounds away in monthly variance. The ten thousand euros in the first example bites only because it sits against a thirty thousand euro fee on one tightly scoped outcome. Move the same absolute number onto a multi-year build and it disappears.
Both walls are properties of the supplier's economics on that account: its margin, its cost base, how much of its quarter the engagement represents. None of that is in the buyer's spreadsheet, which is why sizing done on the buyer's side alone lands at a token percentage or at a punitive number the supplier signs and manages around.
You cannot size what you cannot measure
Suppose the band is found. It still cannot be written down, because the clause is a bet on a number both parties have to read the same way twelve months later.
The same review is clear on this without connecting it back to the pricing recommendation. It advises defining pass and fail criteria and project status definitions jointly and up front, and joint delivery performance indicators so alignment is tracked rather than asserted. And it records the estimation gap plainly: the subsidiary had initially estimated roughly a thousand person-days, while the supplier's figure ran past two and a half thousand. The review sets both figures down without adjudicating either, and elsewhere notes that it found no evidence of the method behind the estimate that was produced. That is the worse half. The two sides were not only far apart, they had no shared way of showing their working, and that is the disagreement that makes an outcome-priced contract unsignable. Parties who cannot agree what the work is cannot agree what finishing it looks like, and the at-risk portion has nothing stable to attach to.
Then there is the entry test. The pricing guidance says the approach works where both parties share responsibility for achieving and measuring the result, and where an established, mature relationship is already in place. The review's findings describe the opposite, and its remedy was to have key staff, architects in particular, spend six to twelve weeks sitting with the other side. A commercial mechanism was being recommended into a relationship that had just failed the mechanism's own precondition: reasonable in a forward-looking recommendation, dangerous to forget when the clause is drafted.
- 3%
- Fixed supplier margin, illustrative model, flexing 2% to 4%
- 50%
- Fees at risk, complex illustrative model
- 120%
- Cap on total rewards, share of contracted price
- ~1,000 vs 2,500+
- Person-day estimates, subsidiary against supplier, as recorded
Why the band narrows as delivery gets cheaper
All of that predates the current round of tooling, which makes the problem harder rather than easier. The marginal cost of producing software is falling. Copilots write the boilerplate. Retrieval-augmented generation puts the specification, the architecture note and the existing codebase in front of a developer without a two-week ramp. Evaluation sets catch regressions that used to reach a human reviewer, and early, carefully scoped autonomous experiments handle narrow, well-fenced tasks at the edges of a build. None of this removes the engineer, and any supplier claiming otherwise is selling. It does mean the person-day is drifting away from the thing the buyer values.
That drift is what pushes buyers toward outcome pricing, and it is also what compresses the band. A supplier whose cost base has shrunk has less absolute margin to put at risk on the same contract, so the survivability ceiling comes down in cash terms at the moment the buyer's appetite goes up. The floor rises at the same time, because the outcomes worth contracting for sit further from the keyboard: a defect rate the business will accept, a process time the operation will feel. Both take longer to observe and cost more to dispute.
What is left to price is accountability, which is not cheap even when production is. Someone has to hold the lineage of the data a model was fitted on, keep evaluation sets honest and versioned, show the system behaves in month nine as it did at acceptance, and answer for it when it does not. Political agreement has been reached on the bloc's AI legislation, and when obligations follow, a buyer in a regulated sector will want a supplier answerable for system behaviour rather than for hours consumed. That is the work that survives the collapse in delivery cost, and the only work worth attaching an at-risk fee to.
The ceiling is set by what the supplier can survive losing. The floor is set by what the supplier will bother noticing. Both are facts about the supplier's economics, and neither is in the buyer's spreadsheet.
How we would size it
Four moves, in our own words, not the review's.
Express the at-risk amount as a share of the supplier's gross margin on the account rather than of contract value, and make the supplier stating that margin a condition of the model. A number set against revenue tells you nothing about whether it bites or breaks.
Set the floor by a salience test rather than by convention. If the at-risk sum would not on its own justify a delivery director changing the staffing, it is decoration.
Attach the fee only to outcomes both parties read off the same system, with the definition and the query that produces it stored together. An outcome each side computes in its own tooling is a future dispute wearing a contract's clothes. Cap the upside too, since an open top invites the supplier to optimise the measured thing at the expense of everything unmeasured.
Start small on the first release and widen on evidence: the first release buys measurement, not savings. Skipping that is how a sound mechanism gets discredited by one bad year.
None of that needs a new contract template. It needs knowing what your numbers mean before you attach money to them, which is where a RealAI Platform engagement starts.
Figures and quoted language are as recorded in an independent post-mortem review of a failed outsourced software build at a European banking group's IT services subsidiary and its appendices. That review produced findings and recommendations, not results, and the illustrative examples were never contracted. Reading its risk and reward recommendation as a sizing problem is ours.
“The ceiling is set by what the supplier can survive losing. The floor is set by what the supplier will bother noticing. Both are facts about the supplier's economics, and neither is in the buyer's spreadsheet.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
