Two teams looked at the same specification for a piece of software and sized it. One came back with about a thousand person-days, the other with more than two and a half thousand. The contract commenced at 2,576 person-days, which is to say the larger number got signed, and the review that followed found no evidence of the method behind either figure.
What keeps me on this is not the ratio. Estimates diverge, and two and a half times is wide but not unheard of on a build with unfamiliar technology and a distributed team. It is what the parties did next: they left the disagreement open and settled it commercially.
The engagement was a failed software build at a European banking group IT services subsidiary, and I led the review afterwards. Of the ten questions put to that review, eight came back no. The one clear yes matters here: the requirement specifications met common practice. The specification was not what failed. What failed sat downstream of a number nobody could defend, inside a contract that put the whole consequence of that number on one side of the table.
What a person-day contract actually prices
Denominating a contract in person-days feels precise because it produces a number with a comma in it. It is not. A day rate prices an hour of somebody's attention and says nothing about how many of those hours the work takes, which leaves the only question with money in it sitting with the buyer. That is defensible when the buyer knows the work better than the supplier. It is indefensible when the buyer has just been told, in writing, that its own sizing is out by a factor of two and a half and neither party can say why.
Look at what it does to incentives once the build is underway. A supplier's good day is a long day. Rework bills. Discovery bills. A design document that never gets written frees capacity for something that does bill. None of that requires bad faith, and the review is explicit that both parties entered in good faith and both spent additional effort trying to make the thing work. The mechanism needs no bad actors, only a contract that pays for input while the buyer is buying output.
The six alternatives, and what each one moves
The review's appendix laid out the constructions available beyond paying by the day. I give them in my own words, and the list reads best as a menu of risk transfers.
Pay on delivery. Fixed fee or event-based, falling due when the client confirms delivery against agreed acceptance criteria, milestones or outcomes. It suits work where scope and deliverables are clearly defined, and that condition is the whole model. It moves estimation risk to the supplier, which is what you want when the supplier is the one claiming the scope is bigger than you think. Undefined scope simply gets priced into the fee, and you pay for the uncertainty anyway.
Pay on business outcome. Payment tracks either an agreed improvement in a business target or a proportion of value delivered. The review sets three preconditions: the supplier's control over delivery, a baseline, and an agreed scorecard. Miss one and the model is unworkable, because you are paying for a movement the supplier cannot cause, measured from a starting point nobody agreed.
Pay with fees at risk. A variant of the outcome model in which a proportion of the fee rides on agreed results. It works, per the review, only where both sides share responsibility for achieving the result and for measuring it, with the formula fixed before mobilisation. The discipline is not in the percentage. It is in agreeing the formula before anybody starts work.
Pay on quality and satisfaction. An agreed proportion up front, the remainder against a quality threshold and how satisfied the client is. Quality can be defined at project level or at individual team member level, which is a choice about how sharp the feedback is.
Pay nothing yet. Work at minimal or no cost as an investment in the relationship, typically pre-sales support or early scoping needed to take a project forward for approval, ideally tied to a commitment for later phases contracted under another model. A financing instrument, not a discount.
Pay a retainer for a service. For work delivered over a prolonged period, a fixed fee retainer based on an estimated amount of time. The supplier manages resources across a standing team for continuity, buying itself staffing flexibility and the client cost certainty.
What the parties did instead
None of those six was used to bridge the estimate gap. The clause that was added split savings from early completion evenly between the parties, with the supplier distributing its share among its project team, the stated intention being to expedite delivery.
It is a reasonable clause. But look at what it does not carry. It rewards finishing early against a baseline that was itself the disputed number, it is generous on the upside and silent on the downside, and it pays out on a schedule event rather than on whether the software worked, which is precisely what did not happen here. Two parties disagreed about how large the work was and answered with a bonus for finishing it quickly. The gap stayed in the contract, waiting.
- ~1,000
- Person-days, the client's own sizing
- >2,500
- Person-days, the supplier's sizing of the same scope
- 2,576
- Person-days at contract commencement
- 140
- Person-days for a design document the review could not find
What at-risk has to mean before it means anything
The appendix carried two worked constructions for putting money at risk, both marked illustrative, so read them as examples rather than terms anybody here signed.
The simpler one takes a supplier paid a margin on spend while also responsible for negotiating the client's discounts from third parties. At three percent, a million of spend earns it thirty thousand and ten million earns three hundred thousand, so its interest runs against the client's. The fix attaches the margin to the discount achieved rather than to the spend: hit the benchmark discount of ten percent and the supplier earns its usual three, come in at five and it earns two, reach fifteen and it earns four. Same money, opposite behaviour.
The more developed one puts half the supplier's fees at risk against the product working and realising its targeted benefits, counted as actuals such as switching off a legacy system or as leading indicators such as process improvement. An initial discount comes off the rates, half of it available back for delivering on time with working functionality, further reward follows for efficiency gains and new benefits, and the rewards in total are capped at a hundred and twenty percent of the day rate or fixed price.
The design point is the cap and the floor together. At-risk money has to be large enough that losing it changes the supplier's behaviour and small enough that the supplier survives losing it. Outside that band it is theatre, and both sides know it by the second month.
A day rate does not price the work. It prices an hour of somebody's attention and leaves the question of how many hours the work takes with the buyer. Every other model on the list moves part of that question back across the table.
The line item that settles the argument
Inside the contracted effort estimate was a line for a system design document at 140 person-days, meant to set out coding standards, design principles and the supplier's reasoning for how the software would be built. No evidence was found that it was ever produced.
I cannot tell you what the final overrun was, because the review does not record how many person-days were consumed before termination and I am not going to invent one. What it does record is a build that never got through user acceptance testing, and a verdict of no against usual procedures, against the contractual obligations on design, implementation and software quality, against the requirement specification, and against maintainability.
Now put the design document back in the frame. Under an effort contract that line was a budget to be spent. Under a delivery-based contract it would have been a deliverable with an acceptance criterion and a payment attached, and its absence would have surfaced as an unpaid invoice rather than as an architectural puzzle after termination. The pricing model does not only allocate risk at the end. It decides which failures are visible while there is still time to act.
What we would change first
Not the rate. Three things come before the rate.
Size the work on both sides with the same named method, and have the supplier apply it first to releases it has already delivered for you, comparing what the method yields against the numbers it quoted without one. The differences are the calibration. Where the figures still diverge, treat the divergence as the finding: two competent teams producing different numbers are describing different systems, and that is a scope conversation, not a haggle.
Then put the estimate risk with whoever claims to understand the work best: if the supplier says the job is two and a half times what you think, invite it to price the job rather than the days.
Then write acceptance criteria for every line item large enough to matter, documentation included, and tie payment to them. That is where a RealAI Platform engagement starts on a sourcing decision.
The estimate will still be wrong. Estimates are wrong. What a contract decides is who finds out first, and who pays for the finding.
Figures are as recorded in a post-completion review of a failed software development contract at a European banking group IT services subsidiary: its contracted effort estimates, its statement of work terms and the pricing constructions in its appendix. The two risk and reward constructions above are marked illustrative in the source and were not terms of this engagement. The review produced findings and recommendations, not results.
“A day rate does not price the work. It prices an hour of somebody's attention and leaves the question of how many hours the work takes entirely with the buyer. Every other model on the list is an attempt to move some part of that question back across the table.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
