Two organisations can look at the same piece of software work, name numbers that differ by a factor of three, and both be reporting honestly. That happens whenever neither number was computed from anything. A disagreement about effort resolves only if the two parties share a unit. Without one, the argument does not converge. It just spends calendar.
A European retail and corporate banking group commissioned an independent review of an application build it had contracted to an external development partner and then brought back in house after quality and schedule failures. The review covered project management, delivery approach and the quality of what was handed over. One finding in it is the cheapest lesson in the whole document, and the one most often waved past, because it presents as a procurement complaint rather than as an engineering defect.
The challenge
The complaint, as it arrived, was that the partner's estimates were too high. Not marginally. The perception recorded inside the client organisation was that estimates sometimes ran two to three times the right amount, and one interviewee recalled a figure five times what had been estimated initially. Both of those numbers have to stay in the register they were collected in. The first is a perception held across a client team. The second is one person's recollection in an interview, and the reviewer's own annotation on the page marks it as an eyebrow-raiser rather than a measurement. Neither was ever reconciled to a ledger.
The mechanism behind the partner's numbers was written down plainly. Estimates were produced by assigning roughly the number of days a piece of functionality or activity was likely to take, based on gut feel. That is a defensible way to quote work when the estimator has done the same work before. It is not a defensible way to defend a quote, because there is nothing underneath it to inspect.
Now the symmetric half, which is the part that makes this a case study rather than a grievance. The phrase the client used was the right amount. No definition of the right amount existed. The review named two candidate causes for inflated estimates, and the first one it named was on the client's side: the client did not use a well-defined methodology to estimate effort, which made its own estimates inconsistent from project to project. The second candidate was that the partner's designs were sub-optimal. Order matters in a list like that, and the reviewer put the client first.
The consequence stated alongside it is the non-obvious one. Absent a defined method on the buyer's side, the supplier is prevented from learning and improving as an organisation. There is no signal to improve against. Every quote is judged by whether it feels large, every rebuttal by whether it sounds confident, and the supplier's estimating capability stays exactly where it started, engagement after engagement. The buyer's measurement gap becomes the supplier's capability gap, and the buyer pays for it twice.
There is one more tell, and it sits in the working notes rather than in the findings. Against the description of how estimates were produced, the reviewer wrote an instruction to ask both sides how they arrived at their numbers. At draft stage, with interviews done and documents read, the mechanism on the client side had still not been established well enough to write down. That is the finding underneath the finding.
The plan movement on the engagement reads differently once you accept all this. The work was allocated 2,576 person days in total, of which 996 covered the development phase, and a later replan put the total above 3,047. In an organisation with an agreed unit, that movement is a measurable event with a cause: scope entered, or the sizing was wrong, and you can say which. Here it is neither evidence of padding nor evidence of honest discovery. It is two uncomputed numbers, one after the other.
The approach
The recommendation was to move to a defined sizing method, function points or use case points. That part is unremarkable and could have been written by anyone. The sequencing around it is where the value sits, and it is a deliberate refusal to hand the client a weapon.
Learn the method first. Apply it to a number of projects from earlier releases. Compare what the method yields against the estimates the partner actually gave. Work out where the differences come from and why. Only then treat it as the standard. Four steps before a single new quote is judged against it.
The stated reason is that estimation methods provide a level of objectivity but still contain elements that depend on judgement, and that judgement assumes expert knowledge. On this engagement the team was working with a front-end technology it had not used before and a delivery approach it had recently adopted, so the judgement calls inside any sizing method would have been made by people with no calibration for either. A method imposed in that state does not produce agreement. It produces a more elaborate argument, with arithmetic in it.
This is the point at which a governance recommendation turns into something we can actually build. The calibration set the review asked for is a data engineering job, not a workshop. Work items, the changes that closed them, the test and deployment records attached to each, the elapsed effort against each: all of it is already emitted by the systems the delivery runs on, and almost none of it is joined. Process mining over that workflow reconstructs what a unit of work has historically cost in this organisation, on this stack, with these teams, which is the only baseline that means anything. Feature and lineage discipline is what makes the resulting number defensible: an estimate you can recompute from stored inputs, and trace back to the history it was derived from, survives a commercial dispute. An estimate produced by a model nobody can re-run is a gut feel with a confidence interval printed on it. Platform work of ours starts at that join, because a sizing method without a calibrated history behind it is a template.
The outcome
What this engagement produced was a diagnosis and a sequence, delivered in draft. The perception figures stayed perceptions. No back-test against earlier releases was run inside the review, no method was adopted, and the two candidate causes for inflated estimates were still candidates when the document closed. Read it as a review of how the disagreement was structured, because that is what it is.
The transferable part has aged well, and one thing about it has sharpened. Machine learning has moved into delivery estimation on the vendor side, and the early and cautious pilots of language models in development work now come with throughput claims attached. Those claims meet the same wall the estimate dispute did. An organisation with no agreed unit of work cannot tell a productivity gain from ordinary variance, cannot tell either from scope drift, and ends where this dispute ended, with two positions on the table and nothing in the room able to adjudicate between them. Our Consult engagements now open by asking what unit the organisation measures work in and where that unit is computed, because the answer decides whether any later number can be argued about at all.
The estimate was never the problem. The absence of anything both sides could compute was the problem, and an argument built on that absence cannot be won, conceded or closed. It can only be carried, at delivery's expense, until someone builds the thing that would settle it.
