Every conversation I have had this quarter about coding assistants ends in the same place. Somebody produces a percentage. Twenty percent faster, thirty, occasionally a figure north of fifty that nobody in the room will defend out loud. Ask what it is faster than and the conversation moves on to tooling.
This is not dishonesty. It is arithmetic with a missing operand. A percentage improvement is a subtraction, and a subtraction needs two numbers. Almost none of the organisations quoting one has the other, because recording what normal looked like is unglamorous work that has to happen before anything interesting starts.
I keep going back to a post-mortem review I led for a European banking group's IT services subsidiary, on an offshore-built advisory application stopped before it could pass acceptance. Nothing in it concerns machine learning. It is the clearest case I have worked on of what a missing baseline costs.
Ten questions and one clean yes
The review ran on documents and people. Sixteen classes of project artefact were pulled and read, from the statement of work and the master agreement through the requirements, the design documents, the plan, the weekly reports from both sides, the risk and issue logs, the test plan, the change log and the defect and code fix logs. Fifteen interviews were run across both organisations. Two builds, taken about ten weeks apart, were read by a technical architect.
That was enough evidence to reach firm conclusions. It was not enough to reach quantitative ones. The closing summary is a list of ten questions with a one-letter answer beside each. Eight no. One yes and no. One yes.
Read that page as a client and it is a verdict. Read it as anyone who has to improve a delivery organisation afterwards and it is a problem, because a verdict does not tell you where to start. Not compliant, by how far? Not maintainable, compared with what the same teams produce on work they do well? The review could say the requirements specification was the one artefact meeting common international practice. It could not say whether that was normal for the subsidiary or unusually good, because nothing similar had ever been assessed.
The measurement that got recommended at the end
Behind the recommendations sits the work that should have run at the beginning. Before you change how software gets built, establish where the building currently stands and put numbers under it, so that anything later claimed about output, schedule, cost, quality and the satisfaction of the people receiving the work has something to be subtracted from.
I want to be exact about the tense, because it is the whole point. Nobody measured those five things on this project. They appear in the closing pack as a recommended first move for the work that would follow, written after a delivery had already failed. There is no before-state in the file and no after-state, and no line in this article reports what the baseline showed, because the baseline was never taken. The recommendation is real. The measurement is absent, and its absence is the finding.
What the five would mean in practice, stated plainly and in our own terms:
Output per unit of effort. Accepted software per person-week, at a scope unit defined once and never renegotiated. Skip it and every later claim about speed is an anecdote with a decimal point.
Schedule accuracy. Planned against actual, per unit of scope, kept as a distribution rather than an average. Skip it and a plan that slips looks like bad luck rather than a systematic estimating bias you could correct.
Cost per unit of that same scope. Not total spend, which only tells you how big the project was. Skip it and a change of supplier, method or tooling can never be priced.
Quality. Defects per unit of scope, where each was found, and what it cost to fix at that stage. Skip it and you cannot separate a build that was rushed from a build that was checked by nobody.
Satisfaction of the people receiving the work. Asked the same way, of the same population, on a fixed cadence. Skip it and morale becomes a rumour raised once the contract is already in dispute.
Five numbers. None of them needs a data scientist, and all of them have to exist before the intervention you want to argue about.
A baseline is mostly a definitions exercise
The instructive part of that review is not that numbers were missing. It is that the numbers present could not be trusted, and always for the same reason: two parties using one word for two things.
The delivered application was described as seventy percent complete. That figure is a claim, not a reading. Seventy percent of the requirements as originally written, of the backlog as it stood, of the screens drawn in the design documents, or of somebody's honest impression on the day, are four different numbers, and the file does not settle which was meant.
Status colour, as above, meant one thing on one side of the contract and another on the other. Both sides reported diligently, monthly, for ten months. The reporting was compliant and the aggregate was meaningless. The recommendation there is definitional: write the definitions down wherever a term could be read two ways, particularly where language and working culture differ across a contract.
Testing split the same way. One side tested each increment against its own slice, in isolation. The other expected end-to-end coverage across dependencies. Both parties could truthfully report that testing had been done. Neither had done what the other assumed.
None of this is exotic. It is what happens by default whenever a metric depends on a human definition and nobody writes the definition down. A baseline is roughly one part instrumentation to nine parts agreeing, in writing, what each counted thing means.
- 10
- Review questions, answered as verdicts
- 1
- Of those answered yes
- 70%
- Completion claimed, denominator never agreed
- 5
- Baseline dimensions recommended, none measured
Which is precisely where AI productivity claims sit
Bring that forward to what most engineering leaders are being asked to approve this year.
A coding assistant goes out to a few hundred developers. Retrieval-augmented generation is stood up over the internal repositories and design documents, with a vector store behind it. A team runs an early and carefully scoped autonomous experiment on a narrow, reversible task. Somebody is asked, three months later, what it bought.
The honest answer in almost every case is that nobody can tell, and the reason is identical to the one in a banking IT post-mortem. Suggestion acceptance rate is not productivity, it is a measure of how often a developer pressed tab. Perceived time saved is a survey with no control. Story points completed are not comparable across a quarter in which the estimating behaviour of the estimators changed, which is exactly what a coding assistant does to them.
The fix is not sophisticated, which is why it keeps not happening. Freeze a scope unit. Record the five dimensions for a quarter with nothing new switched on. Store each definition beside the query that produces it, so a number and its meaning travel together and lineage survives a change of team. That definition store is the first thing RealAI Platform engagements stand up, ahead of any model choice, because it is the one part of the exercise that cannot be added afterwards. Then deploy, and read the same five the same way.
There is a second reason to do this now. The same discipline that produces a delivery baseline produces the evidence trail a model needs: a defined evaluation set, a recorded before-state, lineage from a claim back to the data supporting it. The EU AI Act has reached political agreement, and the documentation habits it points at are ones an MLOps practice should already have. An organisation that cannot describe how fast it builds software will not enjoy being asked to describe how its models behave.
Suggestion acceptance rate is not productivity. It is a measure of how often a developer pressed tab, and it is doing service as evidence in decisions nobody will be able to revisit.
What the missing step actually costs
The subsidiary in that review got a clear answer to the question it asked. What it did not get, and could not have got, was a number describing how far its delivery capability sat from where it needed to be, because the instruments that would have produced one were recommended rather than run.
That is the cost, and it is worth being specific. Not a worse verdict. A verdict instead of a measurement, permanently, and no way to tell afterwards whether the next supplier, the next method or the next tool did any better. Everyone involved goes on arguing from anecdote, competently and in good faith, indefinitely.
The before-reading is cheap, boring and has no demo. It is also the only thing standing between a productivity claim and an opinion, and the window to run it closes the moment the new tool is switched on.
Drawn from a post-mortem review I led of a stopped offshore software delivery for a European banking group's IT services subsidiary: its artefact set, its interviews, its expert code read and its closing summary of answers. That review produced findings and recommendations, not delivered results, and the baseline discussed here was among the recommendations rather than among the measurements. Reading it as the precondition for every AI productivity claim now being made is ours.
“Suggestion acceptance rate is not productivity. It is a measure of how often a developer pressed tab, and it is doing service as evidence in decisions nobody will be able to revisit.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
