When a delivery goes wrong, the first artefact everybody reaches for is the status pack. It is the wrong artefact, and not because people lie in status reports. It is the wrong artefact because a status report was written to tell somebody something, and anything written to tell somebody something is shaped by its reader.
The bills were written to get someone paid. A narrower purpose, and a much stronger guarantee.
I keep coming back to a review I worked on at a European banking group IT services subsidiary, which had put a customer-facing build out to an offshore software supplier and then commissioned an independent look at what had happened. Two exhibits from that review sit a page apart, covering the same ten months of the same project. One is empty. The other is the closest thing to delivery telemetry the engagement produced, and nobody built it on purpose. It fell out of the finance system.
Two exhibits, one window, and only one of them populated
The review obtained fifteen classes of artefact and marked every one both evidenced and reviewed, among them contract and statement of work, requirements, designs, the plan, weekly reports from both parties, the risk, issue, change, defect and code fix logs, the test plan, the communications plan, and the monthly bills for supplier headcount. Access was never the constraint. Nothing on that register was withheld.
So the empty grid is not an evidence gap. Both parties' weekly reports were in the room. What stopped it being filled was that the two sides did not share a definition of what the colours meant, and the review says so directly: differing definitions, differing understandings of degrees of severity. The open question the reviewers wrote against that exhibit, paraphrased only to keep it short, was whether the project was genuinely red and blocked, or whether the supplier was making progress that the ratings did not reflect.
That question cannot be settled from the ratings, because it is a question about the ratings. Any month-over-month series built on top of them inherits the ambiguity and hides it, because the series looks perfectly well formed.
The billing table has no such problem. A billed day is a quantity two organisations agreed on, one of them while writing a cheque. It arrives monthly, on a fixed calendar, in a fixed unit, through a process with controls around it. If you were specifying a delivery event log from scratch you would ask for those properties, and you would not get them.
What the billing table says that the plan did not
Read as a shape, the five populated months tell a story the plan never told. The first month is almost nothing: 18 person-days across four lines, all of them senior. That is a mobilisation month, and mobilisation months are supposed to look like that. The second month is where the build population arrives, and it arrives partially: ten lines billing, 121 days, with most of the new developer lines at 4 to 7 days apiece.
By the third month those same developer lines are at 18 to 20 days each, and across months four and five at 18 to 22, a full month of working days. The monthly total never comes back down. Month four adds a second wave and the total jumps to 228 days: two of the four arrivals bill a full 20 days, the other two 14 and 10. Month five holds at 236, with two more lines appearing at 3 and 2 days each, people joining at the very end of the window.
The curve is not uniform, and the exceptions are worth as much as the trend. One team-lead line runs 22 days, then 12, then 8, then back to 22, while the developers beside it are at full stretch. Whatever that was, it happened in the open, monthly, in a document somebody approved for payment.
Set that against a detail from the document review. The plan existed in two variants, one for the expected case and one for the adverse case, and the team stayed on the expected-case variant. That is a normal thing to do and a normal thing to regret. The bills show what staying on it cost: an effort curve that never flattens, back-loaded so heavily that the last two of five months carry more than the first three combined.
A plan that has not been rebaselined stops being a measurement. It is an intention formed at the start. The bills carry no intention at all, which is what makes them useful.
The line that stopped billing
The most useful cell in the whole table is not a number. It is a dash.
One senior technical line bills in the first two months, 4 days and then 18, then carries a dash for every remaining column. The review's comment against the monthly bills notes that absence, and it is the only remark recorded against that artefact: of everything there was to say about a stack of invoices, the thing worth writing down was who had stopped appearing on them. Whether it also reached the risk log or the issue tracker I cannot tell you from the exhibits in front of me. What I can tell you is which document surfaced it.
That is the pattern I would build an alert on tomorrow, on any programme, in any sector: a capability that was staffed and is now not staffed, detected as a change in the shape of a recurring payment. It needs no cooperation from the delivery team, no new tooling in the supplier's estate, and no agreement on what red means. It needs somebody to notice that a line item that appeared every month has stopped appearing.
Nothing about that is difficult. It is that nobody owns it. Finance reconciles the invoice against the purchase order and the rate card. Delivery reads the status pack. The signal sits in the gap, in a system neither of them thinks of as a delivery system.
- 27
- Individual lines on the supplier's billing roster, across a grid ten months wide
- 762
- Person-days billed in the five months that carry figures
- 13
- Lines billing in each of the two heaviest months, up from 4 in the first
- TBC
- Every populated cell of the status-rating timeline for the same window
Why I file this under data readiness rather than governance
There is a version of this argument that ends at better vendor management. It is not wrong, but it stops short.
The useful property of billing data is that it is definitionally stable in a way almost nothing else in an operational estate is. Every data readiness assessment I run ends up sorting the estate into fields that mean the same thing every month and fields that mean whatever the person filling them in thought at the time. A billed day sits firmly in the first group. It has a unit, an owner, a counterparty with an opposing interest, an approval step, and a retention policy somebody in finance enforces. Its lineage is not a diagram anyone had to draw. It is the payment run.
So supplier billing is one of the few sources in a large organisation where you can build a monthly series across years and trust the comparison across the span. Most of the estate cannot. Categories get redefined, tooling changes, a team renames its stages, and a feature quietly changes meaning halfway through the training window without anything failing.
That is also why process mining took hold on purchase-to-pay before anywhere else. The event log was already there, timestamped and reconciled, because the money forced it to be. Delivery effort is the same class of asset and nobody treats it that way.
A status colour is an opinion somebody typed. A billed day is a number two organisations checked, one of them because it was paying. Only one of those survives contact with a machine learning pipeline.
What it does not tell you
Billing tells you presence, not progress. It has no opinion on whether the work was any good, which is why the reviewers read the code separately. It carries the supplier's own view of who was on the project, so it is only as accurate as their timesheet discipline. And it lags: a monthly invoice is monthly resolution, fine for staffing shape and useless for anything measured in days.
None of that argues against using it. It argues for using it as one channel of several, which is how telemetry works everywhere else. Presence from the bills, change and defect volumes from the tooling, acceptance from the governance record, sentiment from the interviews and held loosely. Build the hard series first and use the soft evidence to explain it.
The modelling end of this is close to trivial. A monthly matrix of lines against days fits in a spreadsheet, and the first useful outputs are a ramp curve, a departure alert and a comparison against the plan the programme is still nominally running to. No pipeline, no deployment, no feature store. That comes later, when you have twenty programmes of history and want to know at month three which of them are ramping the way the ones that landed on time ramped.
Which is the point I would make to anyone budgeting for delivery analytics. The data you need for that is being generated right now, correctly, monthly, by your accounts payable function, and filed into an archive nobody reads. RealAI's Platform team spends much of its first month on engagements like this one recovering series that were never lost, only never assembled.
Figures are as recorded in an independent review of an outsourced software build commissioned by a European banking group IT services subsidiary: its staffing and billing exhibit, its status-rating exhibit and its document review. That review produced findings and recommendations, not delivered results, and the draft I worked from carried unfilled sections, described as such above. Reading a billing table as delivery telemetry is our framing, not the review's.
“A status colour is an opinion somebody typed. A billed day is a number two organisations checked, one of them because it was paying. Only one of those survives contact with a machine learning pipeline.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
