Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsClaims

Counting the Same Benefit Twice

RealAISep 3, 20248 min read
ClaimsPortfolio GovernanceEnergyBenefits CaseDelivery AssuranceAI Strategy

Add every approved business case in a large organisation together and you get a number. In most organisations nobody has ever produced that number, and when somebody finally does, it comes out bigger than the organisation. More cost taken out than the cost base holds. More hours released than the people work. Every case is defensible on its own page. The total is fiction.

That is not an arithmetic error. Arithmetic errors get caught. This one survives because each case is correct within its own boundary, and the boundary is exactly where the checking stops.

The clearest statement of it I have read sits in the appendix of an independent readiness review of a sales and billing platform programme at a European energy retailer. The review was a four-week check on whether the programme was ready to move out of start-up and into delivery, carried out by an outside assurance firm together with an internal audit team, on nineteen interviews and a register of programme documents. Its headline findings were the ones you would expect from that kind of work: scope still open, the chosen platform unproven, no guiding design.

The line I keep going back to is not in the findings at all. It is in a seven-row appendix table on organisational readiness, marked as a working draft, against the last row, the one about how investments are governed above the level of any single project. The answer records that no overall structure of that kind exists. Underneath it sits the risk, in the reviewers' own words:

RISK: 'Double' counting of business benefits in business cases of different project and programs. Our review hasn't included this in the review of the business case.

Two sentences. The first names a failure that can invalidate an entire investment portfolio. The second says the reviewers did not test it.

The finding and its disclaimer

The disclaimer is the more useful half, and I want to defend the reviewers before drawing anything from it.

They had been asked whether one programme was ready. Double counting is not a property of a programme. It is a property of the set of programmes, and nobody had commissioned a review of the set. Inside the scope they had, the only two options were to stay silent or to write the risk down and mark it untested. They wrote it down. Most reviews would have left it out, because a finding you cannot evidence is a finding a client can argue with.

What that leaves is a named, unquantified risk on a page in an appendix of a draft deck. There was no standing body whose job was to pick it up, which is the same absence the finding itself describes. The report can only recommend into the programme it was pointed at.

Five benefits on one programme

The condition that makes double counting likely shows up earlier in the same report, on the context page.

The programme was carrying five benefit ambitions simultaneously: reduce cost to serve, reduce cost to acquire, increase market share, increase top-line revenue, increase profit. One platform build, five levers.

Load five different benefit types onto one initiative and at least some of those levers are shared with every other initiative in the estate. Cost to serve is not owned by a platform rebuild. It is the number that three other efforts are also promising to move, and each of them has a spreadsheet showing how.

Then the tracking problem, from the findings section of the same report: the steering committee had no levers to monitor and control the benefits case yet, and the benefits case had no clear linkage to the programme phases. So no benefit could be bound to a stage and checked when that stage closed. A benefit that cannot be checked at a gate inside one programme has no chance of being reconciled against an identical benefit claimed in another programme's case.

The document register carries the tell. Three separate business case artefacts carry a single document date: a business case, a business case calculation, and a versioned business case document sitting at its forty-seventh version. The development planning document was on version fifty-six. A case on its forty-seventh revision is not being refined. It is being refitted to a scope that has not settled, and every refit quietly changes which benefits it claims and how large they are.

The same failure, in a second currency

One row earlier on the same appendix page, the subject shifts from how the portfolio is governed to how it is actually run. The answer notes that this sits outside the review, and then records a risk: the platform organisation had no projects and no development capacity left, because all available resources had been absorbed into this one programme.

Read those two lines together and you have the same failure twice, denominated differently. The benefits were at risk of being counted more than once because nobody held a portfolio benefit ledger. The engineers had already been counted more than once because nobody held a portfolio capacity plan. In both cases every local plan is sound and the aggregate is impossible.

Capacity double counting is the easier of the two, because it fails loudly and early. At some point a named person cannot attend two stand-ups and somebody escalates. Benefit double counting fails quietly and late. Both initiatives deliver, both declare their saving, the cost base does not move by the sum, and the argument about which number to believe starts long after the funding decision those numbers were written to justify.

7
Rows in the organisational readiness appendix table
5 of 7
Rows carrying an explicit RISK label
2 of 7
Rows showing movement between the two review passes
9 of 12
Recommendations recorded as not addressed yet

The version of this I am seeing now

Nearly every organisation I speak to has more than one generative pilot in flight. A retrieval-augmented assistant over policy and product documents in one function. A drafting copilot in another. A code assistant in engineering. A deflection experiment in the contact centre, plus one or two early and carefully scoped autonomous experiments in a back office. Four or five business cases, written by four or five sponsors.

Three of them, at least, book time saved from the same population of people, in the same unit: hours per week, per head. Sum the promised hours and the adviser's week comes out longer than a week. Nobody notices, because the cases were written against four different baselines, approved in four different forums, and reconciled in none.

This is worse now than it was for a platform programme, for a reason that has nothing to do with the technology. A programme costing tens of millions gets an investment committee, a business case register and an audit trail. A copilot pilot for forty users gets a line in a departmental budget and no portfolio scrutiny whatsoever. Cheap initiatives are precisely the ones that escape the ledger, and there are now a great many more of them.

An honest AI benefit case names three things: the operational number it will move, the baseline it is moving from, and every other initiative currently claiming that same number. The third is the one nobody writes, and it is the only one that makes a portfolio add up.

Two initiatives can book the same saving and both be right inside their own boundary. The boundary is where the checking stops, and the business case ends up adding to more than the business.

What we would fix, and in what order

None of this is analytics, which is the awkward part, because it all has to land before an evaluation set or a vector store is worth arguing about.

One benefit ledger for the whole estate, held above any single programme, with a named owner who is not a programme director. Each benefit line bound to one operational metric and one stated baseline, so that two claims on the same metric collide visibly instead of sitting quietly in separate documents. Each benefit attached to a delivery stage, so it can be checked when that stage closes rather than at the end of everything. A standing rule that no new case is approved until it declares which existing claims it overlaps and how the overlap is split. And the capacity mirror of the same ledger, with named people counted once across the plan, so no team can be fully allocated twice.

That is bookkeeping, not modelling, and it is why we ask which number you want to move before we ask which model you want to build. RealAI's Platform team treats the benefit ledger as part of data readiness rather than as finance homework, because the question is the same lineage question we ask of any figure entering a pipeline: where did this come from, what does it mean, and who else is using it to mean something slightly different.

The reviewers I have been quoting did the right thing inside the scope they were given. They named the risk and marked it untested, which is more than most assessments manage. Nobody had asked them to own the total either, and that is the whole finding.

Details are as recorded in an independent programme readiness review of a sales and billing platform at a European energy retailer, carried out over four weeks by an outside assurance firm together with an internal audit team. That review produced findings and recommendations in draft, not delivered results, and the work was not ours. Reading its portfolio governance line as an investment-arithmetic problem is.

Two initiatives can book the same saving and both be right inside their own boundary. The boundary is where the checking stops, and the business case ends up adding to more than the business.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?