Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

Build It Yourself, Twice

RealAIJun 22, 20247 min read
Delivery AssuranceBuild vs BuyProgramme GovernanceData ReadinessMLOps

Assurance reports mostly open with what is wrong. A findings list sorted by severity, each item traceable to a document that was late or a control that was missing. Read enough and they blur, because a defect list tells you the state of a programme without telling you why that state was ever likely.

An independent readiness check on a European energy retailer's rebuild of its business-customer platform, run over four weeks by internal audit together with an outside assurance firm, does something else. Before it reaches a single defect it names two decisions the retailer had already taken, and says plainly that the programme carried a high risk profile because of them.

Everything after that reads as consequence rather than accusation. It is the best structural move in the document, and a decade on it is the one I would most like to see borrowed, because the fork those two decisions sat on is the fork in front of every enterprise choosing how to build with language models.

Two decisions, named first

The first decision was organisational. The new business line was to be built by a small group of experts working outside the way the company normally ran programmes, which the review notes requires a new way of working, managing and governing. The second was structural. The retailer would build the platform rather than buy a commercially available one, and would form development teams that did the building themselves rather than contract a large system integrator and out-task the work, which the review calls the traditional way of executing large development efforts there.

Its word for the approach is disruptive, paired with high risk in the same sentence. Its later heading on the theme is more useful still: a programme with this profile requires a high risk appetite and measures for running and controlling it. You are allowed to want this. You are then obliged to instrument it.

That framing is why the document does not read as a prosecution. The reviewers are not saying the retailer chose wrongly. They are saying the choice set a risk level, the level was high by design, and the machinery for carrying it had not been built to match.

The case for building, as the programme itself made it

The review is fair to the decision and so should we be. The paper behind it argued that a bespoke platform can provide competitive advantage and may be hard for competitors to copy. That is a real argument, and in this sector it may well have been the right one. The setup was not casual either, and its preconditions had been written down: the very best people drawn from different business units so that cross-functional knowledge and a new culture came with them, a full mandate inside predefined guidelines, a separate location so the team could move fast, more internal IT capacity and a separate IT organisation inside the programme. A considered design, not drift.

What the reviewers could not find was the other half. The choice between doing it in house and doing it with a leading external system integrator was not documented, and no in-depth risk analysis of that choice existed. The strongest argument for the decision sat in a paper written for another purpose, and the risks it created had never been written down anywhere. The consequence is stated flatly: the retailer was now fully in control and responsible for execution, and its contracts with outside parties were input based, meaning it bought people rather than outcomes.

Control and responsibility move together. That is the trade, and a fair one. It becomes a problem only when one half of it is announced and the other half is not.

Every bill that arrived was a do-it-yourself bill

The findings that follow are not random.

The foundation had not been assessed or proven for the load it was about to take, and already generated more than fifty issues a month in its existing form. Extending it was estimated at roughly thirty percent new code, and the review is careful to call that an estimate needing supporting analysis. Velocity under the new way of working was unknown, so the plan rested on an assumed rate, and programme management's own estimate put the resulting planning reliability at about half.

Then the people. A small expert group was running the programme, the number who genuinely knew the foundation system was very limited, and some had already left under pressure and culture friction. The delivery teams were individual external specialists working together for the first time. And the review makes the observation that follows: the learning was being gained by externals, so it would not accumulate inside the retailer.

Buy a system, and the vendor owns reliability. Out-task to an integrator, and the integrator owns velocity and staffing and carries the contractual consequence of missing. Build it yourself, and all three sit on your side of the line, where they have to be governed as programme risks rather than noticed later as delivery noise. Meanwhile the steering committee was still working out how to monitor and control a programme already inside its critical delivery phase. That is the mismatch: the choice was do-it-yourself and the governance was still the governance of something bought.

The same fork, in front of everyone, right now

Strip the sector out and the decision facing enterprises today is identical in shape. Build the retrieval-augmented generation and copilot platform in house with a small expert team. Buy a vendor stack and accept its shape. Or out-task the build and pay for outcomes. All three are legitimate, all three have been right somewhere, and in most companies the choice is being made by default rather than by decision.

What the choice sets, silently, is risk appetite. Build it yourself and you own the evaluation sets, which means you own the definition of a good answer and the discipline of refreshing it as the business changes. You own the vector store and everything behind it: ingestion, refresh cadence, the lineage that lets you say where a retrieved passage came from. You own the MLOps, the rollback, the monitoring of quality drift that nobody outside your building will report to you. You own key-person exposure, and here the expert group is smaller and more mobile than it was then. There is no vendor to escalate to when answer quality degrades, because you are the vendor.

That is not an argument for buying. It is an argument for knowing which of the three you picked, and for keeping the early autonomous experiments carefully scoped while you find out. The EU AI Act has reached political agreement rather than landing as an obligation anyone has yet had to answer to, and the documentation it points toward, data lineage and evaluation evidence and human oversight of anything running unattended, is documentation the do-it-yourself builder needs for their own sake long before a regulator asks.

~30%
New code estimated to extend the chosen foundation, an estimate the review says needs supporting analysis
50+
Issues a month in that foundation before any extension
~50%
Programme management's own estimate of planning reliability, velocity being unknown
12
Recommendations issued, none of them asking the programme to reverse the build decision

What matching governance actually looks like

Most of the twelve recommendations answer the gaps the reviewers had just listed: settle the scope, build the monitoring levers, get the architecture drawn. A few answer the setup itself, and those are the ones worth reading twice.

The reviewers asked for an independent assessment of the foundation the programme was betting on but had not built, and for measures and an alternative route to be designed in advance in case it proved not good enough. They asked for a plan against the platform not arriving on time or at all, and for the main plan to be assessed against the paths that are not the happy one. They asked for quality assurance to be organised both inside the programme and around it. And they asked the programme to organise deliberately for learning the new way of working, for knowledge building, and for retaining the key knowledge workers who held it.

That is a list written for one programme, not a checklist to lift. What travels is the shape of it: the last two items have no equivalent in a bought programme. When you buy, assurance is a contract clause and knowledge retention is the vendor's problem. When you build, both are load-bearing, and a programme that skips them has taken on the responsibility of building without becoming the kind of organisation that can. That is the question RealAI Platform work opens with once a client has decided to build: not what are you going to build, but what have you just made yourself answerable for.

Choosing to build it yourself is not the risk. The risk is choosing to build it yourself and then governing as though you had bought it, because a bought programme lets somebody else carry the things you have just taken on.

Building the platform yourself is the first build. Building the governance a self-built platform needs, the evaluation discipline, the lineage, the independent read on your own foundation, the exit you hope never to use, the retention of the few people who understand any of it, is the second, and it is the one that gets skipped. Do the first without the second and you have not chosen an approach. You have hidden its bill.

Drawn from an independent readiness check on a European energy retailer's rebuild of its business-customer platform: its account of the setup, its findings on the chosen foundation, its closing recommendations. That review produced findings and recommendations ahead of delivery, not delivered results. Reading its opening move as a lesson for how enterprises build with language models is ours.

Choosing to build it yourself is not the risk. The risk is choosing to build it yourself and then governing as though you had bought it, because a bought programme lets somebody else carry the things you have just taken on.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?