Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

Pricing Quality Per Person

RealAIFeb 4, 20248 min read
Delivery AssuranceSourcingSoftware QualityVendor GovernanceBanking

Sourcing negotiations argue about two numbers, the rate and the total. Neither says anything about what the money is attached to, and that omission is where most delivery disputes start.

I was reminded of this reading back over a review I led of a failed offshore software build at a European banking group's IT services subsidiary, which had contracted a customer-facing application to an offshore development supplier. Code quality problems surfaced within months. There were escalations, a grace period, a terminated statement of work, and the build taken back in-house. The review then answered a set of client questions across three lenses: whether the delivery approach worked, whether the project was managed properly, and whether the software met what had been asked for. It produced findings and recommendations. It did not produce a rescued project.

One line in the appendix has stayed with me longer than any finding in the body.

A complaint with nowhere to land

Read the three lenses together and the shape of the failure is clear. The subsidiary was concerned that people who were not experienced enough were being assigned to its work. The review records that concern next to one about transparency, and notes it hardened as backlog accumulated sprint after sprint and trust drained out of the relationship.

Now look at what the client could do about it. It could escalate, which it did, twice. It could invoke a penalty clause, which the review flags as effectively the only control lever available, since supplier accountability was limited to contracted scope. It could ask for a senior architect and wait months. It could accept more bodies and get less coherence.

Not one of those instruments can say what the client wanted to say, which is: these people are fine, those people should not be on this account, and I would like the money to reflect that. A project-level penalty cannot say it. A milestone payment cannot say it. A ramp says the opposite, because it prices headcount when the argument is about capability.

The complaint reached down to who was on the team. Every lever in the commercial arrangement pulled at the level of the project. So a concern the size of one person had only instruments the size of a milestone.

The option in the appendix

The review set out six ways of paying a software supplier as alternatives to how this work had been bought. Pay on delivery, against confirmed acceptance criteria or milestones. Pay on business outcome, against an improvement in a business target or a share of the value delivered. Pay with fees at risk against agreed results. Pay nothing yet, taking early work at low or no cost against a commitment to later phases. Pay a retainer for a service, a fixed fee for sustained support.

And the sixth, paying on quality and satisfaction, set out as a variant of the outcome model. Some agreed share of the fee is paid up front; the rest turns on hitting a quality threshold and on how satisfied the client is with the work. Then the line that matters: quality can be defined at project level or at individual team member level, whichever fits the engagement.

I want to be exact about the status of that. It appeared in an appendix, marked illustrative, as one of six options for future work. It was not a clause in the arrangement under review. Nobody agreed it, priced it or ran it. The review did not claim it would have saved the build, and neither do I. It named the instrument whose absence explains why the loudest complaint in the engagement had no commercial expression.

That generalises, which is worth more than a saved project. Whenever the thing you are unhappy about is smaller than the thing you are paying for, you have no lever, and you find that out under stress.

Selection and payment are two halves of one control

Per-person quality pricing does not work alone, and the review paired it with the other half without ever putting the two on the same slide.

The staffing recommendation told the client to run its own intake rather than take the supplier's bench on trust. Three stages, based on critical success factors agreed with the supplier up front. A short screening interview of the first pool of developers, fifteen to twenty minutes, a straight yes or no. Then deeper competency testing of selected project managers, architects and developers, oral or written, across basic, moderate and advanced skill levels. Then monitoring through onboarding, execution and churn.

Selection without a payment consequence is advisory. Payment consequence without selection is arbitrary. Together they make a control: you had a say in who arrived, so the assessment you make later carries weight, and the supplier knows in advance which assessments cost money. Neither half was in place here. What the client had instead was escalation, and one request for a senior architect sat for two to three months before it was acted on.

What the model needs before it works

The review is equally clear about preconditions, and they are unforgiving. Paying on business outcome rests on three things: how much control the supplier genuinely has over the outcome, whether a baseline exists to measure against, and whether both sides have agreed a scorecard. Putting fees at risk, the review notes, presupposes an established and mature client and supplier relationship. It also names the sizing problem plainly, that a larger pie means bigger risk while a smaller pie means less motivation.

Which produces the trap. The mechanism that would align incentives requires the trust it is supposed to build. A failing build is where you most want outcome-linked payment and least qualify for it. The worked example in the same appendix has a supplier placing half its fees at risk against products that work and benefits that land, with total rewards capped at 120 percent of the day rate or fixed price. Those are illustrative figures on an illustrative construct. No party here was anywhere near able to sign them by the time anyone wanted to.

The practical reading is that per-person quality pricing is a term you write at the start, when relations are good and nobody thinks they will need it, or you do not get it at all.

6
Ways of paying a software supplier set out as options in the review appendix
3
Stated preconditions for paying on business outcome: supplier control, a baseline, an agreed scorecard
15 to 20 min
Recommended first-stage yes or no screen per candidate developer
2 to 3 months
Delay before a request for a senior architect was acted on

The question this now raises

Here is what makes the appendix line current rather than historical. Paying on quality at individual team member level assumes the team is made of individuals. That assumption is quietly coming apart.

Coding copilots are already in enterprise development teams, and retrieval-augmented generation over an internal codebase is a straightforward thing to stand up on a vector store. Suppliers are experimenting, narrowly and carefully, with automating the mechanical work: test scaffolding, migrations. So when a supplier developer submits a change, the question of whose quality is being priced stops having an obvious answer. Assess the person and you are partly assessing a tool that every developer on the account shares. Assess the tool and you have lost the accountability that made the per-person definition useful.

My own view is that the unit has to move from the person to the reviewed change, while accountability stays with the person who submitted it. The same appendix that carries the payment options carries peer review guidance that supports this: keep a review to a few hundred lines, because beyond roughly three to four hundred the defect density drops and the yield stagnates somewhere in the seventy to ninety percent range, and keep the inspection rate under five to six hundred lines an hour. Those thresholds describe the limits of human attention. They did not change when generation got cheap. What changed is that the constraint is now entirely on the review side, which makes review capacity the scarce thing and therefore the thing worth pricing.

None of that is possible without lineage you can actually query: which change, which author, which review, which defect, traced end to end. Most estates cannot answer that today, which is the same data readiness problem that stops so much MLOps work, arriving in commercial clothing. Building that trace is the first block of a RealAI Platform engagement, and we put it before anyone reopens the commercial terms. The clause is easy once the evidence exists, and unenforceable before then.

Findings and figures are as recorded in an independent review of a terminated offshore software development engagement at a European banking group's IT services subsidiary. The pricing models discussed appeared there as illustrative commercial options for future work, not as terms in force. Reading the per-individual quality option as the lever the engagement lacked is ours.

The complaint reached down to who was on the team. Every lever in the commercial arrangement pulled at the level of the project. So a concern the size of one person had only instruments the size of a milestone.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?