Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI
InsightsDelivery Assurance

Pin Every Finding to a Ticket

RealAINov 22, 20238 min read
Delivery AssuranceSoftware QualityGovernanceProcess MiningData Lineage

A review is judged on its findings. A finding is judged on whether anyone can act on it after the reviewer has packed up and gone.

Those two sentences look compatible and they are not, because on the page a finding and an opinion are indistinguishable. Both are a short declarative statement about something being wrong. One of them has an identifier behind it pointing into the system the engineering team opens every morning. The other has a reviewer's judgment behind it, and a judgment travels only as far as the person who made it.

I have both in front of me, in the same report, written in the same week by the same team. It is a review of an outsourced application build at a European banking group's IT services subsidiary, and the technical section I owned states its own limit plainly: a high-level assessment of key aspects rather than a deep-dive audit of the code. Within that limit, the design pages did something the rest of the document did not. They named a property, then pointed at the ticket that evidenced the breach.

What the identifier actually does

An identifier does four things a sentence cannot, and none of them is about rigour for its own sake.

It converts a claim into a task. A finding that reads "coupling between components is too tight" needs someone to decide what to do about it, and that decision starts from zero every time it is discussed. A finding that reads "coupling, and these two items in the tracker" already exists in the system where work is prioritised. It has a state. Somebody can close it, and closing it means something.

It survives the reviewer. This is the part that matters most and gets argued about least. An external review has an end date. Every finding that lives only in the report's prose depends, from that date onward, on someone remembering what the reviewer meant and caring enough to defend it in a meeting where the person who wrote it is not present. Findings that resolve to tracker items do not need a defender. They need a sprint.

It exposes weight. Duplication attracting three separate references is a different statement from duplication appearing once in a reviewer's list of concerns. The count is not a score and it is not evidence of severity, but it is evidence of recurrence, and recurrence is the thing an organisation can least easily argue away.

It makes the finding falsifiable. The supplier can open the item. So can the client. Both can read the same history and disagree about the fix rather than about whether the problem happened. Compare that with an assertion carrying a private note asking whether it should be verified: the assertion cannot settle the question about itself, so the question gets settled by whoever speaks with most authority in the room.

Why the empty column is the normal case

The sixteen blank pages are not carelessness. They are what happens when the evidence rule is written as an aspiration rather than as a precondition.

Whoever built that report knew exactly what good looked like. Every finding page carries the identifier column. A drafting note further into the document asks, in plain words, for more structure and for evidence shown as artefact read and check done. The standard was understood, printed on every page, and executed on two of eighteen.

The reason is cost, and the cost lands at the worst possible moment. Writing the sentence is quick. Finding the identifier that proves it is not, because it means going back into the tracker, searching a system whose vocabulary you picked up late and under pressure, deciding which of several plausible items is the one you meant, and accepting that you might not find anything, in which case you have to either drop the finding or admit it is an opinion. That work arrives when the deadline is closest and the reviewer is most confident they are right. Confidence is exactly what makes the evidence feel redundant.

So the identifier gets skipped, the finding still reads well, and the report ships. Six months later the sentence with the number behind it is a closed ticket and the sentence without one is a disagreement.

The same discipline decides whether your data work is real

This is not only a software assurance point, and the reason I am writing it down now is that the discipline is about to be tested somewhere harder.

Every organisation I work with is standing up machine learning pipelines, and every one of those pipelines generates findings continuously. A feature drifted. A source system changed a field's meaning. A training set contains rows that could not have existed at prediction time. A model's error concentrates in one segment. Each of those is a finding in exactly the sense the review used, and each has the same two possible fates.

If it resolves to an identifier in the tracker the engineering team works in, it is a task, and it is subject to the same prioritisation as everything else the team owes. If it lives in a data science notebook, a monitoring dashboard nobody owns, or a slide in a monthly steering pack, it is an opinion held by whoever built the model, and it expires when that person moves teams. The talent market being what it is, that is not a hypothetical horizon.

Lineage is the same argument made about values rather than about defects. Recording where a number came from is only useful because it lets a later disagreement resolve to a record instead of to a memory. Process mining is that argument again, made about the workflow: reconstructing what actually happened from the traces the systems left, rather than from what the people involved recall having intended. Assembled after the fact by interview, a delivery chain of evidence costs a multi-week engagement. Assembled continuously from the tracker, the build record and the deployment log, it costs a query.

That is the plumbing the RealAI Platform work puts in first, before anything is modelled, because a finding that cannot resolve to an item is not an input to anything downstream.

12
Tracker identifiers cited as evidence in the design findings
7
Design areas carrying at least one identifier
3
Corroborating identifiers on duplication alone, the highest of any area examined
16
Later finding pages with the identifier column present and empty

The rule I would write into the terms of reference

Reviews are commissioned to settle arguments. The way to make one settle an argument for longer than its own duration is to require, as a condition of a finding being written at all, that it resolve to something in a system the receiving organisation will still be running next year.

A finding with no identifier is not thrown away. It is relabelled. It moves out of the findings section and into an observations section, where it is honestly described as what it is: the considered judgment of an experienced person who looked at the thing. That is worth having. It is not worth confusing with the other kind, and the confusion is what makes both weaker, because a reader who cannot tell them apart discounts all of them at the same rate.

The report I have described did that split without meaning to. The design work chose evidence, the vendor-performance work chose judgment, and the difference between them is visible now precisely because nobody hid it.

A finding with an identifier behind it is a task. A finding without one is an opinion, and an opinion expires the day the reviewer stops coming to the building.

Three things to check on your next review

Ask, before the review starts, which system the findings will point into, and confirm that the reviewers have access to it rather than to a monthly export of it. A reviewer who cannot search the tracker cannot cite it, and will write opinions instead, competently.

Ask, at the halfway mark rather than at the end, what proportion of the findings written so far carry a reference. The number is easy to produce and it forecasts the value of the whole exercise. If it is low at the midpoint it will be lower at the deadline, and that is the moment to renegotiate scope rather than to accept a thinner document.

Ask, when the report lands, what each finding turned into. Not whether it was accepted. Findings that became tracker items with owners and dates are the deliverable. Everything else was a conversation, and conversations are worth what the people in them remember.

Drawn from an independent review of an outsourced application build commissioned by a European banking group's IT services subsidiary, in which I owned the technical and code-review workstream. The counts described are of the identifiers and empty columns in that report as drafted. That engagement produced findings and recommendations, not a remediation, and the technical section describes itself as a high-level assessment rather than a deep-dive code audit. Reading the difference between its two halves as a rule about evidence is ours.

A finding with an identifier behind it is a task. A finding without one is an opinion, and an opinion expires the day the reviewer stops coming to the building.

Get in touch

Put RealAI’s applied-AI team on your hardest data problem.

We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.

Next step

Ready to make AI real?