A handling-time target of 315 seconds is not interesting on its own. A target of 315 seconds that says where the missing 135 seconds are supposed to come from is interesting, because the second kind can be shown to be wrong.
The design we wrote for a consumer and SME banking group operating across several European markets sets that target. Average handling time is to fall from a stated baseline of 450 seconds to 315, a thirty percent reduction, and the slide carrying it does not leave the cause implied. It attributes the whole reduction to search: the agent stops opening systems to hunt for the answer, because a retrieval layer puts the relevant policy on the screen along with the document it came from.
Nothing has been run. This is a design deliverable, and every figure in it is a commitment the design makes about what it expects to produce, not a reading taken afterwards. That distinction matters more here than usual, and the deliverable happens to prove it about itself, which is the second half of this piece.
The challenge
The baseline is stated in three lines, and they are worth taking in order. Agents navigate five or more systems to answer basic queries. Cross-sell opportunities are missed because of cognitive load. Churn risk is identified by manual intuition.
The first is a time cost and the other two are opportunity costs, but they have one cause between them. The agent's attention is going into retrieval. Everything that requires attention beyond answering the question, which is to say noticing that this caller is at risk of leaving or that this caller is eligible for something better, has to compete with the search for a fee schedule.
Most contact-centre AI business cases skip past this and quote a percentage, leaving the reader to assume that the model is somehow faster than a person. It is not, and it cannot be. The caller sets the pace of the conversation, and a copilot that drafts a sentence in a fraction of a second saves nothing if the agent was never the bottleneck in speaking. The only place a copilot can take real time out of a call is the silence while the agent is looking something up. Naming search as the mechanism is not a flourish. It is the only mechanism available.
Naming it also exposes the design to a question it has not yet answered. The 450-second baseline is stated. What is not stated, anywhere, is how much of the 450 is search. If retrieval accounts for 200 seconds of an average call, a 135-second saving is aggressive but arguable. If it accounts for 90, the target is arithmetically impossible and no amount of model quality will rescue it. The design names the mechanism. It does not size it. That measurement is a week of call-recording analysis and screen-capture sampling, and it is the cheapest work in the entire programme, because it decides whether the rest of the programme is worth funding.
The approach
The workspace has three legs. Retrieval over policy documents and terms, so the answer arrives with a citation to the document that supports it. A live churn score that surfaces a retention offer during the interaction rather than after it. And drafting at the augmented rung, where the copilot pre-writes the reply or the case note and the agent approves, edits or discards it.
Three controls sit on the same slide as those legs: personal data redaction, human approval, citation. Elsewhere the design makes the wider standard explicit. Names, account identifiers and tax identifiers are masked before inference rather than after, so the worst outcome of a bad prompt is a wrong answer rather than a disclosure. Every claim the copilot makes must link to an internal policy document. Prompts, responses and agent feedback are retained in full, which the design sets as a mandatory control on every assistive deployment rather than as a nice-to-have, and which is also the raw material for everything in the next paragraph.
Citation deserves a second look, because it is doing two jobs and only one of them is safety. As a control it makes a claim traceable to a source a supervisor can read. As an instrument it is the mechanism made observable. Search is only deleted when the agent trusts the retrieved answer enough not to check it independently, and the logs say whether that is happening: how often an agent opens the cited source anyway, how often a draft is accepted rather than rewritten, how often a query returns nothing groundable and the agent falls back to the old five systems. None of those is a satisfaction score. All of them move before a satisfaction score does, and each is a direct reading on the mechanism the target depends on.
This is the point where a design becomes something we build. Choosing to ground a copilot is a short conversation. Wiring retrieval to policy sources that change without warning, keeping the citation honest when a document is superseded, and instrumenting the agent's trust in the answer as a first-class metric is a system. Agentic OS exists for that part, because a copilot that cannot tell you whether its answers were believed is a copilot you cannot improve.
The outcome
What this engagement produced is a design: a workspace specification, a control set, a target for each of three metrics, and a measurement regime to test them against. No copilot was deployed in this phase and no handling time was measured after one. The projections are that handling time reaches 315 seconds, satisfaction moves from 3.8 to 4.4, and acceptance of upgrade offers rises by eight points, from twelve percent to twenty. Every one of those is a designed target. Read as achieved results they would be a fabrication.
The deliverable makes that easy to verify, and we would rather point at it than tidy it away. The satisfaction metric appears twice against two different baselines, 3.8 on the workspace slide and 3.5 in the benefits model. The handling-time reduction appears as 25 percent on the assistive canvas and as 30 percent on the workspace slide, the second being the richer configuration with drafting included. Neither discrepancy is worth hiding, because a number that has been measured does not disagree with itself. Two versions of one figure is what a projection looks like from the outside.
Which is exactly why the mechanism has to be named. A projection that says only "handling time falls thirty percent" cannot be checked before the money is spent. A projection that says "handling time falls thirty percent because agents stop searching" can be checked in three ways before anyone builds anything: measure the search time in the current baseline, run the copilot against a human control group rather than against last quarter, and watch whether the citation is trusted.
The design already carries the apparatus for the second of those. Trials run as randomised comparisons against a human baseline rather than as before-and-after readings. Hypotheses are written before launch and each one names its constraint alongside its improvement, so that faster can never be bought with less accurate. A live trial aborts automatically if the error rate passes 0.5 percent or latency passes two seconds. Tuning decisions happen weekly at working level, scaling decisions quarterly against two criteria: value above cost, risk inside appetite. Our Consult work starts from that regime rather than from the target, because the regime is what turns a number in a deck into something an operations director can defend.
The difference between a designed target and a hopeful one is not the size of the number. It is whether the sentence after the number says why.
