Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesStaffing and HR services

Case study
Staffing and HR servicesA global staffing and HR services group

The business rules were written down first, and then they became the variables the model scored on

A global staffing and HR services group could not predict where placement demand would appear, and could not tell its sales staff which prospects were worth the call. The six-week run that answered it did four things between week zero and week six: turned the sales team's own business rules into data hypotheses over internal and external sources, weighted those rules into an algorithm that emitted a five-star prospect rating, put the ranked calling list in front of twenty sales staff, and captured what happened on those calls to keep tuning the score. Highest-scoring companies rose to the top of the list, the hit rate rose, and satisfaction among the people using it was measured and rose too. The record carries directions, not percentages, and none are invented here.

6 weeksWeek zero to a scoring model in live use
Client
A global staffing and HR services group
Duration
Six weeks, week zero to live use
AI · RIDGE E50 N87.5ρmax 1.00
20Sales staff working the ranked calling list
5-starRating scale the sellers actually consumed
3Inputs the prospect score was built on

A ranking model is worth precisely as much as the ranking someone acts on. Everything upstream of that, the feature engineering, the validation curve, the argument about which algorithm family to use, is a cost until a person changes what they do next because of what the model said. Most scoring projects never reach that point. They reach a notebook.

A global staffing and HR services group had a problem it could state in one line: the hit rate its sales staff achieved was not good enough. The people making the calls could not predict where placement demand was going to appear, and could not tell which prospects deserved the effort. Effectiveness and efficiency both suffered from the same root cause. The default response to that in a large organisation is a demand-forecasting programme measured in quarters, with a data warehouse workstream attached.

What ran instead took six weeks, week zero to week six, and ended with a scoring model in live use.

The challenge

The engagement sat inside a two-track transformation. One track ran top down and dealt with platform, operating model and the multi-country execution question. The other ran bottom up, and its success criterion was written down in advance and was unusually hard: a shipped application, not a study. That constraint did more work than any technology choice made afterwards.

Bottom-up started where these things always start, with an unfiltered wish list. Everything the business wants goes on it, on the understanding that going on the list is not the same as getting funded. The list was then narrowed, in stages, to a deliberately small number of high-impact cases that actually get built. The narrowing instrument was a plain two-axis grid, value of the opportunity against ease of implementation, three bands on each. The offer around it promised the run from an initial ambition to a written playbook in as little as six to eight weeks.

Refusing to fund the whole wish list is the part clients find hardest and the part that makes the clock achievable. A short calendar is not a delivery style. It is a scoping discipline enforced by arithmetic: work that cannot show something in six weeks does not enter the funnel, which forces every candidate to be stated as a decision somebody makes repeatedly rather than as a capability somebody wants.

Prospect prioritisation qualified because it is exactly that shape. A seller decides who to call next, dozens of times a week, and today decides it from memory and habit.

The approach

Four things happened between the two markers on the calendar.

The first was defining data hypotheses. The team asked how the right prospects could be filtered out of internal and external sources, and wrote down the answer as business rules, in the language the sales organisation already used to argue about accounts. The direction of travel is the whole lesson. The business rules came first and then became the variables the algorithm scored on, rather than a feature set being derived from whatever columns the warehouse happened to hold and explained back to the business afterwards.

The second was weighting. Those rules were given weights inside an algorithm, and the output was a five-star rating on a prospect. Not a probability, not a decile, not a score to three decimal places. A rating on a scale a salesperson can read at a glance and either accept or override.

Three signals sat behind it. How big the company is. How fast that company fills its vacancies. How many different types of vacancy it has open. The first is a generic firmographic and would appear on anybody's list. The second and third are the interesting ones, and they are specific to the trade: fill speed is a proxy for how much unmet demand is sitting there, and breadth of open vacancy types is a proxy for how hard the company finds it to serve that demand internally. Both came out of the heads of people who had sold into this market for years, which is why the business-rules-first sequence was not a concession to stakeholders but the actual data-acquisition method.

The third was deployment to a named cohort. The ranked calling list went to twenty sales staff, who used it while reaching out to their contacts. Twenty is a small number, and the size is worth dwelling on. It is large enough to produce signal and small enough that every user is reachable by name when something goes wrong, which in the first fortnight of any scoring deployment is the only support model that works.

The fourth was capture. Outcomes from those calls were recorded and fed back to keep tuning the algorithm. That closes the loop, and it is the step most often designed out of a pilot because it requires the application to write as well as read.

The outcome

Companies with the highest scores rose to the top of the calling list. The hit rate of the sales force went up, and satisfaction among the people using the application went up as well, recorded as a significant increase and measured rather than assumed. Neither outcome carries a percentage anywhere in the record, so neither gets one here. What can be said with confidence is the direction and the mechanism, and the mechanism is that better-ordered work is less demoralising work.

What was put forward next was a capability review, the instrument meant to carry the run from one operating unit to many. It was designed as a self-assessment each unit would fill in and then see set beside its peers, with the result convened rather than issued. That is a design intent and not a delivered result: the record shows the instrument being introduced, not a single unit completing it. The intent is still worth naming, because it decides whether a second deployment is volunteered or imposed, and a scaling instrument that reads as an audit finding gets the second answer.

Read the sequence again with today's vocabulary and almost nothing about it needs replacing. Business rules become features. Ship to a named cohort inside six weeks. Capture outcomes and re-optimise. What has changed is what sits at the end of the ranked list. A five-star rating tells a seller who to call. The same pipeline, with retrieval-augmented generation over vacancy text and account history, can draft what to say on the call, and a copilot of that kind is now a reasonable thing to scope, provided it is scoped narrowly and evaluated against a real evaluation set rather than demonstrated.

That proviso is doing a lot of work. Retrieval only helps if the account history it retrieves has lineage anyone can check, and a generated call opener is far harder to score than a five-star rating, because the outcome arrives days later and through a human. The early autonomous experiments worth funding are the ones where the loop still closes inside a week. Our Consult engagements start by finding out whether it does, and the Platform work that follows is mostly plumbing the capture step, because a score with no outcomes flowing back is a model that will be right once and then quietly drift.

The transferable asset from this engagement was never the algorithm. It was proof, inside one calendar quarter, that the organisation could take a decision its people make daily, encode the rules they already argue about, put the result in twenty pairs of hands and learn from what came back. Everything else the programme later attempted rested on that having happened first.

NEXT STEP

Ready to make AI real?