Skip to content
Hominis Agentic OS · early access program now openJoin the waitlist
RealAI

Case studiesStaffing and HR services

Case study
Staffing and HR servicesA global staffing and HR services group

Two of the three signals behind the ranking measured the employer's hiring, not the employer

A global staffing and HR services group had an unsatisfied sales hit rate and no way to predict which employers would generate placement demand. In a six-week run, business rules written by experienced sellers were weighted into a scoring algorithm that produced a five-star rating on each prospect, and the ranked calling list went live with twenty sales staff who fed their outcomes back into it. The score took three inputs. Only one of them described what a company was; the rest were readings of how its own hiring was going, and that is where the argument sits. The source records the outcome as a direction rather than a magnitude, and this piece keeps it that way.

6 weeksWeek zero to a scored calling list in live use
Client
A global staffing and HR services group
Duration
Six weeks, week zero to a scored calling list in live use
AI · RIDGE E50 N56.3ρmax 1.00
2 of 3Scoring signals that were operational rather than demographic
20Sales staff working the ranked list and feeding outcomes back
4Delivery steps between week zero and week six

A staffing branch always has more employers in its territory than it can call, so the order of the list decides the quarter. For years that order came out of firmographics: headcount band, industry code, turnover, distance from the branch. Those fields are free, they already sit in the CRM, and they have one shared weakness. They describe what a company is. They say almost nothing about what it is about to need.

A global staffing and HR services group came to us with an unsatisfied hit rate from its salesforce and a clear reading of why. The teams could not predict demand and could not separate high-potential prospects from the rest, so effort spread evenly across a list that was not evenly valuable. What came out of the six-week run that followed was not a smarter firmographic cut. It was a scoring model whose useful signals turned out to be operational, and the reason that matters has outlived the engagement.

The challenge

Company size is a stock. Placement demand is a flow, and the two come apart more often than a sales plan assumes. A large employer with stable rosters and a competent internal recruitment function can generate less agency work in a year than a smaller one whose roles turn over constantly and whose managers are short of time. An industry code tells you which sector a company sits in, not whether that sector's local labour market is tight this month, and not whether this particular employer has started losing the race for candidates.

A second problem sat underneath the first, and the method on the record is what gives it away. The score was not learned from a history of converted prospects. It was assembled out of business rules written down by the people who sell, which is what a team does when the labelled history it would rather train on is not there to train on. That is the ordinary starting position, not a failing peculiar to this group, and waiting it out means waiting through a full commercial cycle to collect a training set, which in practice is the same as never.

The approach

Four steps sat between week zero and week six, and the reusable part is the direction the first one runs in.

The first set the data hypotheses. The team worked out how the right prospects could be filtered from internal and external sources, and wrote multiple business rules that then served as the variables the algorithm consumed. The direction of travel is worth stating plainly, because it usually runs the other way in a data science project: business rules became features. The consultants' tacit ranking was made explicit, argued over, and turned into something a machine could apply consistently to every account in the territory rather than to the twenty accounts one person happened to remember.

The second step weighted those rules into an algorithm that produced a five-star rating on each prospect. Choosing a modest output was deliberate. Sellers already understood a star rating, it fitted on a list, and it did not require anyone to trust a probability they could not interpret.

The third step put the ranked calling list in front of a cohort of twenty sales staff, who used it to reach out to their contacts. Their outcomes were captured and fed back to keep tuning the weights, so the thing that went live was a loop rather than a delivery. The fourth read off what the loop had produced.

Now the finding. The score took three inputs. One was the size of the company, the familiar firmographic, and it earned its place. The other two were operational readings of the employer's own hiring: how quickly its open roles were being filled, and how many distinct kinds of role it had open at the same time. Neither is a description of the company. Both are measurements of a company under stress.

Fill speed is the more interesting of the two. An employer whose vacancies close quickly is telling you that its internal channel works, that candidates want to go there, and that an agency is a nice-to-have. An employer whose vacancies sit open is telling you the opposite, in public, before anyone in the branch has picked up a phone. The breadth signal works the same way from a different angle. A single kind of role open in volume is one hiring problem, probably owned by one manager with one budget. Several different kinds of role open at once is an organisation whose demand has outrun whatever it does internally, which is precisely the condition an agency exists to serve.

That is the transferable lesson, and it generalises past staffing. The features that predict whether an account is worth working are rarely the ones that describe the account. They are the ones that describe the account's current difficulty, and they usually live in operational data that nobody has bothered to make available to the commercial side.

The outcome

Be precise about what this produced. The scored list went live, the cohort used it, the highest-scoring companies surfaced at the top of the calling list, and the source records an increased hit rate for the salesforce and a measured increase in satisfaction among the people using the tool. Neither of those is quantified in the record, so neither is quantified here. What is hard is the clock and the scale: four steps, six weeks, a live cohort of twenty, and a feedback path that kept the weights moving after go-live.

Building the same thing now would change the engineering and leave the finding intact.

The scoring itself would still start from written business rules, because the cold-start problem has not gone anywhere and no amount of modern tooling invents a label set that does not exist. What would change is everything around the rules. The two operational signals decay fast, which makes this a data readiness question before it is a modelling question. A fill-speed reading that is a quarter old is worse than no reading, because it carries the authority of a number while describing a situation that has already resolved. That means a scheduled pipeline, freshness checks on every source, and lineage on each score, so that a seller who asks why an account jumped can be shown the three inputs that moved and when they were captured. That is ordinary MLOps discipline, and it decides whether a scoring model survives its first year in the field.

The interface would change too. A ranked list was the right answer for a cohort of twenty. A copilot sitting in the CRM is a better one, drafting the approach for a high-scoring account and pulling the relevant history through retrieval-augmented generation over the group's own notes and placement records, with a vector store standing behind it. That only works with an evaluation set built from the outcomes the original loop was already capturing, which is the same discipline in a new place: the loop was never optional.

We would also keep a person on the call. There is room for early and carefully scoped autonomous experiments here, and the natural first one is refreshing scores and flagging accounts whose fill speed has moved sharply, which is bounded work with a reversible failure mode. Deciding who gets contacted, and how, stays with the seller. That boundary is where our Platform work draws it, and it is the boundary Consult engagements spend most of their time arguing over, because it decides whether anyone in the branch trusts the number.

The finding has held up longer than the technology that carried it. Ranking accounts by what they look like is a habit inherited from the days when firmographics were the only thing anyone could buy. Ranking them by what they are visibly struggling with is available to any team willing to plumb one operational feed into its CRM. Nobody here ran the two rankings side by side, so this piece cannot tell you the size of the gap. What the record does show is that the operational signals were the ones worth building on, and that a seller could act on the result inside six weeks.

NEXT STEP

Ready to make AI real?