Almost all of the traffic on a careers site is anonymous. Somebody reads three job pages, opens a fourth, downloads nothing, fills in nothing and leaves. In most measurement setups that visit is counted once and then discarded, because the funnel is defined as starting at the form, and everything before the form is treated as weather.
I have been reading back through a platform blueprint I worked on for a global staffing and HR services group, and the line that stayed with me sat in an appendix nobody would have read twice. Assessing what a front-end platform could do for the group, it said that such a platform could build a profile of an unknown visitor and, by serving that visitor the most relevant thing it held, improve the odds of turning them into a known one.
Read quickly, that is a personalisation claim of the kind every vendor made at the time. Read slowly, it is a proposition about identity that most organisations still have not absorbed. It says identity is earned rather than collected. You do not ask a stranger who they are. You make it worth their while to tell you.
The mechanism, as it stood then
The version available at the time was content served against a behavioural profile. The system watched which pages a session touched, inferred something about intent from the pattern, and adjusted what it put in front of that session next. No conversation, no question anybody could ask it: it could only reshuffle what it already held, working from an inference over clicks.
Modest, and correct. The bet underneath was that a stranger who has been shown three genuinely relevant roles will eventually give you an email address, because by then it is in their interest to. That reverses the usual arrangement, in which the form arrives first and relevance is promised afterwards, and it is a better arrangement on both sides.
What changed, and what did not
The mechanism can now answer back. A retrieval-augmented assistant sitting on a job corpus lets an anonymous visitor state their own case instead of having it inferred from clicks. They can say they are a night-shift theatre nurse who will not relocate, and get an answer drawn from live vacancies rather than a reshuffled landing page. The visitor supplies the specificity that the older approach had to guess at, and the exchange itself is worth more than the click trail it replaces.
That is a real improvement, and it sits on the same axis as the original claim rather than departing from it. Relevance still earns identity. What improved is its resolution and its speed.
What has not changed at all is where the value is gated.
The gate sits behind the front end, and always has
The blueprint was honest in a way these documents usually are not: it put the blocking conditions on the same page as the capability rather than in an assumptions annex at the back.
The first was consolidation. The group ran separate sites per market, each locally shaped, and the benefit was conditional on merging them onto one platform. The second was capacity, the risk that the local development skills to do that would not be there. The third was the integration into the CV and job databases, and it is the one that matters most, because that is where a recommendation gets its content. Everything upstream of it is presentation.
Note the shape. The strength lived in the front end, where the sponsor and the budget are. The gate lived in the back end, where neither is. That shape has survived every generation of front-end technology since, which is why the capability keeps getting rebought and the conversion keeps not happening.
Why retrieval makes the deferral more expensive
An assistant grounded in retrieval has to retrieve from somewhere. That is the whole design. It is also why the consolidation argument stops being about efficiency and becomes a precondition.
Point a vector store at a set of unmerged local marketing sites and you get an assistant that answers questions about marketing sites. It will be fluent. It will be confidently wrong about which roles are open, because open roles are not in the corpus it was given. Fluency is the failure mode: the older approach failed visibly by showing an irrelevant page, and this one fails invisibly by giving a well-formed answer that no vacancy supports.
So the corpus has to be the job data itself, current, with the CV side reachable when the visitor consents to being matched. Which is the integration the blueprint flagged as difficult, still unbuilt in most groups, now sitting directly on the critical path of the thing everybody wants to pilot. The pilot is cheap. The corpus is the work.
Two further things follow, and they are ordinary engineering rather than strategy. You need an evaluation set built from real questions people ask about roles, paired with the answers your own recruiters would defend, because without it you have no way to tell a good assistant from a fluent one. And you need lineage on every answer, so any recommendation can be traced back to the vacancy record it came from. Both of those are cheap to build early and close to impossible to retrofit.
- Anonymous to known
- The state change the blueprint set as the target for the front end
- Merge the local sites
- Written beside it as the condition on the full benefit
- CV and job databases
- The integration flagged as difficult, and the one holding the value
- Data separated by market
- What country legislation asks of any shared data layer
Who owns the join
Nobody, which is the honest answer and the reason for the deferral.
The front end has a sponsor, usually in marketing, with a budget and a launch date. The CV and job databases have a custodian whose job is that they stay correct and available. The join between them has neither. It has no launch, no demo, and its benefit shows up in somebody else's number.
The blueprint anticipated this better than most, in the sense that its skills plan named the people who would have to exist for the data side to hold: a data steward, a data compliance officer, a data architect. It set out a skills landscape running to dozens of roles the group would need to grow or hire. That was a plan on paper rather than a headcount, and the roles on it that exist to make data trustworthy are reliably the first a programme drops when it needs a working front end by a date.
There is also a constraint that has hardened rather than eased. A shared data layer across markets has to separate data to satisfy country-specific legislation and data protection law, which the blueprint noted in passing and which every year since has made heavier rather than lighter. Candidate data is among the most sensitive a company holds, and the political agreement reached on the EU AI Act places recruitment tooling in its high-risk tier. Consolidating candidate data across markets is now a governance design with a data model attached, not a migration with a governance annex.
Identity was never the thing to capture. It was the thing to earn, and what earns it has always sat behind an integration somebody keeps moving to the next phase.
What I would build first
Not the assistant. One market, one job family, and the plumbing under it.
Make the vacancy data retrievable and current in one place for that one family, and accept that the other markets are out of scope for now. Write the join between an anonymous session and a candidate record explicitly, with the consent that permits it recorded as a field rather than assumed, so that the moment someone becomes known is an event you can point at. Build the evaluation set out of questions your recruiters actually field. Keep lineage from every answer back to the record behind it. Then instrument the one number the whole argument rests on, which is the rate at which an anonymous session chooses to identify itself, and compare it against what a static careers page achieves today.
That sequence is the opening block of a RealAI Platform engagement on this kind of estate, and the order is the point. Every item on it is unglamorous, none of it needs a model to begin, and all of it has to land before an assistant is anything other than a demonstration. Early autonomous behaviour, where anyone attempts it, belongs after that work and inside a narrow scope, because a system acting on a profile it inferred without consent is not a capability. It is an incident.
The blueprint had the proposition right and said plainly what it would cost. What it could not do, and what no document can do, is make the back-end work as attractive as the front-end launch. That remains the actual problem, and it is a governance problem rather than a technical one.
Drawn from an integrated platform blueprint prepared for a global staffing and HR services group, on which I worked. The capability, the conditions attached to it and the skills plan described here were recommendations for a target state: that work produced findings and a direction, not delivered results. Named platform assessments in the source are the authoring firm's view of third-party products and are deliberately not reproduced. Reading the anonymous-to-known claim as a data readiness question is ours.
“Identity was never the thing to capture. It was the thing to earn, and what earns it has always sat behind an integration somebody keeps moving to the next phase.”
Get in touch
Put RealAI’s applied-AI team on your hardest data problem.
We help enterprises move from pilots to production: sovereign models, governed data, and agents you can audit. Start with a value-first assessment.
