Skip to content
Blog

Craft

Before evaluating an AI agent, look at your data

A candidate database loses about 30% of its accuracy a year. An agent reasons soundly over that wrong data, confidently, which costs more than no agent.

Yes, but not all of it, and certainly not beforehand. Clean the scope the agent will touch first, at the moment it touches it. The great preparatory clean-up is a near guaranteed way never to start.

The reason for that is purely arithmetic. A contact database loses roughly 30% of its accuracy per year, and 2.1% of records go invalid every month. On ten thousand contacts, that is close to three thousand stale records in year one, then a further two thousand one hundred the year after: on top of what was already stale. Nobody cleans a database faster than that.

Why stale data costs more with an agent

A person opening an old record can see it. The date of last contact, a client company that no longer exists, an “immediately available” entered eight months ago: they correct mentally and grow cautious.

An agent, by contrast, is not cautious at all. It reads availability as availability, calculates, and hands you an argued, structured, instantly credible shortlist. Nothing in the output signals that three of the five profiles have changed jobs.

That is an uncomfortable inversion: the better the agent, the more bad data costs. A mediocre tool produces results people check out of habit; a tool that is right nine times out of ten builds the trust that lets the tenth through unchecked.

It is also why misunderstood integration is one of the three causes of agent project abandonment. The agent is plugged in, it reasons perfectly over wrong data, the team concludes “the AI gets it wrong”, and nobody looks at the ATS.

The four fields that decide everything

In a recruitment database, value is not spread evenly across fields, and four of them are enough to determine whether an agent will be useful at all.

Availability. The most perishable field and the most decisive one. Availability with no entry date is worth nothing: it is not information at that point, it is somebody’s memory of information. If your ATS does not timestamp it, that is the first setting to change, before anything else.

Contact details. A work email address degrades by 20 to 30% a year; an email verified today has roughly a 70% chance of still being valid in a year. Direct phone numbers and job titles follow, at 15–25%.

Current employer. The field whose staleness is most commercially embarrassing: approaching somebody at a client they have left, or worse, at one they have just joined.

Date of last contact. The only field that tells you what the other three are worth. A profile you know you spoke to three weeks ago is usable even when incomplete; a perfect profile with no date is not.

A profile with those four fields correct is usable. The rest, detailed skills, projects, certifications, is comfort, and it is precisely what an agent can reconstruct from a CV.

The ten-record test

It takes twenty minutes and it beats any audit you could commission.

Pull ten profiles entirely at random from your database, genuinely at random, not the last ten. Check the four fields by hand: a call, a glance at LinkedIn, an email.

Seven out of ten correct and an agent will be useful from day one. Between five and seven, keep it to recent profiles before widening. Below five out of ten, do not deploy yet: you will produce credible and wrong results, and burn the team’s trust over a problem that is not the agent’s.

Run the same test again three months later. It also measures what using the agent does to your data, which is the real subject.

The special case of scattered data

The most common problem is not wrong data: it is correct data that exists and the agent cannot reach.

An availability mentioned in a Teams thread. An interview write-up in somebody’s personal inbox. A client need in a shared file three people maintain and five consult. A skills file in the Drive of a business manager who is on holiday.

That data is perfectly fresh and, from the tool’s point of view, perfectly invisible. The agent will conclude a consultant has no recent write-up, which is false, and propose an interview that already happened.

Two simple reflexes beat a migration project here. Connect the agent where the data actually lives, inbox included, rather than pretending everything is in the ATS. And let it write into the system of record: an agent that files into the ATS what it finds in an inbox absorbs the scattering through use, with no project and no meeting.

That is also a good reason to look at what it connects to before looking at what it can do.

The cleaning order that works

By what the agent touches, in the order it touches it.

Consultants on assignment and their end dates. The smallest scope, the most current, and the one with the most immediate value: anticipating a roll-off by three weeks is worth more than ten hours of sourcing. It is also the fourth week of delegation, so you have three weeks to get to it.

Open client needs. A dozen lines, usually kept current because somebody downstream depends on them. Mostly check that they are in the tool and not in an inbox.

Profiles touched in the last six months. Those profiles are the genuinely active part of the pool. The rest is a dormant asset: it will wake when a search surfaces it, and that is the moment to verify it, not before.

What never pays is cleaning 2022 profiles “to be tidy”. They will decay again before they are useful, and meanwhile the agent is not deployed.

What the agent fixes on its own, and what it does not

It does fix incompleteness, and rather well. A CV exists and the record is empty: it fills it. A job title is wrong and LinkedIn says otherwise: it flags it. An interview write-up sits in an inbox: it files it.

That is already considerable, and it is a virtuous circle: the more the agent works, the fresher the data it touches, because it writes down what it learns.

What it does not fix is absence. If nobody ever noted that a consultant refuses assignments more than an hour from home, no agent will know. And it does not fix data that is wrong but plausible: “immediately available” from last year looks exactly like “immediately available” from today.

Hence the only rule that matters: timestamp everything that perishes. Data with its entry date is never entirely wrong: it is dated, and an agent knows what to do with dated information. Without a date it is indistinguishable from fresh information, and that is what makes it dangerous.

Who owns the tidying

Nobody owns it in most firms, which is precisely why it never happens.

Data quality falls between a business manager who is judged on placements, a recruiter judged on interviews, and an operations lead judged on invoicing. Each depends on the database and none is measured on it. The predictable outcome is that everyone complains about it and nobody spends an afternoon on it.

An agent changes that balance in a useful way, provided you notice it. Because it writes down what it learns, the person who delegates most is also the one improving the data most, without meaning to, and without a project. Tidying stops being a chore and becomes a side effect.

That only holds if somebody watches the trend. Run the ten-record test each quarter and put the number where the team can see it. Not to police anyone: to make visible that the base is improving, which is the only thing that keeps people entering dates.

What it changes about evaluating a vendor

Two questions are worth asking before signing.

What does the tool do with data whose date it does not know? The right answer is that it says so. An agent presenting a shortlist marked “availability unverified for 8 months” is infinitely more useful than one presenting the same list with no note.

Does it write down what it learns? An agent that reads your data without ever enriching it leaves you with the same database six months later. An agent that updates the record after every interview transforms your base instead of merely consuming it, and that is the difference between a search tool and a colleague.

Frequently asked questions

Should we clean the whole database first?

No, and that project almost always fails. A database of ten thousand contacts accumulates around three thousand stale records in year one: you clean more slowly than it decays. Clean the scope the agent touches first, consultants on assignment and open client needs, and let the rest age until it is needed.

Which fields actually matter?

Four: availability, contact details, current employer, and the date of last contact. A profile with those four correct is usable even if the rest is thin; a richly documented profile whose availability dates from last year is not.

Can the agent clean the data itself?

It can flag staleness and cross-check some of it against external sources, which is already a lot. It cannot know a consultant changed jobs if nobody, anywhere, wrote it down, and it will not guess.

How do we know if our data is good enough?

Take ten records at random and check the four fields by hand. If seven out of ten are right, an agent will be useful from day one. Below five, it will produce credible and wrong results, which is the worst case.

Sources

  1. Keepsync, CRM data decay statistics 2026keepsync.io
  2. ZoomInfo, B2B data decay: rates, costs and how to stop itpipeline.zoominfo.com

Read next

€100 in credits when you sign up

Join the waitlist.

Leave your email address and we will let you know as soon as Balt can join your team.

Already 247 staffing firms on the waitlist