Skip to content
Blog

Product

The AI match score is the wrong instrument

A score of 87 % cannot be argued with, only applied. 75 % of companies let an AI reject without review, and 26 % of candidates call that fair.

No, and the problem is not that the score is wrong: it is that it cannot be argued with. A number does not say what it measured, so a recruiter who finds it doubtful has nothing to push back with, and ends up reading the CV anyway, which cancels the very saving the tool was bought for.

The argument is usually treated as an ethical concern. It is operational first: 75 % of companies now let an AI reject an application without any person having read it, and 57 % of those same companies name screening out qualified candidates as one of their main risks. They are describing a tool they have handed a decision they cannot check.

What a score of 87 % does not tell you

It does not tell you what produced the 87. Genuinely relevant experience, a keyword well placed by a candidate who worked out the system, or a variable correlated with age or postcode: all three yield the same number, and nothing in the interface separates them.

Nor does it tell you what is missing, which is the most useful thing a recruiter could know. Learning that a profile sits at 87 teaches you nothing about what to probe in the interview, whereas learning that they have eight years on the right technology but no trace of the client’s regulatory constraint hands you your first question.

Finally, it says nothing about what the tool could not look at. A CV silent on a subject and a CV that contradicts the requirement produce similar scores, when one calls for a ten-minute phone call and the other for a decline. That confusion between absent information and negative information is the costliest weakness of the score model, and it stays invisible for as long as you look at the ranking rather than the files.

Why the score won anyway

Because it demonstrates well. A single number fits in a column, sorts, charts and presents in a meeting, whereas a paragraph of explanation demands screen space and reading time. An interface constraint decided the shape of the product, and nobody really decided that a rejection should fit into two digits.

There is a more honest reason, and it deserves acknowledging. Facing several hundred applications for one role, a growing share of them generated automatically, a recruiter needs a reading order. The need is real, and the loop in which both sides armed themselves with AI made volume unmanageable faster than tools learned to answer it.

Except that the need is for a reading order, not a verdict. The slide from one to the other happened without discussion, largely because a tool that ranks and a tool that screens out look very much alike on screen and nothing alike in court.

What ranking does better than scoring

A ranking assumes it will be read. It orders a list somebody will go through, and the person going through it can decide at the third line that the order is wrong, pull a file up, drop the first one. The score, by contrast, is built to avoid the reading, which is exactly what you ask of it the moment you wire a threshold to it.

The difference shows up in the vocabulary, which is not incidental. We write that an agent ranks and explains, never that it sorts, screens out or selects, because the verb always ends up bleeding into the product. A team that says “the tool screens out” will build a tool that screens out, and the limit will have been lost in a design meeting before it is lost in the code.

What we display next to each profile is therefore a sentence, not a number: what was found, what is missing, what could not be verified. It is heavier to read than a percentage, and it is the only form a recruiter can contradict, which is the definition of a tool you can actually work with.

Mobley v. Workday is the sharpest demonstration of what an unexplainable decision costs. On 17 February 2026, the federal court for the Northern District of California authorised notice of a collective action open to anyone over forty whose application was processed through the platform since September 2020.

The element that tipped the case is worth remembering for anyone deploying a screening tool. The plaintiffs argued that rejections arrived minutes after the application, therefore too fast for a person to have read anything. The timestamp became the evidence, without anyone needing to open the model.

The practical consequence is easy to state. The day a decision is challenged, you have to be able to show what was evaluated and who took it, and a score of 87 % answers neither question. That is why the validation line sits on whatever leaves the company, including and especially a refusal.

What candidates do with it, and why it concerns you

Only 26 % of candidates trust an AI to evaluate them fairly, across a 2026 survey of 2,950 responses, and 67 % say they are uneasy about a process led by an AI. Those numbers are usually read as an image problem, which is a misjudgement in a market where the profiles you want are the ones choosing.

Behaviour follows the distrust, and that is where it turns into a cost. A candidate who assumes a machine is reading their file writes for the machine, which degrades the quality of the information you receive, and you answer that degradation by hardening the screening tool. The spiral is well known, and nobody has ever escaped it by adding another layer.

A written explanation breaks the spiral at one precise point. A refusal that says what was missing is a refusal a candidate understands, does not come back on, and does not add to the general conviction that the process is arbitrary. It is also, incidentally, what lets you contact the same person again in six months without starting from nothing.

How to check a tool in ten minutes

Ask to see a poorly scored profile, not a strong one. A demo is always built around an excellent candidate scoring 94 %, whereas the value of a tool is judged on what it says about an average file, the one you would have hesitated to call.

Then demand the reasons behind the score, at a level of detail you can use. “Technical skills: 8/10” is not a reason, it is the same score cut into pieces, and the next question exposes it: ask what would have changed the score. A tool that can answer knows what it measures; a tool that hesitates has a model producing a plausible number without knowing where it came from.

Finally, ask about the threshold. Is there a value below which an application disappears from the screen, who set it, and where is that written down? If the threshold exists and any user can move it without leaving a trace, the tool lets anyone decide alone at what point you stop looking at people. The next question, once the tool is chosen, is what it must simply not be able to do.

Frequently asked questions

Is a match score forbidden?

No, nothing forbids it as such. The problem is the use: a score that triggers an automatic rejection is a decision taken without explanation, which article 22 of the GDPR has governed since 2018. A score displayed next to its reasons, followed by a human decision, is a legitimate sorting tool.

What is the difference between ranking and screening out?

Ranking means ordering a list somebody is going to read, from most to least relevant. Screening out means removing applications from that list before a person sees them. The first saves time, the second takes a decision, and only the second commits the company.

What does a usable explanation contain?

What was found, what is missing, and what could not be verified. “Eight years of Python, four years of distributed architecture, no trace of your messaging stack” can be refuted in one sentence by a recruiter who knows the role. “87 %” cannot.

How do you test explainability during a demo?

Ask to see a poorly scored candidate and demand the reasons behind the score. Then ask what would have changed it: a tool that cannot answer does not know what it is measuring. Those two questions take ten minutes and eliminate most products.

Sources

  1. Employer Branding News, AI in hiring statistics 2026: adoption, bias, trust and regulationemployerbranding.news
  2. HeyMilo, Explainable AI for recruiting: the match score problemheymilo.ai
  3. Wiggins Childs, Federal court authorizes notice in Mobley v. Workday, February 2026wigginschilds.com

Read next

€100 in credits when you sign up

Join the waitlist.

Leave your email address and we will let you know as soon as Balt can join your team.

Already 247 staffing firms on the waitlist