Skip to content
Blog

Delegation

How to test an AI agent before you sign

With thirty cases from your own history, whose right answer you already know, replayed three times each. A demo proves none of that.

With thirty cases drawn from your own history, whose right answer you already know, replayed three times each. It is the only method that tells you something about your work rather than about a model’s general capability, and it takes about a day to prepare.

The stakes justify the day. Quality remains the first obstacle to deploying an agent, cited by around a third of the organisations surveyed, well ahead of latency at 20 %, even though 57 % of them already have agents in production. Most have deployed, in other words, and most are not satisfied.

Why a demo proves nothing

A demo is a case chosen by the seller, on data they know, following a sequence they have played a hundred times. It establishes that a scenario is possible, which was never the question: the question is what happens to the shaky file on a Tuesday afternoon, the one with incomplete information and an impatient client.

The gap shows up in the industry’s own numbers, and it is wider than you would guess. 89 % of teams have put something in place to watch their agents in production, but only 52.4 % run evaluations against a set of cases, and 37.3 % assess behaviour continuously. Most are therefore watching their agent work without having defined what good work would be.

Public leaderboards do not fill that hole, whatever the brochures say. They saturate, teams optimise against them, and above all they contain neither your clients, nor your internal rules, nor the way your data is actually entered. They serve to eliminate a product that is visibly behind, never to choose one.

The thirty cases, and how you pick them

Take them from the last six months, from your own tools, with the data as it stood. A case built for the occasion is a clean case, and your data is not clean: that is exactly why an excellent agent returns confident wrong answers on a record nobody has touched in two years.

The distribution matters more than the count. Allow half for ordinary cases, the ones representing the bulk of the volume and where the gain will come from; a quarter for edge cases, where the right answer is to ask rather than decide; and a quarter for trapped cases, where the necessary information is missing or contradictory.

It is the last quarter that separates products, and it is the one nobody prepares. An agent that invents a plausible answer on an incomplete file is more dangerous than a mediocre agent, because nothing in the answer signals that it was fabricated. On those cases, the good mark goes to the one that refuses and says what it is missing.

Write the right answer down before you run the first trial. The rule looks obvious and is almost always broken: the moment you have seen the agent’s output, you adjust your idea of the expected result without noticing, and the evaluation stops measuring anything.

Replaying three times, and what the spread measures

Put every case three times, in separate sessions, and compare the three outputs. An agent that answers three different things to the same question is unusable, even if all three answers are defensible, because you will never be able to tell your team what it does.

That variance is the most predictive indicator and the least examined. It tells you whether the product was built with guardrails or whether it lets the model improvise, and it correlates directly with the unpleasant surprises of the third month. Two products with the same accuracy rate but opposite variances are not comparable.

Record the path as well as the result. An agent that arrives at the right place after eleven tool calls, seven of them pointless, will cost you money, for the mechanical reason that every step rereads everything before it, and it will fail on the first slightly different case because it had not understood the question.

What to record, beyond the right answer

Behaviour under doubt. On your trapped cases, does the agent ask or does it decide? A reasoned refusal beats an answer that happens to be right, and it is the only way to know what will happen on a file you had not anticipated.

What it does without authorisation. Check what leaves for the outside world during the trial, message by message. An agent that emails a candidate because an instruction was ambiguous has just shown you where its validation line runs, and that line belongs in the product rather than in an instruction.

Cost per finished task. Ask for it, write it down, compare it between products across the same thirty cases. It is the only honest unit of comparison, and many vendors have never calculated it for themselves.

The calendar of a trial worth running

Week one. Assemble the thirty cases and write their right answers, wire up read-only access, and ask the agent for nothing else. This week also measures how long a vendor takes to open an access for you, which says a good deal about what follows.

Weeks two and three. Real use, on a deliberately narrow scope and on the work whose result you can check in ten seconds. Two people are enough, provided they genuinely work with the agent rather than trying it between meetings.

Pick the two people carefully, because the trial measures them as much as the product. The right pair is one person who does the work every day and one who is mildly sceptical, not the two most enthusiastic volunteers: an agent evaluated only by people who want it to succeed comes out of every trial well, and comes out of the sixth month badly.

Week four. Replay the thirty cases, compare against the answers written in week one, and look at what has moved. A product that learns from your context should be better in week four than in week one on the same cases; if it is identical, it is learning nothing from you.

Do not conclude before the end of the fourth week. A delegation takes three to four weeks to settle, and a trial stopped after five days measures novelty, in one direction or the other, without measuring the product at all.

What a vendor should accept without argument

Your thirty cases, unseen before execution. A vendor who asks to prepare their data or adjust their instructions in the light of the cases is not selling you a product, they are selling you an integration engagement, which is another matter and another price.

Cost per task, given plainly. A figure and a spread, not a speech about model pricing.

A clean exit at the end of the trial, with your data recoverable in a readable format. It is the best test of good faith there is, it costs an honest vendor nothing, and it exposes the others immediately.

Those three demands eliminate a lot of people, and that is their function. Gartner expects over 40 % of agentic AI projects to be abandoned before the end of 2027, and the causes are known and avoidable: a ROI nobody managed to measure, a governance nobody set, an integration everyone underestimated. A four-week trial run seriously handles all three before signature, which remains the cheapest place to handle them.

Frequently asked questions

How many cases make a trial meaningful?

Thirty is enough to separate two products, provided they come from your own history and you know the right answer for each. Beyond fifty, the cost of grading exceeds the information gained. Below twenty, a single unusual case distorts the ranking.

Can you rely on public agent leaderboards?

They tell you a model is capable, not that a product will do your work. Public evaluations saturate, the teams optimising against them know them, and none of them contains your clients, your rules and your data. They serve to rule out, never to choose.

What do you measure besides accuracy?

The path taken, the behaviour under doubt, and the cost per task. An agent that reaches the right answer after eleven pointless tool calls is an agent that will cost you money and fail on the first slightly different case.

How long should a trial run?

Four weeks, structured. One week to assemble the cases and wire up access, two weeks of real use on a narrow scope, one week to replay the cases and compare. A shorter trial measures novelty rather than the product.

Sources

  1. LangChain, State of agent engineering 2026 (1,340 responses, November-December 2025)langchain.com
  2. Automation Anywhere, AI agent benchmarks: the 2026 enterprise evaluation guideautomationanywhere.com
  3. Kili Technology, AI benchmarks 2026: top evaluations and their limitskili-technology.com

Read next

€100 in credits when you sign up

Join the waitlist.

Leave your email address and we will let you know as soon as Balt can join your team.

Already 247 staffing firms on the waitlist