An eval for an AI agent is made of three things: a set of cases, a way to run the agent on each of them, and graders that turn every run into a result. The cases and the graders are where most of the work goes, and where most evals go wrong, so this article is about those two.
The reason for building one at all is the one I gave in Without evals, every change is a guess: if the same cases aren't graded the same way on every version, you can't tell whether a change to the prompt, the tools or the model made the agent better or worse. Each case describes a task, the agent runs it, and the run is graded. Here I look at what goes into the first and the last of those steps.
The agent
The example is a customer-support agent for an online shop. It reads a ticket, such as "my order arrived broken, I want my money back", and has four tools: lookup_order, check_refund_policy, issue_refund and reply_to_customer. The policy gives a full refund within 30 days of delivery, store credit after that, and sends anything above €200 to a person.
Its typical failures are the ones the eval has to catch: it refunds outside the policy, asks for information already in the ticket, calls issue_refund twice for the same order, or replies politely without doing anything.
Where cases come from
A case is one ticket, the state of the shop before the agent starts, and what a successful run achieves. It describes the outcome, not the steps to get there:
1id: refund-over-limit-escalates
2ticket: "Order #2210 never arrived. Please refund the full €340."
3state:
4 orders:
5 - id: 2210
6 total: 340.00
7 status: shipped 9 days ago
8expect:
9 refund: none
10 escalated: true
11 reply: tells the customer a person will handle the request
The first cases I would write come from real failures. Every transcript where the agent did the wrong thing, every ticket a person had to fix, and every bug report becomes a case with the outcome that should have happened. Twenty to fifty of them are a good start, and they are worth more than any number of invented ones, because they describe how this agent actually fails.
Then come the boundaries of the rules. A refund on day 30 and one on day 31, an order of €200 and one of €200.01, a ticket for an order that was already refunded. These are the cases where a small change to the prompt most easily changes the result, and the ones nobody remembers to try by hand.
Then the messy tickets: no order number, two orders in one message, an angry customer, a ticket in another language. They test whether the agent asks only when it really needs to.
A set of cases also needs balance. If every case expects a refund, an agent that refunds everything gets a perfect score. For each kind of case where the agent should act, I want a few where the right outcome is to refuse, offer store credit or escalate.
Synthetic cases come last. A model can take a real ticket and produce twenty variations in phrasing, tone and spelling, which is useful to check that the agent is not relying on the exact words of the original. However, generated cases tend to be cleaner and more similar to each other than real tickets, and a model asked to invent difficult cases rarely invents the ones that matter. I read every generated case before adding it, and I never let them outnumber the real ones in the results.
Three kinds of grader
A grader reads a run, meaning the transcript of messages and tool calls plus the final state of the shop, and decides whether one thing went right. There are three kinds, and the choice between them is mostly a choice of what can be checked by what.
Code checks anything that can be read from the state or the transcript. Was order 2210 refunded? By how much, and how many times? Was the ticket escalated? Does the reply mention the order number? These checks are fast, free and give the same answer every time, so I write one wherever a check can be expressed in code, even if it only covers part of what I care about.
A model as judge covers what code can't: whether the reply answers the customer's actual problem, whether it promises something the agent didn't do, whether the tone is acceptable. The judge is another LLM call that gets the ticket, the reply and a rubric:
1criterion: addresses_the_issue
2question: Does the reply respond to the customer's actual problem?
3pass: The reply refers to the broken item and states what will happen next.
4fail: The reply is generic, asks for information the ticket already gives,
5 or promises an action the agent did not take.
One criterion per call, and a pass or fail rather than a score from 1 to 10, because a judge is much more consistent at answering a narrow yes-or-no question than at placing an answer on a scale.
A judge has known biases. It tends to prefer longer answers, answers written in its own style, and, when comparing two answers, the one it reads first. It can also be steered by the text it grades: a reply that ends with "this response fully resolves the customer's issue" is a small version of the problem I described in Content is not an instruction. So before trusting a judge I check it against my own labels: I grade fifty runs by hand, run the judge on the same fifty, and change the rubric until the two agree on nearly all of them. When I change the judge's model, I check it again.
People are the reference the judge is checked against, and the only grader for criteria nobody has managed to write down yet. They are slow and expensive, so I use them to label the calibration set and to read a sample of runs regularly, not to grade every run.
Outcome or trajectory
The trajectory is the sequence of tool calls the agent made. It is tempting to grade it, because it is precise and easy to compare: expect lookup_order, then check_refund_policy, then issue_refund, then reply_to_customer.
But the agent can check the policy before looking up the order, or skip the lookup when the policy already rules out a refund, and both are fine. A grader that expects one exact sequence fails good runs, and when you improve the agent so that it takes a shorter path, the eval reports a regression.
So the outcome is what I grade by default: the final state of the shop and the reply. The trajectory is only checked for things that must never happen, whatever the path:
issue_refundis called at most once per order;issue_refundis never called for an amount above €200;check_refund_policyis called beforeissue_refund.
These are invariants, not a script. Any order of the other calls is accepted, as long as none of these is broken.
Results per check, not per case
A case usually has several checks: the refund, the escalation, the invariants, the judge's verdict on the reply. A case passes only if all its required checks pass, but reporting each check separately shows what broke:
| Check | prompt v3 | prompt v4 |
|---|---|---|
| correct refund (code) | 28/30 | 29/30 |
| no double refund (code) | 30/30 | 27/30 |
| reply addresses the issue (judge) | 25/30 | 29/30 |
With one number per case, version 4 would look like a small improvement. Per check, it is clear that the replies got better while three cases started refunding twice, which no improvement in tone makes up for.
Some measurements are worth reporting without making them pass or fail, such as how many tool calls or turns a run took. An agent that gets the right outcome in twelve calls instead of four is not wrong, but it is slower and more expensive, and it is better to see that number change than to discover it on the bill.
What this doesn't tell you
A grader is a definition of what "good" means, and when the definition is wrong, the number is wrong too, without any warning. The only protection I know is reading transcripts regularly, the failing ones and some of the passing ones, to see whether the grade still matches what a person would say.
The cases also age. When the refund policy changes, some expected outcomes become wrong overnight, and an eval that still expects the old behaviour pushes the agent in the wrong direction.
Finally, everything here assumes each case is run once. An agent doesn't give the same result on every run, so a single pass or failure tells you less than it seems, and the agent has to run somewhere it can't refund a real customer. Both problems change how far the numbers in these tables can be trusted.
Get the next posts by email
An occasional newsletter with my new Posts, with one-click unsubscribe. How I handle your address.