Skip to content

Posts › Development

Development

Without evals, every change is a guess

Why an AI agent needs evals before its next prompt or model change, what makes an agent harder to evaluate than a single LLM call, and the task, run and grade loop every eval is built on.

7 min read

You change one sentence in an agent's system prompt, try it on the three tickets you remember, and the answers look better. You ship it. A week later someone notices that the agent has been refunding orders it should have escalated. Was it that sentence? There is no way to tell, because nothing measured the agent before the change or after it.

That is the problem evals solve. An eval is a set of tasks, run against the agent and graded automatically, that produces a result you can compare between two versions. Without one, every change to the prompt, the tools or the model is a guess, however carefully you look at the output. In this article I explain why trying an agent by hand stops working, what makes agents harder to evaluate than a single model call, and the basic loop every eval is built on.

The agent

Throughout the article I use a customer-support agent for an online shop. It reads a ticket, such as "my order arrived broken, I want my money back", and has four tools: lookup_order, check_refund_policy, issue_refund and reply_to_customer. The policy is simple: a full refund within 30 days of delivery, store credit after that, and anything above €200 goes to a person.

It is a small agent, but it fails in the ways real ones do. It refunds orders the policy says it shouldn't. It asks the customer for an order number that is already in the ticket. It calls issue_refund twice for the same order. Or it writes a polite, well-formed reply and never actually does anything.

Trying it by hand

Everyone starts by trying the agent on a few tickets and reading the answers. At the beginning this is the right thing to do: you learn what the agent does and what a good answer looks like. The trouble starts when it becomes the only check.

The tickets you try are the ones you remember, usually the ones from the last bug, so the cases nobody has thought about lately are never tried again. There are only a handful of them, because reading transcripts takes time. And you judge them by eye, which means the same answer can pass on Monday and look wrong on Friday.

Changes also interact. A sentence that makes the agent escalate large refunds more reliably can make it hesitant on small ones, and nobody tries a €40 refund while testing the €200 rule. The same happens without touching the prompt at all: when the provider releases a new model and you switch to it, every behaviour of the agent can shift a little, and three tickets are not enough to tell whether the new one is better.

What is missing is a baseline. If the same cases are graded the same way on every version, the question "is this change an improvement?" gets an answer, and so does "what did it break?".

Why agents are harder than a single call

Evaluating a single LLM call is already familiar to many developers: send an input, compare the output with an expected one, or check its structure. An agent is a loop instead. The model reads the ticket, decides to call a tool, reads the result, decides again, and stops when it thinks the job is done. Several things change because of that loop.

There are many steps. A mistake in the second step, like reading the wrong order, may only become visible in the fifth, or not at all. The final reply can look perfect while the refund went to the wrong order.

Actions have side effects. For a support agent the reply is the least important output. What matters is whether money moved, how much, and to whom. An eval has to look at the state of the world after the run, not only at the text the agent wrote.

There are many valid paths. The agent can look up the order and then check the policy, or the other way round. It can ask a clarifying question or work it out from the ticket. An eval that expects one exact sequence of tool calls fails agents that are doing the job well, and fails them again when they find a shorter way.

The same input gives different runs. Run the same ticket five times and you may get five different transcripts, and occasionally a different outcome. A single run, passing or failing, tells you less than it seems.

Runs are expensive. One run is several model calls and tool calls, so an eval of a hundred cases costs money and minutes, which limits how often it can run.

The consequence I take from this list is the main idea of agent evals: by default, grade the outcome, not the path. Check what the agent achieved, and check the path only for the things that must never happen along the way.

The loop: task, run, grade

Every eval, however large, repeats the same three steps.

The task describes one case: the input, the state of the world before the agent starts, and what success looks like. For the support agent it fits in a few lines:

yaml
 1id: broken-item-within-30-days
 2ticket: "My order #1042 arrived broken. I want my money back."
 3state:
 4  orders:
 5    - id: 1042
 6      total: 89.00
 7      delivered: 12 days ago
 8expect:
 9  refund:
10    order: 1042
11    amount: 89.00
12  reply: acknowledges the damage and confirms the refund

The task says what the outcome should be, not which tools to call or in what order.

The run puts the agent in an environment that holds that state, lets it work until it stops, and records everything: each message, each tool call with its arguments and result, and the final state. The environment matters: issue_refund must change a fake order in a test store, not send money to a real customer.

The grade turns the run into a result. Here there are two checks of a different kind. Whether order 1042 was refunded €89, exactly once, can be checked with ordinary code that reads the final state. Whether the reply acknowledges the damage needs judgment, from a person or from another model. Choosing between those graders, and deciding what else is worth checking, deserves a post of its own.

Repeat the loop over all the cases and you get a table you can compare:

Version Cases Passed Pass rate
prompt v1 30 24 80%
prompt v2 30 27 90%

The overall number is the least interesting part of that table. Version 2 might fix four cases and break one, and the broken one might be the €200 escalation, which matters more than the other four together. So I always read the cases that changed result between two versions before reading the total.

Where to start

I would not start with a framework or a large dataset. I would start with 20 to 50 cases written by hand, taken from real failures: the tickets where the agent did the wrong thing, plus the few that must never break, like the escalation limit and the double refund. Writing them down before the next change gives you the baseline the rest depends on.

Cases written from real failures are worth more than many generated ones, because they describe what actually goes wrong. Generating more cases, covering the edges of the policy and grading the answers well are the next steps, and each has its own trade-offs.

What this doesn't tell you

A passing eval says the agent handles the cases you wrote, graded the way you graded them. It says nothing about the tickets nobody has seen yet, so the quality of the cases sets the ceiling of what the eval can tell you.

The table above also hides a problem. The agent doesn't give the same result on every run, so the difference between 24 and 27 might be noise. Telling a real improvement from a lucky run takes more than one run per case. Measuring how much the result moves on its own, building a test environment the agent can't damage, and running evals on every change are what make the numbers worth trusting.