Evaluating AI agents
How to tell whether a change made your agent better or worse: why agents need evals, how to build cases and graders for one, and how far to trust the numbers when every run can differ.
-
Part 1
Without evals, every change is a guess
Why an AI agent needs evals before its next prompt or model change, what makes an agent harder to evaluate than a single LLM call, and the task, run and grade loop every eval is built on.