Skip to content

Series

Evaluating AI agents

How to tell whether a change made your agent better or worse: why agents need evals, how to build cases and graders for one, and how far to trust the numbers when every run can differ.

  1. Part 1

    Without evals, every change is a guess

    Why an AI agent needs evals before its next prompt or model change, what makes an agent harder to evaluate than a single LLM call, and the task, run and grade loop every eval is built on.