Without evals, every change is a guess
Why an AI agent needs evals before its next prompt or model change, what makes an agent harder to evaluate than a single LLM call, and the task, run and grade loop every eval is built on.
tags /
Why an AI agent needs evals before its next prompt or model change, what makes an agent harder to evaluate than a single LLM call, and the task, run and grade loop every eval is built on.