1 part
Evaluating AI agents
How to tell whether a change made your agent better or worse: why agents need evals, how to build cases and graders for one, and how far to trust the numbers when every run can differ.
Posts meant to be read in order, from the first part to the last.
1 part
How to tell whether a change made your agent better or worse: why agents need evals, how to build cases and graders for one, and how far to trust the numbers when every run can differ.
2 parts
What happens when the content an LLM application reads tries to give it orders, and how to keep it in check: first by keeping instructions apart and guarding each step of the call, then by having a separate judge read the input before the model does.