A graded bank of test tasks you rerun on every change, so that “did this edit make the agent better or worse?” has an answer you would defend with a number. Grown most honestly from real failures found in traces (Chapters 15 and 16). Previouseval Nextevaluator–optimizer