Home / Learn / Evaluating and observing agents

Guide

How to Evaluate AI Agents: Traces, Evals, Judges and Gates

How to evaluate AI agents: trace every run, grow an eval set from real failures, read pass@k and pass^k, calibrate the judge, gate releases. Start here.

By Enrique Gutiérrez · Last reviewed

Evaluating an AI agent means measuring, across repeated runs, whether its outcomes and the paths that produced them meet a standard you have written down. This guide shows how to evaluate AI agents: traces record what the agent did; evals say whether it should have. Without both, a passing demo is an anecdote and a quiet regression is invisible until a user reports it.

The method is Part V of the book, Chapter 15, Observability and Debugging and Chapter 16, Evaluating Agents (both in the full book), where the thesis bites hardest: an agent is only as trustworthy as the signal you can use to verify it. Chapter 1 makes what signal tells you it worked? the first bearing of the compass, and evaluation is that bearing given a budget and a schedule.

Why is evaluating an AI agent harder than testing software?

Evaluating an AI agent is harder than testing software because the component that matters is a distribution, not a function. The same input produces different outputs on different runs, each output is often one of many acceptable answers, and a multi-step agent produces a whole path whose early differences cascade into different routes through the task.

Chapter 2, free online, explains the mechanism: generation is a weighted lottery at every token. Chapter 16 adds three properties that break the unit-test analogy: there is often no single right answer, so exact-match assertions fail good outputs; correctness is often a judgment call; and the ground shifts: “Tests over a codebase move only when someone moves them; tests over a model move on their own.”

Agents add a fourth. An agent produces a trajectory, the full record of one run, and the grade has to attach to something. Grade only the final answer and an agent asked to deactivate an account that deletes it instead, then reports the account “no longer active,” scores full marks. And compounding errors bite: at 95% per step, a twenty-step task comes home about one time in three (try the compounding error calculator). Hence the book’s warning that “a single end-to-end success rate is an average, and an average is exactly the wrong summary for a system whose product is its worst runs.”

The cheapest honest response fits in an afternoon. Give the agent the same task three times, read all three outputs end to end, and ask of each only “would I accept this?” Everything else on this page scales that ritual up.

The cheapest eval you can run this afternoon.
Figure 16.2 The cheapest eval you can run this afternoon. Give the agent the same task three times, read all three outputs, and ask of each only “would I accept this?” Here runs 1 and 2 land on the same shape of answer while run 3 diverges (in accent); the spread you read by hand is the reliability envelope of the previous figure, measured with your own eyes. Repeated trials, a human grader, an acceptance criterion—the whole discipline in miniature. Reuse this diagram

What should a trace of an agent run contain?

A trace of an agent run should contain the complete, structured record of everything between the user’s request and the final answer: a tree of spans, one per model call and one per tool call, with each subagent’s run as a subtree, carrying timings, token counts, costs, versions and the full inputs and outputs.

The anatomy is borrowed from distributed tracing. The root span is the user’s request; each pass through the agent loop hangs a model-call span beneath it; each tool call hangs off the model call that asked for it.

The most common mistake is storing the prompt template and its variables instead of the rendered prompt, the assembled text the model actually received. A wrong retrieved document, a turn dropped by truncation, a variable that rendered empty: all invisible in the template, all plain in the rendered text. Hash the rendered prompt on each span, and one comparison shows whether the model’s input changed.

The rest of the inventory: tool name, full arguments and full result; tokens in and out; latency; cost; the model you requested and the model that answered; the reason generation stopped; and a structured error. Add a session summary span at the end of each run with steps, cost and final status, so “every run that failed yesterday, sorted by cost” is one query. The model’s narration is useful too, with a caveat the book states exactly: “In a trace, the tool calls and results are ground truth about what happened; the narration is testimony about why.”

One agent run drawn as a span waterfall.
Figure 15.2 One agent run drawn as a span waterfall. Each bar is a span: where it starts is when it happened, how long it runs is how long it took, and the indentation of its row records who called whom—the tool spans nest under the model call that asked for them, and everything nests under the root run. Every span carries the same record, the attributes bracketed over the first model call: the rendered prompt, the tokens, the cost, the latency. The bug is not the final answer but the first wrong span—here the long database call that erred, in accent—and a waterfall makes it the thing your eye lands on. The timings and ordering are illustrative. Reuse this diagram

Capture has a price in privacy: rendered prompts and tool results hold users’ data, so make content capture redactable, off by default and time-limited. Instrumentation is converging on a vendor-neutral standard (LLM tracing with OpenTelemetry); agent observability covers the tooling category.

Which agent bugs show up in a trace, and where?

Most agent bugs show up in a trace as one of six repeating shapes, Chapter 15’s bestiary, and almost every fix lands in the harness, the tools or the context rather than the model.

Bug shape Where it shows in the trace Where the fix lives
Stuck loop Session summary: 40 tool calls where healthy runs take 9; the same tool repeated after an error or empty result Retryability flag on errors, repeat-call detector, step budgets
Hallucinated tool arguments An argument whose value appears in neither the user’s text nor any earlier span Tighter types and enums, “do not guess, ask” in the tool description
Lost or poisoned context The rendered prompt of the failing step lacks the needed turn, or holds the early wrong fact Context assembly and compaction
Swallowed error or silent truncation A clean tool call followed by a clean, confident reply; result size jumps or a JSON array is cut mid-way Structured errors, truncate-and-say-so, result size on every span
Wrong stop condition Step counts pinned at the cap, or early exits with the goal visibly unfinished Stop conditions, tested against a scripted model stub
Silent degradation across deploys A good old trace and a bad new one disagree at some first step; responding model, prompt version or prompt hash differs Pin models, version prompts, record both on every span

The method never changes: read the trace forward to the first step where something is wrong, not the last step where the wrongness became visible, because that first step is the bug. When a failure refuses to happen twice, record every crossing of the nondeterministic boundary and replay it, the subject of how to debug an AI agent. The fuller field taxonomy is in AI agent failure modes.

How do you build an eval set for an agent?

You build an eval set for an agent by mining real failures from traces, naming the shapes they fall into, and writing each shape as a task: an input plus a way to grade the output, whether a reference answer, a checkable condition or a rubric. Every task must be solvable, and a reference solution is how you prove it.

First decide precisely what “would I accept this?” means, because “summarize the meeting well” lets two graders split on pass or fail and turns your metric into a measure of the grader’s mood. The book’s test is agreement: if two domain experts can disagree in good faith about whether an output passed, sharpen the criterion, not the agent. Anthropic’s guide to agent evals adds a useful diagnostic: with frontier models, “a 0% pass rate across many trials (i.e. 0% pass@100) is most often a signal of a broken task, not an incapable agent” (Anthropic, “Demystifying evals for AI agents,” January 2026).

Where should eval tasks come from?

Eval tasks should come from your own traces: traces reveal failures, failures become tasks, tasks guard against regression, and the agent returns to production to fail in newer ways. When you fix a failure, the trace in your hand is both the reproduction and the seed of an eval case. A suite grown that way tracks what your agent actually gets wrong; a suite written from imagination tracks what you feared, and the two overlap less than you would hope. The step-by-step version is building a golden dataset from real traces.

The trace-to-dataset flywheel, the loop that turns firefighting into improvement.
Figure 15.5 The trace-to-dataset flywheel, the loop that turns firefighting into improvement. Production runs are recorded as traces; the failures in them are sorted and become new eval cases (in accent, the handoff to the next chapter); those cases join the regression suite; the suite makes every later change safer; and safer changes go back to production to be traced again. Turn the loop once and a resolved failure becomes a test that guards every future release. Reuse this diagram

Include negative cases, tasks asserting what the agent should not do. A suite with only positive cases steers you toward an agent that, say, searches the web for everything, because nothing said not to. And read the failures: Chapter 16’s test is that they should seem fair, so a low score points at the agent rather than a grader bug. Keep timeouts and crashed sandboxes out of the quality number entirely.

How many tasks does an agent eval set need?

An agent eval set needs fewer tasks than most teams think to start and more than they think to compare versions. In the book’s words, “Twenty to fifty tasks you have hand-checked and believe in beat a few hundred synthetic ones nobody has read,” because early defects are large and a small net catches them.

Anthropic’s guide gives the same starting range and the reason: “In reality, 20-50 simple tasks drawn from real failures is a great start,” since early changes have large effects and “this large effect size means small sample sizes suffice.” The catch arrives when you ask whether a prompt edit moved the pass rate by a few points: a proportion measured on fifty tasks has a margin of roughly ten points either way. The eval sample-size calculator computes the interval for your own numbers, and how many eval examples you need works through the trade.

How to evaluate AI agents: what should you measure?

When you evaluate an AI agent, measure the outcome first, the state of the world after the run, and reserve trajectory checks for efficiency and safety. Track cost, latency and step count on the same fixed bank of tasks, and report pass rates over several trials per task, never single runs.

“Grade the outcome” has a sharper form: grade the state, not the prose. An agent that says the meeting is scheduled has produced a sentence; query the calendar. The book puts it this way: “The transcript can claim success while the world is unchanged, and only one of them is your product.” Avoid asserting the path, the exact sequence of tool calls, because agents keep finding legitimate routes the test author never imagined, and the over-specified check fails good behavior until nobody believes the suite. A path check earns its place in two cases only: a three-step task that took fifty steps, and a destructive tool that must never be touched. Agent trajectory evaluation shows how to grade the path without scripting it.

For grading, climb a ladder and stop as low as you can. Programmatic checks come first: the output parses, the tests pass, no internal identifiers leaked. Next come rubrics decomposed into yes-or-no questions, with several small graders each owning one dimension, so a failure tells you what broke. A model as grader sits at the top, for outputs no code can judge. AI agent evaluation metrics catalogs what to count and what to leave out.

Dimension What it asks Cheapest reliable grader When it misleads
Outcome Did the world end up in the right state? A state check: read back the row, run the code When only the final reply is checked, not the state
Trajectory Was the path efficient and safe? Step count, forbidden-tool assertions When it asserts one “correct” tool sequence
Quality of open output Is the answer good by a written standard? A rubric of binary checks, then a calibrated judge When the judge is uncalibrated or grades by length
Cost and latency What did the run consume? Span attributes summed per run When reported as a mean instead of percentiles
Reliability How often does it succeed, every time? pass@k and passk over repeated trials When a single run per task is treated as a rate

What is the difference between pass@k and passk?

The difference is which bound your product lives on: pass@k is the chance that at least one of k attempts succeeds, passk the chance that all k do. Equal at k = 1, they diverge as k grows.

pass@k fits when a cheap verifier picks a good answer out of the pile, such as code run against tests; passk fits an agent that acts unattended and must be right every time. The passk metric was introduced with the τ-bench customer-service benchmark, whose authors reported in 2024 that the strongest function-calling agents they tested succeeded on fewer than half the tasks and were “quite inconsistent (pass8 < 25% in retail)” (Yao et al., 2024). The figures describe the models of that year; the gap they expose does not.

The space between the two curves is what the book calls the reliability envelope: “A narrow envelope is a steady agent. A wide one is a talented gambler: capable of the task, not to be trusted with it.” At 90% per attempt, ten attempts almost certainly include a success, while all ten succeed only about 35% of the time. Both descriptions are true; quote the one your product depends on. The pass@k calculator draws the envelope for your own runs, and pass@k vs passk goes deeper. Glossary: pass@k and passk.

One agent, two honest summaries of the same ten attempts, drawn here for an illustrative per-attempt success rate of 90 percent.
Figure 16.1 One agent, two honest summaries of the same ten attempts, drawn here for an illustrative per-attempt success rate of 90 percent. The chance that at least one of k attempts succeeds (pass@k) climbs toward certainty; the chance that all k succeed (passk) decays toward zero. The widening gap between them is the reliability envelope: the wider it is, the more capable-but-erratic the agent—and a demo only ever shows you the top curve. Reuse this diagram

Is LLM-as-a-judge reliable enough to grade agents?

LLM-as-a-judge is reliable enough to grade agents once you calibrate it against human labels, and not before. A strong judge approximates a human grader about as well as a second human does, no better, and it arrives with documented biases in known directions. Treat it as an instrument, never as ground truth.

The widely cited study by Zheng and colleagues found that strong judges “can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans” (Zheng et al., NeurIPS 2023). The same paper catalogued position, verbosity and self-enhancement bias. Chapter 16 adds leniency and a mitigation for each: run pairwise comparisons both ways and keep only verdicts that survive the swap, check scores for a correlation with word count, judge with a different model than the one that generated, and reserve the top grade for flawless.

Write the rubric for a new hire with a worked pass and fail, have the judge justify before it rules, give it the reference answer, and demand a parseable verdict. Then calibrate: hand-label a few dozen real outputs, run the judge over the same sample, and report Cohen’s kappa beside raw accuracy, because a judge that answers “good” every time on a set that is 70% good scores 70% accuracy with zero information. Split out precision and recall for pass/fail verdicts. The book’s summary is one line: “A judge you have not calibrated is an opinion you have automated.”

When should you skip the judge?

Skip the judge whenever a code check can decide the question; the check is cheaper, deterministic and immune to every bias above. The calibration method in full is in is LLM-as-a-judge reliable?, rubric design in writing an LLM-as-a-judge rubric (with its preference for a small ordinal scale), and the division of labor in LLM-as-a-judge vs human evaluation. The same judge returns as the critic in the evaluator–optimizer pattern, where an oracle outside the model is what keeps the loop honest.

How do evals gate a release?

Evals gate a release when a threshold is enforced instead of glanced at: every change to a prompt, a tool definition or the model reruns the suite, and below a set pass rate on the regression tasks, or on any failure of a safety case, the change does not merge. That enforced threshold is a regression gate.

The gate ends whack-a-mole: a prompt is read at every step, so any edit is global; fix refunds and cancellations quietly break. Chapter 16 calls the exit eval-driven development: the suite, like a test suite, defines progress. It needs two things: prompts and tool definitions versioned in the repository, so a regression can be attributed to a change, and comparisons of distributions rather than single runs.

A healthy suite splits in two. Capability evals are hard tasks the agent mostly fails, the hill you are climbing. Regression evals are tasks it reliably passes, held near 100%. Stable capability tasks are promoted to the regression suite, and harder ones replace them, which keeps the suite from saturating into permanent green. The economics follow slow test suites: a smoke subset on every commit, the full multi-trial battery nightly and before release. Details are in regression testing for LLM applications and how to test AI agents; the gate then hands off to a shadow deploy and canary, the rollout ladder in shadow deploy, then canary.

Eval-driven development as plumbing.
Figure 16.4 Eval-driven development as plumbing. Every change runs against two suites: the capability evals it is meant to climb, and the regression evals it must not break; tasks that stabilize are promoted from the first into the second. The change reaches a gate (in accent) where a threshold is enforced, not glanced at—the numbers, not moods, decide whether it merges or is blocked. Reuse this diagram

What should you monitor once the agent is live?

Once the agent is live, monitor the signals a user would recognize as the product working, plus the ones that predict your invoice: task success and output quality, cost per run, latency at the percentiles, the step-count distribution, error and retry rates per tool, and the reasons generation stopped.

The signal no single run can show is drift: the users, the model behind an alias, or the retrieved documents change, and quality sags with no commit to blame. Because drift is a trend, alert on week-over-week change as well as absolute thresholds. Online signals (feedback, retries, a calibrated judge on sampled traffic) surface failures your offline set never imagined; each becomes a new task. Read traces on a schedule, worst first. Chapter 15’s line is the reason: “a flagged trace is an eval case that has not been written down yet.” See how to monitor AI agents in production and how do you know if an AI agent is working?.

Worked example: evaluating a refund agent from trace to gate

A support agent handles refunds. The numbers are illustrative; the statistics are standard.

Trace. A customer reports a refund confirmed but never issued. The trace shows a clean process_refund span returning an error object that a middleware layer flattened to an empty result, followed by a confident reply. That is the swallowed-error shape. The fix is a structured error, and the trace becomes a task: same input, graded by reading back the refund record, not the reply.

Eval set. Three weeks of fixes like this produce 50 solvable tasks, each with a reference solution, including eight negative cases (refunds the policy forbids). Each task runs five times.

Reading the result. The agent passes 41 of 50 tasks on the majority of trials: 82%. The 95% Wilson interval is 69.2% to 90.2%, a twenty-point band. A prompt edit raises the score to 44 of 50, or 88%. A two-proportion test puts that six-point gap at z ≈ 0.84, nowhere near significance, so the gate should treat it as “no regression,” not “improvement.” Repeated trials of one task are not independent, which makes the real evidence weaker still.

Reliability. Per attempt, refunds succeed about 90% of the time. pass@5 is effectively 100%; pass5 is 0.9⁵ ≈ 59%. Refunds run unattended, so the agent lives on pass5: five runs of the same refund include at least one failure about 41% of the time.

Can the team trust its tone judge?

Not yet, and raw accuracy would have hidden it. Tone is graded by an LLM judge. The team hand-labels 60 real replies: 42 good, 18 bad. The judge agrees on 52 of 60, or 87% accuracy. Its confusion counts are 39 good–good, 13 bad–bad, 5 bad replies passed and 3 good replies failed. Expected chance agreement is (42×44 + 18×16) / 60² ≈ 0.59, so kappa = (0.867 − 0.593) / (1 − 0.593) ≈ 0.67. The judge catches 13 of 18 bad replies (72% recall). The book’s rough yardstick for a judge standing in for a human is a kappa around 0.7 to 0.8, so this judge needs a sharper rubric before it grades unsupervised.

Gate. The regression suite blocks any merge whose pass rate falls below the current floor and stops on any failure among the eight forbidden-refund cases, with no override in the normal path.

Where does this method stop working?

This method stops working where its inputs stop holding. A suite measures the failures you already found, so a green suite says nothing about the ones you have not. Graders have bugs, and agents occasionally learn a loophole in one. Calibration decays when the judge’s underlying model is updated or your traffic shifts. The book’s phrase is the right posture: “a green suite is evidence, never proof.”

The cost is real too. Multi-trial evaluation multiplies the compute bill by the trial count, a model judge adds a second bill, and calibration takes senior human hours. Below some scale, full statistical rigor costs more than the failures it prevents; Chapter 16’s advice is to buy as much certainty as the cost of being wrong justifies, and no more. Some domains also lack a cheap check altogether, the verification gap of research agents, and there the judge and human review carry more weight than this page’s ladder would suggest.

What survives every scale is the stance. Whether the agent is good is an empirical question, settled by running it and counting: change one thing, measure, compare, keep or revert. Or, as Chapter 16 has it, “Intuition proposes; measurement decides.”

Which chapter, tool or post covers each piece?

Each piece maps to one section of the book and one or two pages here: the chapter argues, the tool computes, the post goes deeper. Every eval concept here traces back to Chapter 15 or 16.

Concept Book Tool Deeper post
Traces, spans, rendered prompts Chapter 15, “Traces and Spans” Agent observability, LLM tracing with OpenTelemetry
Record and replay, bug shapes Chapter 15, “Reproducing Nondeterministic Failures” and “A Taxonomy of Common Agent Bugs” How to debug an AI agent, AI agent failure modes
Drift and production monitoring Chapter 15, “Production Monitoring” How to monitor AI agents in production
Eval sets and sample size Chapter 16, “Building an Eval Set” Eval sample-size calculator How many eval examples do you need?, Golden dataset from real traces
pass@k, passk, reliability envelope Chapter 16, “Why Evaluation Is Hard” pass@k calculator pass@k vs passk
LLM-as-a-judge calibration Chapter 16, “LLM-as-a-Judge” Is LLM-as-a-judge reliable?, LLM-as-a-judge rubric
Regression gates Chapter 16, “Eval-Driven Development and Regression Gates” Regression testing for LLM applications, How to test AI agents
Why runs differ at all Chapter 2 (free) Compounding error calculator Why agent errors compound

Chapter 15 closes on the image that ties the part together: “The traces are the flight recorder.” Chapter 16 is the flight plan. Read both and you can say, with a number you would defend, whether your agent is good and whether yesterday’s change made it better or worse.

The chapters behind this guide

  1. Chapter 15: Observability and Debugging In the full book
  2. Chapter 16: Evaluating Agents In the full book

Tools and explainers for this topic

Tool

Compounding error calculator

A free compounding error calculator for AI agents: whole-run success from per-step reliability, and the reliability a long task needs.

Tool

Eval sample-size calculator

A free eval sample size calculator for AI agents: confidence intervals for a pass rate and the number of tasks you need, computed in your browser.

Tool

pass@k and pass^k calculator

Compute pass@k and pass^k for an AI agent from its per-attempt success rate or your own runs, and see the reliability envelope. Free pass@k calculator.

Explainer · 3 min

The agent loop: four beats and three exits

A three-minute animated explainer of the agent loop: the four beats of every pass, the history that is the agent's only memory, and three exits ranked by trust.

Articles in this cluster

Questions readers ask

How many eval examples do I need for an AI agent?
To find the large defects of a young agent, twenty to fifty hand-checked tasks drawn from real failures are enough. To tell whether a change moved the pass rate by a few points you need hundreds of independent trials, because a pass rate on fifty tasks carries a margin of roughly ten points either way.
Is LLM-as-a-judge reliable?
It can be, once calibrated. A strong judge agrees with human preferences about as often as two humans agree with each other, but it has documented position, verbosity, self-preference and leniency biases. Hand-label a few dozen real outputs, measure agreement with Cohen’s kappa, and use a code check instead wherever one can decide the question.
What is the difference between pass@k and pass^k?
pass@k is the probability that at least one of k attempts succeeds, the right number when a cheap check picks the good answer. pass^k is the probability that all k attempts succeed, the right number for an agent that acts unattended. They are equal at k = 1 and pull apart as k grows.
How do you test an AI agent when every run is different?
Split the system at the seam between model output and your own code. Test tools, parsers and loop control with exact assertions and a scripted model stub. Test model behavior with property checks and pass rates over several trials per task, never single runs.
What should I monitor for an AI agent in production?
Task success and output quality, cost per run, latency at the percentiles, the step-count distribution, error and retry rates per tool, and the reasons generation stopped. Alert on week-over-week trends as well as thresholds, because drift moves through every threshold slowly.