AI agent evaluation metrics say whether an agent did its job, how it got there, what that took, what it must never do, whether it would do the same tomorrow, and whether the grader can be believed. No single score covers all six. Report four numbers, each with its denominator and its interval, and treat the rest as diagnostics.
The four answer four questions: did the run reach the right end state, does it do so every time, did it break a rule, and what does a success cost? Choosing them is this post’s synthesis. The list does not come from the book this site belongs to, and it is not a published standard, so the sections below give the reasons and you can disagree with them.
Which AI agent evaluation metrics should you report first?
Report four: task success rate graded on the end state, passk, a policy-violation count, and cost per successful task. Each answers one question a release decision turns on, and each exposes a failure the other three cannot see. The panel is this post’s choice; the book supplies the parts.
Chapter 16 of AI Agents, Engineered asks you to “track cost, latency, and step count on the same fixed bank,” and sets the safety rule for a regression gate: “any failure on the safety cases, the pipeline stops.” Its warning about a lone rate is that “an average is exactly the wrong summary for a system whose product is its worst runs.”
| Question | Number | Denominator | State beside it | What it still hides |
|---|---|---|---|---|
| Did it reach the right end state? | Task success rate | Scored trials; infrastructure errors excluded and reported | Trials per task; a 95% interval computed across tasks | Rule-breaking successes; run-to-run spread |
| Does it do so every time? | passk (pass@k only where a test or a person picks the winner) | Tasks, each run n ≥ k times | k, n, and an interval on the share of tasks | Why a task flips |
| Did it break a rule? | Policy-violation count | Every run checked, negative cases included | The rules checked; the number of runs; if the count is zero, the 95% upper bound | Harm through an action no rule covers |
| What does a success cost? | Cost per successful task | Trials graded successful (the numerator is the cost of all trials) | Step count and tail latency on the same task set | Every flaw of the success grader; human review time |
Read the third row as a count with a gate at zero, and remember that zero violations in n runs is not a zero rate. If runs are independent, the 95 percent upper bound on the true rate solves (1 − p)n = 0.05, which is about 3 ÷ n. Two hundred clean runs still allow about 1.5 percent, and 40 allow about 7 (standard statistics, computed for this post).
Keep it a count all the same. A rate invites someone to average it into the other three, and that is the one thing the panel must prevent.
Worked example: the same four numbers for two agents
Two illustrative agents show how the panel changes shape while the four questions stay fixed. A support agent that issues refunds writes to a system you can query. A research agent that writes reports leaves a document and a trace, and nothing else.
| Number | Support agent that issues refunds | Research agent that writes reports |
|---|---|---|
| Task success rate | Trials where the ledger shows one refund, for the right order and amount, and no other record changed ÷ scored trials. Grader: a code check on the ledger. 5 trials per task. | Trials where the report passes every binary check (each claim cites a source the run fetched; every required section is present) ÷ scored trials. Grader: code for citations, a calibrated judge for the rest. 5 trials per task. |
| passk | pass5, from 5 trials per task. The agent acts unattended, so all five trials of a task must pass. | pass5 if the report goes out unread. pass@3 only if a person really picks one of three drafts. |
| Policy-violation count | Runs that refunded above the limit, skipped the identity check, or refunded an order outside the window. Gate at zero. | Runs that fetched from a blocked domain, sent the user’s data to an outside service, or exceeded the fetch budget. Gate at zero. |
| Cost per successful task | Cost of all trials ÷ trials that passed the ledger check. | Cost of all trials ÷ trials that passed every check. |
The research agent has no end state to query, so its task success rate is a pass-all rate over binary checks, and part of its grader is a model. That makes the judge-agreement row of the reference table mandatory for it. Its violation rules are ones the success checks cannot see, so the third number stays independent of the first.
What does each agent metric measure, and how does it mislead?
Each of the AI agent evaluation metrics below answers one of six questions: outcome, path, cost, safety, stability, or whether the grader itself can be believed. The table has 23 rows, and the buttons narrow it to the question you came with. The six family names are this post’s grouping.
Formulas are given in words as their source states them. Where the “Defined in” column says no published origin was found, the formula is this post’s wording of common usage, and tools will differ. The “How it misleads” column is my analysis unless it names a source.
| Metric | What it measures | Formula in words | Needs | How it misleads | Defined in | Answers |
|---|---|---|---|---|---|---|
| Task success rate | Whether the job got done | Mean over tasks of passed trials ÷ trials run | A grader per task, ideally a state check | Graded on the final message, it counts claimed and rule-breaking runs as wins; one trial hides the spread | No single origin; stated as E[c/n] by Yao et al., 2024 | outcome |
| Pass-all rate | Whether an example passed every check | Examples that pass every check ÷ examples | A set of binary checks | Hides which check failed unless you can drill in; moves when a check is added | Husain and Shankar, 2026 | outcome |
| Completion rate | Whether the run finished | Runs that ended with a final answer ÷ runs started | Run status only | Finishing is not succeeding; one benchmark uses the same name for task success (Levy et al., 2024) | No published origin found | outcome |
| Completion under policy | Success that broke no rule | Tasks where all success checks pass and policy violations total zero ÷ tasks | A written policy list; a violation check | Only as complete as the policy list | Levy et al., 2024 | outcome, safety |
| Progress rate | How far a run got | The largest share of hand-labeled subgoals matched at any one step so far (or the best state-match score, where matching is continuous) | Subgoals labeled per task | Credits an unfinished job; can rise while task success stays flat | Ma et al., 2024 | outcome, path |
| User signals | What users did next | Share of sessions with an immediate rephrase, a retry, an abandon or explicit feedback | Product telemetry | Silence is not satisfaction; skewed toward users who bother to react | No published formula; named in Chapter 16 | outcome |
| pass@k | Chance that at least one of k attempts works | Mean over tasks of 1 − C(n−c, k) ÷ C(n, k), from n ≥ k trials of which c passed | n ≥ k trials per task; something that picks the winner | Climbs with k toward the share of tasks the agent can ever solve; means little with no picker | Chen et al., 2021; idea in Kulal et al., 2019 | stability |
| passk | Chance that all k attempts work | Mean over tasks of C(c, k) ÷ C(n, k) | n ≥ k independent trials per task | Unreadable without k; pk on a pooled rate is a different number | Yao et al., 2024 | stability |
| Per-task pass threshold | Whether each task holds up under repeats | Tasks passing at least m of n trials ÷ tasks | Repeats per task | With m = n it is passk under another name; a pooled threshold lets one always-failing task hide | This post’s name; the m = n case (“all-pass@k”) is Levy et al., 2024 | stability |
| Tool-call precision and recall | Whether the expected calls were made, and only those | Precision: calls matching name and parameters ÷ (those + unexpected calls). Recall: the same matches ÷ (those + expected calls not made) | A reference set of calls per task | Precision is perfect for a run that stopped early; matching rules differ by tool (names only, or names and parameters; repeated calls counted each time, or matched one-to-one) | One open-source library’s documentation (Ragas, read 2026); no paper origin found | path |
| Required and forbidden calls | Whether evidence was fetched and the never-list respected | Runs missing a required call, or making a forbidden one ÷ runs | Two short lists per task | Harm through an allowed tool; a required call whose result was ignored | No published origin found | path, safety |
| Step count | Effort and wandering | Model turns or tool calls per run, read as a distribution | Traces | The mean hides the tail; fewer steps is worse if a check was skipped | No formula to cite; Chapter 15 reads it as a distribution | path, cost |
| Tool error and retry rate | Friction at the tool boundary | Failed or retried calls ÷ calls, per tool | Span attributes | A retry counts as one operation or two attempts, depending on the tool | No published origin found; named in Chapter 15 | path |
| Cost per task | What one run takes | Tokens or money summed over one run, retries and subagent calls included | Usage metering per run | Falls when the agent gives up early | Named, not defined, in Anthropic’s 2026 guide | cost |
| Cost per successful task | What one correct result takes | Cost of all runs, failed ones included ÷ runs graded successful | Cost and a success grade on the same runs | Inherits every flaw of the success grader; leaves out human review time | Nearest published form: cost-of-pass, per problem, for language models (Erol et al., 2025, preprint); the pooled agent form is practitioner usage | cost |
| Latency | How long someone waits | Time to first token and total completion time, read at percentiles | Timestamps | The mean hides the tail; streaming hides total time | Chapter 19 | cost |
| Human intervention rate | How often a person steps in | Interventions ÷ a unit you must state (turn, task, session) | Approval, escalation and takeover events | Can rise as experience rises; falls when people stop looking | No published definition found | cost, safety |
| Policy-violation count | What it did that it must not | Runs with at least one violation of a written rule, counted per rule | Negative cases; a violation check | Zero in n runs is not a zero rate; turned into a rate, it gets averaged away; a model-based violation check has its own error | No single origin; this post’s wording | safety |
| Attack success rate | Whether an adversary can steer it | Security cases where the attacker’s goal is met ÷ security cases | An adversarial test set | Pooled over repeats, it understates an attacker who retries; a static set ages | Debenedetti et al., 2024 | safety |
| Judge agreement | Whether a model grader matches people | Cohen’s kappa between judge verdicts and human labels, with precision and recall | A few dozen human labels | Raw agreement flatters; ten labels give “a random number” (Chapter 16) | Standard statistics; Chapter 16 asks for it | grader |
| Confidence interval | How wide a rate really is | A range around a rate, computed from the passes and the total | The count and the total | Related tasks make it too narrow; says nothing about a wrong grader | Standard statistics | grader |
| Do-nothing baseline | Whether the grader passes an idle agent | Success rate of an agent that returns an empty reply and calls no tool, on the same tasks and graders | A stub agent; the full task set | A zero clears one grader fault only; a task whose right answer is to do nothing needs a positive assertion | This post’s name; the check is from Zhu et al., 2025 | grader |
| Coverage | How many runs the number rests on | Counts of scored, errored and unscored runs | Run statuses | Dropped runs leave the denominator silently | No published origin found; Chapter 16 asks for a separate status | grader |
A table of metric names is the kind of generic list some practitioners warn about. Read it as vocabulary for definitions you write against your own failures; Husain and Shankar’s evals FAQ (2026) allows generic metrics only as “exploration signals to identify traces worth reviewing.”
How do the six families fit together?
Outcome says whether the job was done, and the other five families say how far to believe that. Confusion about AI agent evaluation metrics starts, in my reading, when a number from one family is taken as the answer to another family’s question.
Outcome and path
Outcome is the family the panel starts with, and the book’s rule for it is “grade the state, not the prose.” Progress metrics belong here as development signals: one benchmark scores half the share of checkpoint points earned plus half the full-completion flag (Xu et al., 2024), which is useful below 100 percent and wrong as a gate.
Path metrics explain why an outcome moved. A tool call can be scored against a reference set, checked against required and forbidden lists, or simply counted. The matchers, and what each lets through, are a subject of their own in agent trajectory evaluation; why per-step accuracy is not task success is the subject of why agent errors compound.
Cost and safety
Cost covers money, time and human attention. Chapter 15 asks for “latency at the percentiles rather than the mean, because agents produce long tails,” and calls the step-count distribution “your loop detector.”
Human intervention rate has no published definition that I could find. In one vendor’s 2026 telemetry on its own coding product, new users interrupted in about 5 percent of turns and experienced users in about 9 percent, while sessions run with full auto-approval rose from roughly 20 percent to over 40 percent (Anthropic, 2026). The rate went up as experience and auto-approval went up.
Safety numbers describe a tail, so they are counted and gated, never averaged. Debenedetti and colleagues (2024) define targeted attack success rate as “the fraction of security cases where the attacker’s goal is met,” where a security case pairs a user task with an injection task. The source reports it beside benign utility and utility under attack, since an agent that refuses everything scores a perfect zero on attack success.
Stability and the grader
Stability asks whether the number would survive a rerun. A 2026 paper that evaluated 15 models on two benchmarks reports that “outcome consistency remains low across all models,” meaning “agents that can solve a task often fail to do so consistently” (Rabanser et al.). pass@k and passk are the two standard summaries, and the next section defines them.
Grader metrics measure the instrument. If an LLM judge scores any row, its agreement with human labels is part of the metric’s definition; the method is in whether an LLM judge is reliable, and the judge agreement calculator does the kappa arithmetic. Intervals and sample sizes belong to how many eval examples you need.
What is pass@k, and how does passk differ?
pass@k is the probability that at least one of k attempts at a task succeeds; passk is the probability that all k succeed. Chapter 16 calls the first “the optimistic bound” and the second “the pessimistic bound,” and the gap between them the reliability envelope.
The name pass@k and its estimator come from code generation. Chen and colleagues (2021) “generate n ≥ k samples per task,” count “the number of correct samples c ≤ n which pass unit tests,” and report the mean over problems of 1 − C(n−c, k) ÷ C(n, k). Kulal and colleagues (2019) had reported the same idea as “success rate at B.”
Chen’s paper also says how to read the number: as “the best out of k samples, where the best sample is picked by an oracle with prior knowledge of the unit tests.” Without a picker, pass@k describes a product you do not have.
passk was introduced for agents. Yao and colleagues (2024) define it as “the chance that all k i.i.d. task trials are successful, averaged across tasks,” estimated as the mean over tasks of C(c, k) ÷ C(n, k). In that paper, one frontier model of 2024 passed about 61 percent of one domain’s tasks on a single trial, and its pass8 fell to about 25 percent.
One worked number
Take an illustrative rate. Suppose every attempt succeeds 90 percent of the time, independently. Then pass10 = 0.910 = 34.9 percent, and pass@10 = 1 − (1 − 0.9)10, which the calculator below prints as “> 99.9%”. Chapter 16 calls this agent “a pass@k triumph and a passk catastrophe, and both descriptions are true.”
With JavaScript on, the pass@k and passk calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the pass@k and passk calculator on its own page to share a result by link.
The sliders use those closed forms, which assume one shared success rate and independent attempts. Under “Use your own runs instead,” enter the attempts and passes for one task, and for any k up to n the tool switches to the two estimators above. Which number a product should quote is the whole question of pass@k versus passk.
The k is a product decision: the number of consecutive unattended runs that must all be right. Write it into the metric’s definition and run n ≥ k trials per task. With n = k each task scores 0 or 1, so passk on 40 tasks is exactly as precise as a pass rate on 40 tasks, and it needs the same interval.
How do you know whether an agent metric can be trusted?
A number can be compared with last week’s when three things are stated beside it: the trials per task, the interval, and what was counted as an infrastructure error. Without them, a change in the number has explanations other than the agent, and you cannot tell them apart.
Sampling. A suite that reports 90 percent from 45 passes on 50 tasks, each run once, has a 95 percent Wilson interval of 78.6 to 95.7 percent (standard statistics, computed for this post; the eval sample size calculator reproduces it). Chapter 16 puts the consequence plainly: “a two-point shift on a fifty-task suite is well inside the noise.”
The environment. One lab’s 2026 experiment on its own model held the model, harness and tasks fixed and changed only the resource limits of a coding benchmark’s sandbox. Infrastructure error rates fell “from 5.8% at strict enforcement to 0.5% when uncapped,” and the score rose by 6 percentage points between the strict setting and no cap at all (p < 0.01). The author’s advice is that “leaderboard differences below 3 percentage points deserve skepticism until the eval configuration is documented and matched” (Anthropic, 2026). Excluding errored runs has its own risk: if errors fall mostly on long or hard tasks, the rate is flattered, so show it with errors counted as fails beside it.
The grader. When a model does the grading, a score can rise because outputs drifted toward what the judge prefers. Pin the judge’s version in the definition, and re-measure its agreement with human labels whenever either changes.
Why does the headline number stay green while the agent breaks?
A headline success rate stays green because it grades one thing, usually the final state or the final message, and the break happened somewhere it does not look. The published measurements differ in denominator, so read each with its own.
- Successes that broke a rule. A 2026 preprint ran three frontier models on two domains of one customer-service benchmark, four trials per task, and found that “27–78% of benchmark reported successes are corrupt successes concealing violations across interaction and integrity” (Cao, Driouich and Thomas). That is a share of credited successes, with violations identified by a model judge the authors checked by hand.
- Success under policy. Levy and colleagues (2024) scored three open web agents and report that “their average CuP is less than two-thirds of their nominal completion rate.” Both rates share one denominator, all tasks.
- A grader that passes an idle agent. Zhu and colleagues (2025) found that on one tool-use benchmark “a trivial agent that returns empty responses is considered successful on intentionally impossible tasks” and “achieves a 38% success rate.”
- Wrong arguments. In a July 2026 forum post from an account named after an evaluation vendor, the author writes that “Task-success stayed green until we added argument-level scoring.” It is one anecdote.
The panel catches some of these and not others. Its policy-violation count would have moved in the first two cases while the success rate stayed put. The third is a grader fault that none of the four numbers sees, which is why the reference table carries a do-nothing baseline in the grader family.
The fourth depends on the task. Where a wrong argument leaves a wrong record, grading the end state catches it, which is why the first number is defined on state. Where the call only reads, it takes an argument-level path check.
How do you define a metric so two people compute the same number?
Write down the unit, the numerator, the denominator, the grader, the trials and the aggregation before anyone computes anything. One name hides two formulas whenever one of those is left to the tool’s default, and the defaults differ.
pass@k is the clean example. The estimator’s own authors warn against the shortcut of plugging the empirical pass@1 rate into 1 − (1 − p)k: “we show that it is biased” (Chen et al., 2021, section 2.1). Their Appendix A gives the direction: the shortcut “results in a consistent underestimate.” Two dashboards can both say pass@10 and disagree, one using the per-task estimator and one the shortcut.
Units cause the same split. An August 2026 discussion in one open-source tracing project asks how to count an agent that retries one call and then succeeds: “1 operation, 2 attempts (operation success 100%, attempt success 50%)” or “2 operations with 50% success.” Denominators do too: a July 2026 issue in an open-source evaluation framework reports that “a run with 1 success and 9 errors” under a tolerant error setting was shown as plain accuracy 1.0. Both are reports about particular tools at a date, cited as examples of the category.
The template below has fourteen fields. Copy it once per metric on your panel.
metric: <name, exactly as it appears on the dashboard>
question: <the one question this number answers>
gate: <decision it feeds: ship | roll back | investigate>; threshold: <value, or "zero">
unit: <task | trial | turn | tool call | session>; a retry <is | is not> a new unit
numerator: <what is counted>
denominator: <what it is divided by>; excluded: <what, and where it is reported>
grader: <code check | rubric | judge | human>, version <id>
if a judge: agreement with human labels <kappa, n labels, date>
trials_per_task: <n>, each from a clean environment
errored trials: <rerun up to r times | dropped>
a task left with fewer than <k> scored trials is <excluded and listed | scored as a fail>
aggregation: <mean over tasks of passed/trials | pass^k, k=<k> | pass@k, k=<k>
| tasks passing >= m of n | count>
uncertainty: <interval across tasks: method and level>
or <count out of n runs; if zero, the 95% upper bound, about 3/n>
environment: <agent version: code, prompt, model, tool schemas>
limits: <steps, time, resources>; a run that hits a limit is scored <fail | infra_error>
infra_error: <what counts>; reported separately, never as a fail
task_set: <name and version, n tasks>; covers <request types>; leaves out <request types>
or live traffic: <sampling rule, window>
blind_spot: <the bad run this number still passes>; caught by <companion metric>
owner: <who>; grader last checked against read transcripts on <date>
Example: task success rate for the refund agent
Here the template is filled for the support agent’s first number. Every value is illustrative.
metric: Task success rate (refund agent)
question: Did the refund end in the right state?
gate: ship; threshold: lower end of the 95% interval at or above 80%
unit: trial; a retry inside a run is not a new unit
numerator: trials where the ledger shows one refund, for the right order
and amount, and no other record changed: 180
denominator: scored trials: 200; excluded: 4 errored executions,
listed under infra_error
grader: code check against the refund ledger, version 7; no judge
trials_per_task: 5, each from a reset sandbox
errored trials: rerun once (4 reruns, all scored)
a task left with fewer than 5 scored trials is excluded and listed (none)
aggregation: mean over the 40 tasks of passed trials / 5
30 tasks at 5 of 5, 10 tasks at 3 of 5:
(30 x 1.0 + 10 x 0.6) / 40 = 90.0%
uncertainty: 95% interval across the 40 per-task rates, mean +/- 1.96 x SE
SE = 0.175 / sqrt(40) = 0.028, so 84.6% to 95.4%
(not the pooled Wilson interval on 180 of 200: 85.1% to 93.4%)
environment: agent build 41 (code, prompt, model, tool schemas pinned)
limits: 12 steps, 120 seconds; a run that hits a limit is scored fail
infra_error: sandbox crash, tool or model endpoint not responding, rate limit
task_set: refunds-v3, 40 tasks from real tickets; covers single-order
refunds; leaves out partial refunds and chargebacks
blind_spot: a correct refund issued without the identity check;
caught by the policy-violation count
owner: <name>; grader last checked against 20 read transcripts on <date>
The uncertainty line follows the rule in how many eval examples you need: the task is the unit, so average each task’s trials into one score and compute the interval across tasks. The 40 per-task rates have a standard deviation of 0.175, so the standard error is 0.175 ÷ √40 = 0.028 and the interval is 90 ± 1.96 × 2.8 points, or 84.6 to 95.4 percent. That is a normal approximation; at 40 tasks, cross-check it with a bootstrap over tasks.
The pooled Wilson interval, 85.1 to 93.4, treats 200 trials as independent and is too narrow. Here the across-task standard error is 0.028 ÷ 0.021 = 1.3 times the pooled one, and the ratio depends on the data. Miller’s 2024 paper on error bars reports clustered standard errors “over 3X larger than naive standard errors,” but that figure is for one reading benchmark whose questions share passages (3.05 times; 1.10 and 1.88 on two others), not for repeated trials.
The second and third numbers
For passk, k = 5 is the product’s choice, five unattended refunds in a row, and n = k, so a task scores 1 only when every trial passed. That gives pass5 = 30 ÷ 40 = 75 percent, with a 95 percent Wilson interval on 30 of 40 tasks of 59.8 to 85.8 percent.
A dashboard that takes the pooled shortcut instead reports 0.95 = 59 percent. That sits just under the interval’s lower end, so 40 tasks do not measure the sixteen-point gap well. What the example shows is the direction: when tasks differ in difficulty, the pooled shortcut understates passk.
The third number is a count: zero violations in 200 runs, which bounds the rate at about 1.5 percent only if the runs were independent. Five trials of one task are not, so the bound is optimistic.
What does a successful task cost?
Cost per successful task is the cost of all runs, failed ones included, divided by the number of runs graded successful.
The nearest published definition is a 2025 preprint on language models. It defines cost-of-pass for one problem as “the expected monetary cost to obtain one correct solution”: the expected cost of one attempt divided by the probability of a correct answer (Erol et al.). The pooled form for agent tasks, used here, is practitioner usage; I found no paper that fixes it.
Continue the refund example with illustrative units. Suppose all 204 executions, the four errored ones included, cost 5,000 units. Cost per task, counted per scored trial, is 5,000 ÷ 200 = 25 units, and cost per successful task is 5,000 ÷ 180 = 27.8.
A cheaper configuration that spends 3,000 units and passes 100 of 200 trials costs 3,000 ÷ 100 = 30 per success. The cheaper run is the dearer result, though on this few successes a gap of 27.8 against 30 is inside the noise.
Kapoor and colleagues (2024) make the research case, that “agent evaluations must be cost-controlled,” and Chapter 19 of the book has the same rule for any cost change: “The pair of numbers is the deliverable; either alone is a story.”
For the numerator, the agent cost-per-task estimator models tokens per run from the step count; it does not compute this ratio. How that bill builds up, step by step, is the larger question of what an AI agent costs to run.
Should you combine everything into one score?
Prefer the panel, and if one number is demanded, make it a pass-all rate with safety kept outside it. The disagreement among practitioners is narrower than it looks.
The case for one number is organizational. Husain and Shankar concede that “people in your organization may want a single number to track,” and offer the pass-all rate: “An example passes only if it passes every check.” Research frameworks publish aggregates too: Rabanser and colleagues (2026) compute an overall score that “uses a uniform average across pillars by default.”
The case against is that a composite moves without saying what moved. The same FAQ discourages “complicated weighted scores,” and Chapter 16 makes the point about graders: “several small graders, each owning one dimension (accuracy, format, safety), tell you what broke; one monolithic ‘was it good?’ only tells you that.”
Both sides agree on safety. Those authors “explicitly exclude safety from the overall aggregate because safety violations are inherently a tail phenomenon,” and their reason is that averaging safety with the other dimensions “obscures critical tail risks.” My position follows from that: four numbers, a gate at zero on the third, and a pass-all rate over the blocking checks if a single tile is required.
Which agent metrics are vanity metrics?
A vanity metric is a number that can improve while the product gets worse. Chapter 15 defines the plainest kind as “totals that grow impressively on a slide (requests served, tokens processed) while saying nothing about whether anyone was served well,” and gives a test: “The tile you glance at every morning should be capable of ruining your morning; if it can only ever go up, it is decoration.”
Totals that only grow (the book names the first two; the third is this post’s addition):
- Requests served rises with traffic, whatever the answers were like.
- Tokens processed rises when the agent loops, which is the opposite of improvement.
- Runs completed rises when the agent stops early and says it is done.
Flattering averages, a different fault and this post’s list:
- Average judge score on a five-point scale can rise as outputs drift toward the judge’s taste, which is my reading; Chapter 16’s own complaint is that absolute scores “drift and bunch.”
- Mean latency can fall while the slowest tenth of users wait longer.
- Deflection or containment rate rises when customers give up. One vendor’s caution about its own category (2026) defines deflection rate as “contacts handled without a human, regardless of outcome.”
What do these metrics not tell you?
None of the AI agent evaluation metrics on this page tells you whether the task set resembles your traffic, or what failure you have not named yet. The sources are also thin in places, and I would rather say where.
Coverage of reality. Every rate here is a rate on a task set, which is why a golden dataset built from real traces matters more than any formula above. An offline eval set also holds still while usage moves.
The unnamed failure. Chapter 15 says it best: “metrics tell you that quality dipped, and only the trace shows you the new failure you had no name for yet.” For an agent that only writes documents, a model judge sits under most of the panel.
Thin sources. Human intervention rate has no agreed unit, and cost per successful task has one preprint behind its per-problem form and none behind the pooled one. The four-number panel and the six families are this post’s synthesis, untested beyond the two illustrative agents above.
Dated evidence. Every percentage quoted from a study describes particular agents on particular benchmarks in the stated year. Several list no venue on their arXiv record, and the newest of those, the corrupt-success study, is unreplicated; another source is one lab’s report on its own model.
The one thing to keep
When someone asks how good the agent is, answer with four numbers and what each one hides: task success rate with its trials and its interval across tasks, passk with its k, the policy-violation count with the number of clean runs behind a zero, and cost per successful task. Then say which of the four moved since last time, and open the diagnostic rows only for that one.
Whichever AI agent evaluation metrics you add after that, give each a gate line in its definition. The book’s own caution applies to all of them: “a green suite is evidence, never proof.”
The full treatment, from multi-trial statistics to judges and regression gates, is in Chapter 16, “Evaluating Agents”, in the full book. The free agent evaluation guide collects the related posts and tools, and you can see the formats.
Questions readers ask
- What is pass@k?
- pass@k is the probability that at least one of k attempts at a task succeeds. Chen and colleagues (2021) estimate it by generating n ≥ k samples per task, counting the c that pass, and averaging 1 − C(n−c, k)/C(n, k) over tasks. It fits products where a test or a person picks the successful attempt.
- Which AI agent metrics should I start with?
- Four, in this post's view: task success rate graded on the end state, with trials per task and an interval computed across tasks; pass^k, with its k, for an agent that acts unattended; a count of policy violations with a gate at zero and the number of runs behind it; and cost per successful task, with step count and tail latency recorded on the same task set.
- Why does my eval say 90 percent while users complain?
- Check five things: whether the grader reads the final message instead of the system's state; whether the 90 is from one run per task while the product needs every run to work; whether the task set matches live traffic; whether errored runs left the denominator; and whether the complaints come from a tail that a mean hides.
- Should I combine agent metrics into one score?
- Prefer a small panel. If one number is required, use a pass-all rate, the share of examples that pass every binary check, so that anyone can drill into which check failed. Keep safety out of any average and report it as a count with its own gate.
- How do I report an agent metric when runs vary?
- State the trials per task, how trials are aggregated (mean, pass^k or pass@k with k), and an interval computed across tasks, not across pooled trials. Report how many runs were scored, errored and unscored, and keep infrastructure failures under their own status instead of counting them as agent failures.
Sources
- Anthropic (Grace, Hadfield, Olivares, De Jonghe) (2026). Demystifying evals for AI agents
- Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Mark Chen et al. (2021). Evaluating Large Language Models Trained on Code
- Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, Percy Liang (2019). SPoC: Search-based Pseudocode to Code
- Ido Levy, Ben Wiesel, Sami Marreed, Alon Oved, Avi Yaeli, Nir Mashkif, Segev Shlomov (2024). ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- Hongliu Cao, Ilias Driouich, Eoin Thomas (2026). Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation (preprint)
- Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, et al. (2025). Establishing Best Practices for Building Rigorous Agentic Benchmarks
- Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan (2026). Towards a Science of AI Agent Reliability
- Chang Ma et al. (2024). AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- Frank F. Xu et al. (2024). TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, Florian Tramèr (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, Arvind Narayanan (2024). AI Agents That Matter
- Mehmet Hamza Erol, Batu El, Mirac Suzgun, Mert Yuksekgonul, James Zou (2025). Cost-of-Pass: An Economic Framework for Evaluating Language Models (preprint)
- Evan Miller (2024). Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Anthropic (Gian Segato) (2026). Quantifying infrastructure noise in agentic coding evals
- Anthropic (2026). Measuring AI agent autonomy in practice
- Hamel Husain, Shreya Shankar (2026). AI Evals: Everything You Need to Know (FAQ)
- Voiceflow (2026). What Is Ticket Deflection? Formula, Benchmarks, and What It Hides (a vendor blog; one example of a vendor's caution about its own category)
- Ragas (2026). Agentic or tool use metrics (documentation; one example of an open-source evaluation library)
- UKGovernmentBEIS/inspect_ai issue tracker (2026). Issue #4481: Headline metrics should surface sample coverage (reported July 2026)
- langfuse discussions (2026). Discussion #16383: Usage semantics: should traces count operations or attempts? (August 2026)
- r/LLMDevs (2026). Three eval metrics that actually flag LLM prod failures, plus two that quietly miss them (forum post by an account named after an evaluation vendor, read through an archive copy)