Home / Blog / Evaluating and observing agents / How Many Eval Examples Do I Need? Sample Sizes …

Evaluating and observing agents

How Many Eval Examples Do I Need? Sample Sizes and Intervals

How many eval examples do I need? Twenty to fifty at first, hundreds to compare versions. The arithmetic, paired tests, and a worked example to copy.

By Enrique Gutiérrez · Published · 15 min read

How many eval examples do I need? Twenty to fifty hand-checked tasks to start, because the defects of a young agent are large and a small net catches them. A few hundred independent tasks once you need to see a five- or ten-point change between versions, and around a thousand to resolve two or three points. The number follows the question the eval set is answering this month.

If you only want the arithmetic for your own numbers, the eval sample size calculator computes intervals, required sample sizes and A-versus-B tests in your browser. This post is about everything around the arithmetic: why the small number and the large number are both right, which tricks shrink the large one, why some of your tasks count for less than you think, and how an eval set actually grows over the life of a project.

How many eval examples do I need, by question?

The number of eval examples you need is set by the decision the eval feeds, so start by naming the decision. An eval set (a curated collection of tasks, each an input plus a way to grade the output) answers different questions at different stages, and each question has its own price in tasks.

The question you are asking Typical size (illustrative) What drives the number
Does the agent have large, obvious defects? 20–50 tasks, a few runs each Probability that a common defect shows up at least once
Did this change break anything that used to work? The regression suite, near 100% pass Coverage of behaviors someone cares about
Is version B better than A by 10 points? Roughly 100–250 tasks, paired Statistical power at the gap you care about
Is B better than A by 3 to 5 points? Several hundred to about 1,000 tasks Power falls with the square of the gap
Is a dangerous action rare enough to ship? Hundreds of trials with zero failures Upper bound on a rate you never observe

The book’s own answer covers the first row, and it is blunt. Chapter 16 says that “twenty to fifty tasks you have hand-checked and believe in beat a few hundred synthetic ones nobody has read, and the twenty are affordable precisely because early in a project the defects are large and a small net catches them.” The statistics pages that rank for this question mostly cover rows three and four. Both camps are right about their own row, and the mistake is applying one camp’s number to the other camp’s question.

Why are twenty to fifty tasks enough at the start?

Twenty to fifty tasks are enough at the start because the job of an early eval set is to find defects, and a defect that affects many tasks is almost certain to appear in a small random sample. Finding a problem takes far fewer tasks than measuring a pass rate precisely.

The arithmetic is one line of standard probability. If a defect breaks a fraction f of the tasks your users send, and you draw n tasks at random from that mix, the chance that at least one of them exposes it is 1 − (1 − f)ⁿ.

Defect affects 20 tasks 30 tasks 50 tasks 100 tasks
1 task in 5 99% >99% >99% >99%
1 task in 10 88% 96% >99% >99%
1 task in 20 64% 79% 92% >99%
1 task in 100 18% 26% 39% 63%

Read the table by rows. A prototype’s defects live in the top rows: the tool that fails on dates, the instruction ignored after ten steps, the refusal on any request mentioning money. Thirty tasks catch each of those nearly every time. The bottom row is the failure that hits one user in a hundred, and no small set will find it reliably; that is production monitoring’s job, which the post on monitoring AI agents in production takes up.

Anthropic’s engineering guide to agent evals reaches the same conclusion from practice: “In reality, 20-50 simple tasks drawn from real failures is a great start,” because “in early agent development, each change to the system often has a clear, noticeable impact, and this large effect size means small sample sizes suffice.” The phrase that matters is effect size. Small sets work when the effects are large, and they stop working the day the effects you are chasing become small.

What does a pass rate on fifty tasks hide?

A pass rate on fifty tasks hides a margin of roughly ten points on each side. At an observed 80 percent, the 95 percent Wilson interval on fifty independent tasks runs from about 67 to 89 percent; on twenty tasks it runs from about 58 to 92. The single number is the center of a wide band, and any two versions inside that band are indistinguishable.

The half-width shrinks with the square root of the number of tasks, so precision gets expensive fast. At 80 percent observed: about ±17 points at 20 tasks, ±11 at 50, ±8 at 100, ±5.5 at 200, ±4 at 400 and ±2.5 at 1,000. This is the arithmetic behind the book’s warning that “a two-point shift on a fifty-task suite is well inside the noise,” and behind its sharper line that a celebrated “two-point improvement” may be “noise wearing a suit.”

The interval method matters most where the sets are smallest. Bowyer, Aitchison and Ivanova argued at ICML 2025 that intervals built on the central limit theorem “perform very poorly” on small eval sets, “usually dramatically underestimating uncertainty.” In their simulations, nominal 95 percent intervals on 100 questions covered the truth only 92.5 percent of the time, and the textbook interval collapses to zero width when an agent passes everything. They recommend the Wilson score interval or a simple Bayesian interval instead, which is what the calculator uses.

When do you need hundreds of tasks?

You need hundreds of tasks when the decision depends on a difference of a few points between two versions. The number of tasks required grows with the inverse square of the gap you want to detect: halve the gap and you need four times the tasks.

Sizing for a comparison has a different shape from sizing for a margin. You pick the smallest difference that would change your decision (the minimum detectable effect), the false-alarm rate you accept (conventionally 5 percent) and the chance of catching a real difference (conventionally 80 percent, called power). For two independent samples at a baseline of 80 percent, standard power arithmetic gives about 200 tasks per version to detect a move to 90 percent, and about 900 per version to detect a move to 85.

Evan Miller’s 2024 paper on error bars for evals runs the same calculation for model benchmarks and concludes that “new evals should contain at least 1,000 questions in order to have good signaling ability.” That figure is for distinguishing models a few points apart, and it is the right target for a mature product where every remaining improvement is small. It would be the wrong target for a prototype, and nobody who has tried to hand-check a thousand agent trajectories in week two needs to be told why.

How does a paired comparison shrink the number?

A paired comparison shrinks the number by running both versions on the same tasks and testing only the tasks where they disagree. Tasks where both pass or both fail carry no information about which version is better, and removing them removes the largest source of noise in most evals: some tasks are simply harder than others.

The test is McNemar’s, from 1947, and it is short enough to do by hand. Count the tasks B fixed (A failed, B passed) and the tasks B broke (A passed, B failed). If the versions were equally good, each disagreement would be a coin flip, so you ask how surprising your split is. Miller recommends the paired analysis “wherever practicable” and calls it a “free” reduction in variance, because question scores are positively correlated across models: a task that one version finds hard, another tends to find hard too.

How much it saves depends on the discordance rate, the share of tasks where the versions disagree. For 80 percent power at a 5 percent false-alarm rate:

Gap to detect 10% of tasks disagree 20% disagree 30% disagree
10 points ~76 tasks ~155 tasks ~233 tasks
5 points ~312 tasks ~626 tasks ~940 tasks

Similar versions disagree on few tasks, which is exactly when pairing helps most. A one-paragraph prompt edit usually lives in the left column; a model swap can land in the right one. A pilot run of a few dozen tasks tells you which.

Should you run each task several times or add more tasks?

Run each task a few times, then spend the rest of the budget on more tasks. Repeated runs reduce the noise inside each task, where the agent’s randomness lives; only more tasks reduce the variation between tasks, which is usually larger and is what generalization means.

The book asks for both. Run each task “a handful at minimum, a couple of dozen when the stakes justify it; the exact count is an economics question, and any figure is illustrative.” Miller’s paper quantifies the trade in a stylized case with binary scores and task difficulty spread evenly: going from one run to two cuts the variance of the pass rate by a third, four runs cut it by half, and no number of runs can cut it by more than two thirds. In another of his examples, raising the runs per question from one to ten on a 198-question eval shrinks the minimum detectable effect from 13.2 to 7.5 points. Useful, and bounded.

Two cautions follow from the same paper. The task remains the unit: average each task’s runs into one score and compute the interval across tasks, because pooling all the runs as if they were independent overstates your precision. And resist lowering the sampling temperature to make results steadier; Miller’s section on it is titled “Don’t touch the thermostat!”, since it can move the noise somewhere the statistics cannot see. The runs also buy something the pass rate cannot show: the spread between pass@k and passk, which the pass@k calculator computes and the book calls the agent’s reliability envelope.

Why do correlated tasks count for less?

Correlated tasks count for less because the interval formulas assume every task is an independent draw, and tasks that share a source tend to pass or fail together. Thirty tasks built from one customer’s account are closer to one observation repeated thirty times than to thirty observations.

Miller’s paper borrows the fix from the social sciences: clustered standard errors, which treat each group of related questions as the independent unit. On the real benchmarks he measured, the clustered standard error was 1.1 times the naive one on one reading-comprehension set, 1.9 times on a multilingual set and 3.05 times on another, the last meaning the honest interval was three times as wide as the one usually reported. His suggested report lists the cluster count beside the question count.

Agent eval sets are clustered by construction, which makes this the point where agents differ most from textbook classification. The book’s method for growing a suite is to “mine the recent failures, name the shapes they fall into, and give each shape a task in the set.” That produces groups: several tasks from one long trace, several variants of one failure shape, several requests from one heavy user.

Survey statisticians summarize the cost with the design effect, 1 + (m − 1)ρ, where m is the average cluster size and ρ the correlation within a cluster. With seven tasks per cluster and an illustrative ρ of 0.2, the design effect is about 2.2, and four hundred tasks carry the information of fewer than two hundred independent ones. The cheapest protection is diversity: when you build a golden dataset from real traces, cap how many tasks any single source contributes.

What makes a task count at all?

A task counts only if it can be passed and its grade can be trusted. Sample size multiplies whatever each task contributes, and a task that measures nothing contributes nothing at any scale.

The book makes solvability non-negotiable: “every task must be solvable, and the way to know is to include a reference solution that proves it.” An unsolvable task (information missing, constraints contradictory, a grader expecting a file path the instructions never mention) fails every version forever. It caps your pass rate below 100 percent, hides real progress, and adds noise that no amount of extra data averages away. Writing the reference solution is also the fastest way to discover that a task is ambiguous.

The grader is the other half. If a model grades your tasks, the interval on the pass rate covers sampling noise only, and a judge that disagrees with careful humans adds error the interval cannot see. Calibrating an LLM-as-a-judge is its own sample-size problem, and the book’s warning is direct: “Kappa computed on ten examples is a random number; you need a few dozen labels before it means anything.” The arithmetic of agreement statistics and judge calibration is close to that of any classifier, which is why the guide to building an LLM classifier reuses it.

Worked example: growing an eval set over a project’s life

Here is one eval set traced from the first afternoon to a year in production. The agent is hypothetical, a support agent that reads tickets, looks up accounts and drafts replies; every number is illustrative and computed with the standard formulas above.

Week 1: three tasks, three runs each. Before any infrastructure, the book’s ritual: “give the agent the same task three times, read all three outputs end to end, and ask of each one only ‘would I accept this?’” Two of the three tasks produce wildly different outputs across runs. No statistics yet, and none needed: you have met the distribution.

Month 1: thirty tasks, three runs each, from real failures. You mine the first week of internal traffic, write a reference reply for each task, and grade by hand. The agent passes 18 of 30 tasks (per-task verdict by majority of runs). The 95 percent Wilson interval is 42 to 75 percent, which is wide and also beside the point. What mattered was the reading: three failure shapes covered half the failures, each would have shown up in almost any thirty-task sample, and each gets a fix.

Month 3: 120 tasks and the first real comparison. Two prompt versions are close, and you need to choose. Run separately, A passes 78 of 120 and B passes 88 of 120. An unpaired two-proportion test gives p ≈ 0.16, which settles nothing. Paired on the same tasks, the picture changes: B fixed 14 tasks that A failed and broke 4 that A passed. McNemar’s exact test on that 14-to-4 split gives p ≈ 0.03. Same tasks, same runs, a clear answer, because the 102 tasks where the versions agreed were removed from the noise.

Launch: the safety cases. A separate set asserts what the agent must never do, such as issue a refund above its limit. Here you never observe the event you care about, so the right tool is the rule of three, as Hanley and Lippman-Hand set it out for clinicians in 1983: with zero failures in n independent trials, the 95 percent upper bound on the true rate is about 3/n. Zero failures in 50 trials still allows a rate near 6 percent. Zero in 300 trials, spread over distinct tasks, bounds it near 1 percent. Whether 1 percent is acceptable is a product decision, and the statistics only tell you what you have earned the right to claim.

Month 12: 400 tasks, two suites, one gate. The set has split the way the book describes: capability evals the agent mostly fails, regression evals it reliably passes, and tasks promoted from the first to the second as they stabilize. A smoke subset runs on every commit and the full battery nightly, feeding a regression gate where the numbers decide whether a change merges.

Eval-driven development as plumbing.
Figure 16.4 Eval-driven development as plumbing. Every change runs against two suites: the capability evals it is meant to climb, and the regression evals it must not break; tasks that stabilize are promoted from the first into the second. The change reaches a gate (in accent) where a threshold is enforced, not glanced at—the numbers, not moods, decide whether it merges or is blocked. Reuse this diagram

The 400 tasks came from about 60 customer accounts, so they are clustered. With roughly seven tasks per account and an illustrative within-account correlation of 0.2, the effective sample is under 200 independent tasks, enough to see a ten-point paired change and too few for a three-point one. The suite tells you so, as long as you report the cluster count.

Stage Tasks Runs per task Question answered Statistic that matters
Week 1 3 3 Is the agent steady at all? None; read the outputs
Month 1 30 3 What are the big defects? Chance a defect appears
Month 3 120 3 Is B better than A? Paired test on disagreements
Launch +300 safety trials 1+ Is the bad event rare enough? Rule of three
Month 12 400 (≈60 clusters) 3 nightly Did anything regress? Clustered, paired intervals

None of these stages is the “right” size. Each one is the size that answered the question of its month, and the set grew because the questions got harder. For the testing side of the same story, including how to write assertions for outputs that differ on every run, see the guide on how to test AI agents.

How much certainty is worth paying for?

Certainty is worth paying for up to the cost of being wrong and no further: in the book’s words, “buy as much certainty as the cost of being wrong justifies, and no more.” The eval sample size calculator walks through what that certainty costs, stage by stage.

In practice that becomes a short rule. Write down the smallest change that would alter your decision before you size anything; if a five-point drop would stop a release, size the suite to see five points, and if nothing short of fifteen would, a smaller suite is the honest choice. When the required size is unaffordable, say so in the report (“this suite can detect ten-point changes, not three”) instead of presenting the number as if it were sharper than it is. An underpowered eval that admits it is far more useful than one that reports to a decimal place.

What are the limits of this arithmetic?

The limits are that every interval here measures sampling noise on the tasks you chose, and nothing else. It cannot see a grader that is wrong, a task mix that no longer matches what users send, or a defect that none of your tasks exercise.

The defect table above assumes your tasks are drawn from the same mix as real traffic; a suite written from imagination fails that assumption silently. The power figures assume the effect you sized for is the one that matters. And a green suite, as Chapter 16 says, “is evidence, never proof.” The arithmetic tells you how much to trust a number. Whether the number measures the right thing is decided by reading traces, which no formula does for you.

The one thing to keep

How many eval examples do I need is the wrong first question; the right one is what decision the eval must support this month. Start with twenty to fifty solvable, hand-checked tasks and read the failures. Move to hundreds, paired and diversified across sources, when the differences you care about shrink below ten points. Use the eval sample size calculator to put an interval on every number you report.

The full argument, from the three-run ritual to judges and regression gates, is in Chapter 16, “Evaluating Agents”, in the full book. Chapters 0 to 2 are free to read; see the formats.

Questions readers ask

How many eval examples do I need to start?
Twenty to fifty tasks you have checked by hand, each with a reference solution, drawn from real failures where you have them. Early defects are large, so a small set catches them. Run each task a few times and read the outputs.
How many examples do I need to compare two prompts or models?
Decide the smallest difference that would change your decision first. Detecting a ten-point gap reliably takes roughly one to two hundred tasks with a paired design; a five-point gap takes several hundred; two or three points takes around a thousand.
Is it better to run each task many times or to add more tasks?
More tasks, once you have a few runs per task. Repeated runs only reduce the noise inside each task; the variation between tasks shrinks only when you add tasks. A few runs per task capture most of the benefit of resampling.
Can I claim my agent never does something dangerous if it passed every safety case?
Only up to a bound. With zero failures in n independent trials, the 95% upper limit on the failure rate is about 3/n, the rule of three. Zero failures in fifty trials still allows a rate near 6%; in three hundred trials, about 1%.
What confidence interval should I report for an LLM eval?
A Wilson score interval or a Bayesian Beta-Bernoulli interval for a pass rate, clustered when tasks come in related groups, and a paired interval for a difference between versions. The textbook normal-approximation interval is too narrow on small eval sets.

Sources

  1. Evan Miller (2024). Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
  2. Sam Bowyer, Laurence Aitchison, Desi R. Ivanova (2025). Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
  3. Anthropic (2026). Demystifying evals for AI agents
  4. James A. Hanley, Abby Lippman-Hand (1983). If nothing goes wrong, is everything all right? Interpreting zero numerators
  5. Quinn McNemar (1947). Note on the sampling error of the difference between correlated proportions or percentages