What does this eval sample size calculator answer?
This eval sample size calculator answers the three questions an eval owner asks: how sure a pass rate is, how many tasks a target margin needs, and whether version B really beats version A. The number of tasks you need depends on the question you are asking. To find the large defects of a young agent, twenty to fifty hand-checked tasks are enough. To tell whether a change moved the pass rate by a few points, you need hundreds of independent trials, because a pass rate measured on fifty tasks carries a margin of roughly ten points either way.
The book is blunt about the first number. In Chapter 16 it says that “twenty to fifty tasks you have hand-checked and believe in beat a few hundred synthetic ones nobody has read,” and that the twenty are affordable because “early in a project the defects are large and a small net catches them.” This calculator is for the moment after that, when the defects are smaller than the noise and you need to know which is which. The longer argument, with the book’s reasoning about what an eval set is for, is in the article How many eval examples do you need?
What does a pass rate of 41 out of 50 actually tell you?
A pass rate of 41 out of 50 tells you that the agent’s true rate on tasks like these is probably between 69 and 90 percent. That is the 95 percent Wilson interval, the default in the first tab. The observed 82 percent is the most likely single value, but the data are compatible with an agent that fails one task in three and with one that fails one in ten.
This is why the book warns that “a two-point shift on a fifty-task suite is well inside the noise.” The arithmetic behind the sentence is the width of that interval. Two versions of an agent that score 80 and 82 percent on fifty tasks have intervals that overlap almost entirely, and no honest reading of the two numbers can rank them.
The calculator reports two intervals:
| Method | What it guarantees | When to use it |
|---|---|---|
| Wilson score | Coverage close to the stated confidence, sensible near 0% and 100% | Everyday reporting of eval results |
| Clopper–Pearson | Coverage never below the stated confidence | When an interval that is too narrow would be costly |
The naive interval taught in introductory courses, the observed rate plus or minus 1.96 standard errors, is not offered. Brown, Cai and DasGupta showed in 2001 that its coverage is erratic even at moderate sample sizes and collapses near the extremes, which is exactly where a good agent’s pass rate lives.
How many tasks do you need for a given margin?
The number of tasks you need grows with the square of the precision you want. At 95 percent confidence and a pass rate near the middle, a margin of 10 points takes about 97 tasks, 5 points takes 385, and 3 points takes 1,068. Halving the margin costs four times the tasks, and the second tab computes the figure for your own margin and expected rate.
Two details change the answer. A pass rate far from 50 percent needs fewer tasks for the same margin, because the variance of a proportion is largest in the middle; if your agent passes around 90 percent, you can enter that and the requirement drops. And the formula assumes independent trials. Running one task twenty times is not the same evidence as running twenty different tasks once, because the runs of one task share whatever makes that task easy or hard.
That second detail is why the book separates two practices. Repeated runs answer “how steady is the agent on this task?”, which is the reliability envelope of Chapter 16 and the subject of the pass@k and passk statistics. More tasks answer “how well does it generalize?” The book recommends both, and it also recommends starting with three runs read by hand before any infrastructure: “give the agent the same task three times, read all three outputs end to end, and ask of each one only ‘would I accept this?’”
Is version B really better than version A?
Version B is really better than version A only when the gap between them is larger than the noise in both measurements. The third tab runs a two-proportion test on the two pass rates and reports how likely a gap this size would be if the versions were equally good. When the gap is not convincing, it reports how many tasks per version you would need to detect a gap that size reliably.
With A at 38 of 50 and B at 44 of 50, the twelve-point gap looks decisive and is not: a gap this large would appear about one time in eight between two identical agents. Detecting a twelve-point difference with 80 percent power at the usual 5 percent false-alarm rate takes about 160 tasks per version.
Two practices make the comparison cheaper and more honest:
- Pair the runs. Run both versions on the same tasks and look only at the tasks where they disagree. The paired comparison removes the variation between tasks, which is usually the largest source of noise, and needs fewer tasks than the unpaired estimate shown here.
- Track more than the pass rate. The book’s rule for regression testing is that “every change to a prompt, a tool definition, or the model reruns the suite, and the numbers decide,” with numbers in the plural: cost, latency and step count beside the pass rate, because “an agent two points more accurate and three times slower is not obviously an upgrade.”
How much certainty should you buy?
You should buy as much certainty as the cost of being wrong justifies, and no more. The book states this without hedging: “Multi-trial evaluation multiplies your compute bill by the trial count,” model grading “adds a second bill on top,” and “there is a scale below which full statistical rigor costs more than the occasional production failure it would prevent.”
A practical reading for most teams:
- Early prototype. Twenty to fifty hand-checked tasks, three runs each, read by a person. The interval will be wide; that is fine, because the defects you are hunting are wider.
- Before a launch. Enough tasks that the margin on the pass rate is smaller than the drop you would refuse to ship. If a five-point regression matters, plan for a few hundred trials across distinct tasks.
- Comparing two versions. Size the suite for the smallest difference that would change your decision, run the versions on the same tasks, and keep a smoke subset for every commit with the full battery nightly, as Chapter 16 recommends.
The intervals here measure sampling noise only. They cannot see a grader that is wrong. If your pass or fail verdicts come from a model acting as judge, the judge’s disagreement with humans is a second source of error that no sample size removes, and the book’s advice is to calibrate the judge on “a few dozen outputs at minimum, drawn from real traffic, including some truly bad ones” before trusting any interval built on its verdicts.
Where do the formulas come from?
The formulas come from standard statistics, not from the book. The Wilson interval dates from 1927 and the Clopper–Pearson exact interval from 1934; both are computed here directly, the second through the inverse of the regularized incomplete beta function. The sample-size figure uses the normal approximation for a proportion, and the comparison uses the pooled two-proportion z-test with the standard power formula.
Evan Miller’s 2024 paper on error bars for language-model evals makes the same case for model benchmarks that this page makes for agents: report intervals, cluster correlated questions, and pair comparisons when the same items are used. What the book adds is the engineering frame around the arithmetic: which questions an eval set must answer, why every task needs a reference solution, and when the cost of rigor stops paying for itself. Those chapters are where an eval set becomes a regression gate.
Questions readers ask
- The book says twenty to fifty tasks. Why does the calculator ask for hundreds?
- They answer different questions. Twenty to fifty hand-checked tasks catch the large defects of an early agent, which is the book's point. Measuring a small difference between two versions to within a few points is a statistics question, and the arithmetic of a binomial proportion needs hundreds of independent trials for that.
- Should I run each task several times or add more tasks?
- Both, for different reasons. Repeated runs of the same task measure how much the agent varies on it; more tasks measure how well it generalizes. The intervals here assume independent trials, so treat repeated runs of one task as weaker evidence than the same number of different tasks.
- Which interval should I report, Wilson or Clopper–Pearson?
- Wilson for everyday reporting: it stays sensible near 0 and 100 percent and is close to the nominal coverage. Clopper–Pearson when you need a guarantee that the interval is never too narrow, at the cost of being wider.
- Does this work for scores graded by an LLM judge?
- Yes, if the judge's verdict is pass or fail, but the interval only covers sampling noise. A judge that disagrees with humans adds error the interval cannot see, so calibrate the judge first with the judge-agreement checks the book describes.
Sources
- Edwin B. Wilson (1927). Probable inference, the law of succession, and statistical inference
- C. J. Clopper and E. S. Pearson (1934). The use of confidence or fiducial limits illustrated in the case of the binomial
- Lawrence D. Brown, T. Tony Cai and Anirban DasGupta (2001). Interval estimation for a binomial proportion
- Evan Miller (2024). Adding error bars to evals: a statistical approach to language model evaluations
- Anthropic (2026). Demystifying evals for AI agents