Home / Tools / pass@k and pass^k calculator

Free tool · runs in your browser · from Chapter 16

pass@k and pass^k calculator

Compute pass@k and pass^k for an AI agent from its per-attempt success rate or your own runs, and see the reliability envelope. Free pass@k calculator.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The definitions and the reliability envelope come from Chapter 16; the closed forms assume independent attempts, and the estimators are the standard ones from the code-generation literature.

What does a pass@k calculator compute?

A pass@k calculator computes the probability that at least one of k attempts at a task succeeds, given how often a single attempt succeeds. This one also computes passk, the probability that all k attempts succeed. The two numbers describe the same agent from opposite ends: pass@k is the optimistic bound and passk the pessimistic one, and the gap between them is what the book calls the reliability envelope.

Chapter 16 calls the contrast between the two “the single most useful idea in this section.” Its example is a single agent that succeeds 90 percent of the time per attempt. Across ten attempts, at least one success “is all but guaranteed,” while all ten succeeding “happens only about one time in three.” The same agent is, in the book’s words, “a pass@k triumph and a passk catastrophe, and both descriptions are true.”

How are pass@k and passk calculated?

pass@k and passk are calculated from the per-attempt success rate p. If attempts are independent, the chance that all k fail is (1 − p) raised to k, so pass@k is one minus that. The chance that all k succeed is p raised to k. At k = 1 both equal p; as k grows they pull apart.

Per-attempt p pass@5 pass5 pass@10 pass10
99% 100% 95.1% 100% 90.4%
90% 100% 59.0% 100% 34.9%
70% 99.8% 16.8% 100% 2.8%
50% 96.9% 3.1% 99.9% 0.1%

The calculator draws both curves across attempts and shades the band between them. The teal curve is pass@k climbing toward certainty; the ochre one is passk decaying toward zero.

One agent, two honest summaries of the same ten attempts, drawn here for an illustrative per-attempt success rate of 90 percent.
Figure 16.1 One agent, two honest summaries of the same ten attempts, drawn here for an illustrative per-attempt success rate of 90 percent. The chance that at least one of k attempts succeeds (pass@k) climbs toward certainty; the chance that all k succeed (passk) decays toward zero. The widening gap between them is the reliability envelope: the wider it is, the more capable-but-erratic the agent—and a demo only ever shows you the top curve. Reuse this diagram

When you have real runs instead of a guessed p, open “Use your own runs instead” and enter how many attempts you ran (n) and how many passed (c). For any k up to n, the tool switches to the unbiased estimators: pass@k = 1 − C(n − c, k) / C(n, k), introduced by Chen and colleagues in 2021 for code-generation benchmarks, and its counterpart passk = C(c, k) / C(n, k). Plugging c/n into the closed form instead overstates pass@k on small samples, which is why the estimator exists.

Which number should you report?

You should report the number that matches how your product uses the agent. The book’s rule is to “ask which bound your product actually lives on before you quote any number at all.” The two cases are rarely ambiguous once you ask the question:

  • pass@k is “the right number when one good answer is all you need because a cheap verifier will pick it out of the pile”: code you will run against tests, or a draft a person will skim and choose from several.
  • passk is “the right number for an agent that must be correct every time it acts unattended”: a support agent answering customers, a workflow that files changes, anything where each attempt has consequences.

Quoting pass@k for an unattended agent is the most common way an evaluation flatters a system. The demo shows one good run; production experiences all of them. Chapter 16 states the distinction in two sentences: “The demo tells you what the agent can do. Shipping asks what it always does.”

What does the reliability envelope tell you?

The reliability envelope tells you how steady the agent is. “A narrow envelope is a steady agent. A wide one is a talented gambler: capable of the task, not to be trusted with it.” The calculator reports the width of the envelope at your k and reads it back: over 50 points is a gambler, under 20 a steady agent, and anything between needs the question above answered before either number is quoted.

The envelope is also the formal version of a ritual the book recommends before any tooling: “give the agent the same task three times, read all three outputs end to end, and ask of each one only ‘would I accept this?’” Three similar, acceptable outputs suggest a narrow envelope. Three wildly different ones mean, as the chapter puts it, that “you have met your distribution personally.”

Widening k shows something the per-attempt number hides. An agent at 90 percent is a fine assistant if a person reviews every output, and a liability if it runs ten actions in a row unsupervised. The per-attempt rate did not change between those two products; the bound did.

What does a worked example with real runs look like?

A worked example with real runs shows why the estimator matters. Suppose you ran one task twenty times and seventeen runs passed. The naive per-attempt rate is 85 percent, and the closed form gives pass5 = 0.85⁵ ≈ 44 percent. The unbiased estimator, which asks how often five runs drawn from your twenty would all have passed, gives C(17, 5) / C(20, 5) ≈ 40 percent.

Four points is not a rounding error when the question is whether to let an agent act five times without review. The gap grows as k approaches n and as the sample shrinks, which is exactly the regime of a small team’s first eval suite. Enter n = 20 and c = 17 in the tool to see both numbers, then raise k toward 20 and watch passk fall to zero: twenty runs drawn from twenty, three of which failed, can never all pass.

Why do these numbers depend on independence?

These numbers depend on independence because the closed forms multiply probabilities, which is only valid when one attempt’s outcome says nothing about the next. Attempts at the same task are often correlated: a task with an ambiguous instruction tends to fail every time, and an easy one tends to pass every time. Correlation pulls pass@k down and passk up, narrowing the envelope.

That is why the tool’s estimators from observed runs are the better input once you have them. They make no independence assumption across tasks; they only count what happened. The same caution applies to the arithmetic of long chains, which the compounding error calculator shows from the other direction: steps of one run fail together more often than independence predicts, and that makes long runs worse, not better.

How many runs do you need for these numbers to mean something?

You need more runs than most teams start with. A per-attempt rate estimated from five runs of one task can be off by twenty points or more, and passk inherits that error raised to the power k. The book recommends running each task “a handful at minimum, a couple of dozen when the stakes justify it,” and comparing distributions rather than single runs.

Two habits make the numbers trustworthy. Run every task in your eval set several times, so each task gets its own c and n. And size the set itself so the margin on the pass rate is smaller than the difference you care about; the eval sample-size calculator computes that margin for your numbers.

Questions readers ask

What is the difference between pass@k and pass^k?
pass@k is the probability that at least one of k attempts succeeds; pass^k is the probability that all k succeed. At k = 1 they are the same number, and as k grows pass@k climbs toward 100 percent while pass^k falls toward zero.
Which one should I report?
Report pass@k when a cheap check will pick a good answer out of several, such as code that runs against tests. Report pass^k when the agent acts unattended and has to be right every time it acts.
Why does the calculator use a different formula for my own runs?
With observed runs, plugging c/n into the closed form is biased. The estimator 1 − C(n−c, k)/C(n, k), introduced for code-generation benchmarks, gives an unbiased pass@k from n samples, and C(c, k)/C(n, k) does the same for pass^k.
Are real attempts independent?
Usually not. Attempts on the same task share whatever makes it easy or hard, so the envelope in practice is often narrower than the independent case suggests. Treat the closed forms as the clearest picture of the idea, not a forecast.

Sources

  1. Chen et al. (2021). Evaluating Large Language Models Trained on Code
  2. Yao et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
  3. Anthropic (2026). Demystifying evals for AI agents