pass@k vs passk is a choice between two questions asked of the same repeated runs. pass@k is the chance that at least one of k attempts at a task succeeds: can the agent ever do it? passk is the chance that all k succeed: does it always? The gap between them is the agent’s reliability envelope.
I wrote this for the engineer who has just run every task in a suite several times and watched one success rate split into two numbers that disagree by forty points. This post says which number your product lives on, how to compute both without the two shortcuts that bend them, and how many trials a reported k needs behind it. Every worked number is illustrative and comes from one small suite, computed for this post.
pass@k vs passk: what is the difference?
The difference is the word between the attempts: pass@k joins them with “or” and passk joins them with “and.” Take one task and run it k times. pass@k is the probability that at least one run succeeds. passk is the probability that every run does. At k = 1 they are the same number, the plain success rate.
Chapter 16 of the book, in the section “Why Evaluation Is Hard,” calls pass@k “the optimistic bound” and passk “the pessimistic bound,” and names the gap between them the reliability envelope. Its reading of the gap is the one this post builds on: “A narrow envelope is a steady agent. A wide one is a talented gambler: capable of the task, not to be trusted with it.”
The figure is drawn for one task with one success rate, and that is the case the closed forms cover. If a task succeeds with probability p on each independent attempt, pass@k = 1 − (1 − p)k and passk = pk. A real suite has many tasks with different rates, and most of this post is about what that does to the picture.
Where do the two numbers come from?
pass@k comes from code generation and passk from agent evaluation, five years apart. Kulal and colleagues (2019) reported “success rate at B: the fraction of test examples where the system generates an accepted program under the budget of B trials.” Chen and colleagues (2021) called it pass@k and gave the estimator in use today.
Yao and colleagues (2024) proposed the second number for “real-world agent tasks requiring reliability and consistency like customer service.” They define passk as “the chance that all k i.i.d. task trials are successful, averaged across tasks.” The last three words of that definition matter more than they look, and the pooling section below returns to them.
Which number does your product live on?
Your product lives on pass@k if something picks the winning attempt, and on passk if every attempt acts and nobody checks. The book’s instruction is to “ask which bound your product actually lives on before you quote any number at all,” and one question settles it: who or what looks at an attempt before it takes effect?
The requirement for a picker is in the source. Chen’s paper says pass@k “can also be interpreted as the result of evaluating the best out of k samples, where the best sample is picked by an oracle with prior knowledge of the unit tests.” An oracle is a check outside the model that knows the right answer. With no such check in the product, pass@k describes a system you have not built.
One lab’s guide to agent evals (2026) gives the short form of the rule: “pass@k for tools where one success matters, passk for agents where consistency is essential” (Anthropic). The table below is this post’s longer form of the pass@k vs passk choice. It adds two situations that the two-line rule skips: a person who reviews every output, and a batch of different jobs. The buttons filter by who checks.
| Your situation | The question you are asking | Number to quote | State beside it | Who checks the output |
|---|---|---|---|---|
| The agent produces k candidates for one task and an automated test selects one that passes | Is a passing answer among k tries? | pass@k, with k = the candidates the product really generates | Trials per task; that the product’s test is as strict as the eval’s grader | a check picks |
| The agent retries when a failure is detected, up to k attempts | Does it get there within the retry cap? | pass@k, with k = the cap | What detects the failure and how often that detector is wrong; the cost of k attempts | a check picks |
| The agent shows k drafts and a person chooses one | Is a good draft in the pile? | pass@k, as a ceiling | How often the person picks the good draft is a separate measurement | a person picks |
| One attempt per request, and a person reads each output before it takes effect | How often is the first attempt right? | Success rate (pass@1, which equals pass1) | Trials per task; an interval across tasks | a person reviews each |
| One attempt per request, acting with no review, and the same kind of request keeps arriving | Does the same request get the right outcome every time? | passk per task, with k fixed before the run | k, trials per task, an interval across tasks | nobody checks |
| A batch of k different jobs, one attempt each, that must all be right | Do all k jobs come out right? | Mean success rate raised to the k, labeled a batch estimate | That jobs are assumed independent draws from the task mix; this is not passk | nobody checks |
| Comparing two versions of the agent | Did capability or consistency move? | The row above that matches the product, computed the same way on both sides | Same tasks, same k, same trials, same estimator, both intervals | a check picks, a person picks, a person reviews each, nobody checks |
What about a batch of different jobs?
A batch of k different jobs that must all succeed is a third question, and per-task passk does not answer it. passk asks whether the same task holds up under repetition. A nightly run over fifty different invoices asks whether fifty separate tasks each pass once.
If the jobs are independent draws from the mix of tasks your suite represents, the chance that all k are right is the mean success rate raised to the k. That is the arithmetic of a chain, which why agent errors compound works through. Chapter 16’s own example, “across ten unattended actions,” uses one shared rate of 90 percent, where the two questions give the same answer. They separate as soon as tasks differ, so say which one you computed.
What does the envelope tell you that one success rate hides?
The envelope tells you how the failures are distributed: whether the agent fails a little on every task or completely on a few. A single success rate cannot show this. Chapter 16 puts the complaint in one sentence: “a single end-to-end success rate is an average, and an average is exactly the wrong summary for a system whose product is its worst runs.”
Here is pass@k vs passk on a worked example with illustrative numbers. A suite has 20 tasks, each run 10 times. Ten tasks pass all 10 runs, four pass 7, three pass 5, two pass 2, and one passes none. That is 147 passes in 200 runs, a success rate of 73.5 percent.
| k | pass@k | passk | Envelope (points) |
|---|---|---|---|
| 1 | 73.5% | 73.5% | 0.0 |
| 3 | 88.9% | 57.1% | 31.8 |
| 5 | 92.7% | 51.7% | 41.0 |
| 10 | 95.0% | 50.0% | 45.0 |
Each cell is a mean over the 20 tasks of the per-task estimators described two sections down. The pair at k = 5 says something the 73.5 percent did not. With five tries and a picker, this agent delivers on about 93 percent of tasks. Left alone for five repeats of a task, it gets all five right on about 52 percent.
How do you read the envelope at k = n?
At k equal to the number of trials, the envelope is a head count, and it sorts the suite into three piles. pass10 is the share of tasks that passed every run: 10 of 20. One minus pass@10 is the share that never passed: 1 of 20. The envelope, 45 points, is the share that did both: 9 of 20 tasks passed at least once and failed at least once.
The three piles call for different work, which is my reading and not a source’s. The never-pass pile is a capability gap or a broken task, and more retries will not help it. The middle pile is where the agent can do the job and sometimes doesn’t, so retries with a picker will help and unattended use will hurt. Rabanser and colleagues (2026) describe that pile at research scale: “agents that can solve a task often fail to do so consistently.”
One caution about the always-pass pile: ten clean runs prove less than they seem to. Ten passes in ten runs rule out, at one-sided 95 percent confidence, only failure rates above 25.9 percent. A task whose true success rate is 90 percent goes ten for ten about one time in three (34.9 percent).
Why can’t you pool the runs and raise the rate to the k?
You can’t, because tasks differ in difficulty, and averaging first and raising to the k second gives a different number from doing it in the other order. The definition already says which order is right: “averaged across tasks.” The same lab guide quoted above makes the premise explicit: “Each task has its own success rate—maybe 90% on one task, 50% on another.”
On the worked suite, the pooled shortcut takes 73.5 percent and computes 0.7355 = 21.5 percent for pass5. The per-task figure is 51.7 percent. On the other side the shortcut gives 1 − (1 − 0.735)5 = 99.9 percent for pass@5, against a per-task 92.7 percent.
| At k = 5 | pass@5 | pass5 | Envelope (points) |
|---|---|---|---|
| Per-task estimators, averaged over tasks | 92.7% | 51.7% | 41.0 |
| Pooled rate in the closed forms | 99.9% | 21.5% | 78.4 |
Both errors push outward. For tasks whose true rates you know, the pooled picture is the widest envelope a given mean rate allows. One task’s estimate from a few runs can land outside it: the 7-of-10 task below shows 92 points where the closed forms at 70 percent give 83. The pooled picture is the picture of a suite where every task is equally flaky. The narrowest envelope for the same 73.5 percent is zero: a suite where each task either always passes or always fails. A real suite sits between, and where it sits is the information.
A published result shows the same gap. In the paper that introduced passk, one frontier model of 2024 scored “∼61%” on a single trial in one domain and fell “as low as ∼25% for pass8” there. If every task shared that 61 percent rate, pass8 would be 0.618, about 1.9 percent. That comparison is this post’s arithmetic; the paper does not make it.
Twenty-five percent is what a suite looks like when a block of tasks nearly always passes.
Why not plug each task’s pass rate into the formula?
Because an observed rate from a few trials is noisy, and raising a noisy number to the k does not average out. Per task, the unbiased estimators count subsets instead. If c of n trials passed, passk is C(c, k) ÷ C(n, k): of all the ways to draw k of your n runs, the share in which every drawn run passed. pass@k is 1 − C(n − c, k) ÷ C(n, k).
Both formulas are in Yao’s paper, and the second is Chen’s. Chen’s paper warns against the shortcut directly: “One may be tempted to estimate pass@k with 1−(1−p̂)k where p̂ is the empirical estimate of pass@1, but we show that it is biased.” Its Appendix A gives the direction, “a consistent underestimate,” and adds: “The gap doesn’t fully close even when n > 5k.”
The bias runs the other way for passk. The figures below are exact expectations over a binomial count, computed for this post, at k = 5.
| True rate of one task | True value | Shortcut’s average at n = 5 | At n = 10 | At n = 20 |
|---|---|---|---|---|
| 70%, pass5 | 16.8% | 31.1% | 24.0% | 20.4% |
| 10%, pass@5 | 41.0% | 29.6% | 34.9% | 37.8% |
The shortcut overstates passk and understates pass@k, and the combinatorial estimators average to the true value in every one of these cases. One of the suite’s 7-of-10 tasks shows it on a single count. The calculator below opens on that task: it prints a pass@5 of 100.0 percent, a pass5 of 8.3 percent, and an envelope 92 points wide. The shortcut would have said 0.75 = 16.8 percent, twice as much.
With JavaScript on, the pass@k and passk calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the pass@k and passk calculator on its own page to share a result by link.
The pass@k calculator handles one task at a time: the sliders use the closed forms for a rate you assume, and “Use your own runs instead” switches to the estimators for any k up to n. For a suite, compute each task and average.
On the worked suite the two shortcuts miss in opposite directions. Averaging per-task plug-ins gives a pass5 of 53.8 percent, two points over the estimator’s 51.7. The pooled rate gives 21.5 percent, thirty points under. Two dashboards that both print “pass5” can therefore disagree by thirty points on the same runs.
How many trials do you need before k means anything?
You need at least k trials per task, because the unbiased estimators do not exist for a k larger than n. Drawing five runs from three is not possible. Anything reported for k above the trial count is an extrapolation from an assumed rate, and it should be labeled as one.
With n = k, each task scores 0 or 1, and one task’s estimate is very coarse. For a task with a true rate of 70 percent, the standard deviation of its pass5 estimate is 37 points at n = 5, 21 points at n = 10 and 13 points at n = 20 (exact, computed for this post). Per-task values are for sorting tasks into piles. They are too rough to rank tasks against each other.
The suite number is steadier, and its main uncertainty comes from the tasks. Resampling the 20 tasks of the worked suite (a bootstrap, 20,000 resamples) puts a 95 percent interval on pass5 of 31.4 to 71.7 percent, and on pass@5 of 81.6 to 99.9 percent. Noise from the trials alone is much smaller: about 3 points of standard deviation at n = 10, for a suite with underlying rates like this one.
That comparison is the reason for this post’s rule of thumb. Take k from the product and fix it before the run. Run n ≥ k trials per task, about 2k where you can afford it. Then spend what is left on more tasks, because twenty tasks leave pass5 known only to within about twenty points either way.
The bill is tasks times trials: 200 agent runs for this suite. The book is direct about it: “Multi-trial evaluation multiplies your compute bill by the trial count.” Its guidance on the count is “a handful at minimum, a couple of dozen when the stakes justify it.” How many tasks the interval needs is the subject of how many eval examples you need, and the eval sample size calculator does that arithmetic.
How should you report the pair?
Report pass@k vs passk as a pair, with everything a second person needs to recompute it: tasks, trials per task, k, the estimator, an interval across tasks, and who picks the winner. A bare “pass5 = 52%” hides at least four choices. On this suite the estimator alone moves the number by thirty points, and k moves it by more than twenty.
suite: <name and version>, <T> tasks
trials per task: n = <n>, independent runs, fresh state each
k: <k>, chosen because <the product's reason>, fixed before the run
estimator: per task, then mean over tasks
pass@k = 1 - C(n-c, k) / C(n, k)
pass^k = C(c, k) / C(n, k)
success rate: <passes> / <runs> = <x>%
pass@k: <x>% (95% interval across tasks: <lo> to <hi>, <method>)
pass^k: <x>% (95% interval across tasks: <lo> to <hi>, <method>)
envelope at k: <pass@k minus pass^k> points
piles at k = n: always pass <a>/<T> | flaky <f>/<T> | never pass <z>/<T>
who picks: <the test or person that selects an attempt | nobody>
quote to users: <pass@k | pass^k | success rate | batch estimate>, because <who checks>
not covered: <k above n; batch of different jobs; grader error; traffic the suite lacks>
Filled in for the worked suite, with a support agent that acts unattended, it reads like this:
suite: returns-v1 (illustrative), 20 tasks
trials per task: n = 10, independent runs, fresh state each
k: 5, chosen because the same request type recurs unreviewed; fixed before the run
estimator: per task, then mean over tasks
success rate: 147 / 200 = 73.5%
pass@5: 92.7% (95% interval across tasks: 81.6 to 99.9, bootstrap over tasks)
pass^5: 51.7% (95% interval across tasks: 31.4 to 71.7, bootstrap over tasks)
envelope at 5: 41.0 points
piles at k = 10: always pass 10/20 | flaky 9/20 | never pass 1/20
who picks: nobody
quote to users: pass^5, because nobody checks
not covered: k above 10; a batch of different returns; grader error
The “quote to users” line is the decision table’s answer for this product, and the pass@5 line stays in the report anyway. Together they say the agent can do most of these tasks and does about half of them every time. That is a different engineering problem from an agent that cannot do them at all. The wider set of numbers a release decision needs is in the post on AI agent evaluation metrics.
Where does this reading break?
The pass@k vs passk reading breaks where its assumptions do: independent trials, a grader that is right, and a suite that resembles the traffic. Each one can move the envelope without the agent changing.
Trials that are not independent. The estimators assume each run of a task is a fresh draw. Runs that share a cached response, a leftover database row or a warm session agree with each other more than fresh runs would. That inflates passk and narrows the envelope.
A flaky grader. A grader that disagrees with itself turns steady tasks into flaky ones on paper, which widens the envelope. Before trusting a middle pile, rerun the grader on stored outputs. If a model does the grading, whether an LLM judge is reliable is the prior question.
A picker that is weaker than the eval’s. pass@k credits the product with the eval’s grader as its picker. If production has a weaker test, or a hurried person, the product’s real number sits below pass@k by an amount this metric cannot see.
The suite is not the traffic. Both numbers are means over the tasks you wrote. A suite that over-represents easy tasks has a large always-pass pile and a flattering passk. The work of testing an AI agent starts with which tasks go in.
Small suites. An interval of 31 to 72 percent on pass5 cannot tell a good release from a bad one. With 20 tasks, read the piles and the transcripts in the middle pile, and treat the headline pair as a rough position.
The one thing to keep
On pass@k vs passk, neither number is the honest one and neither is the flattering one. Each is a correct answer to a different question, and the mistake is quoting one where the product asks the other. Pick by who checks the output, compute per task with at least k trials, and report the pair with its interval and its three piles.
The pair also feeds decisions further out. Whether a passk is good enough to put in front of customers is a question about AI agent reliability for an enterprise, and the suite is one of several signals for knowing whether an AI agent is working once it is live. The book’s line for both is the same: “The demo tells you what the agent can do. Shipping asks what it always does.”
The full treatment, from multi-trial statistics to judges and regression gates, is in Chapter 16, “Evaluating Agents”, in the full book. The free agent evaluation guide collects the related posts and tools, and you can see the formats.
Questions readers ask
- What is the difference between pass@k and pass^k?
- pass@k is the probability that at least one of k attempts at a task succeeds; pass^k is the probability that all k succeed. They are equal at k = 1. As k grows, pass@k rises toward the share of tasks the agent can ever solve and pass^k falls toward the share it never fails.
- Which one should I report, pass@k or pass^k?
- Ask who checks the output. If a test or a person picks one attempt out of k, report pass@k with the product's real k. If the agent acts on every attempt with no review, report pass^k. If a person reviews every single output, the plain success rate with its interval is enough. Report both when you want to show how flaky the agent is.
- Is pass^k just the success rate raised to the power k?
- Only for one task whose true success rate you know. From real trials, compute C(c, k) / C(n, k) for each task, where c of n trials passed, and average over tasks. Raising a suite's pooled success rate to the k understates pass^k when tasks differ in difficulty; raising one task's observed rate to the k overstates it.
- How many trials per task do I need for pass^k?
- At least k, because the unbiased estimators are undefined for k larger than the number of trials. With exactly k trials each task scores 0 or 1. This post's rule of thumb is to take k from the product, run about twice k trials per task when you can afford it, and spend the rest of the budget on more tasks.
- How do you pronounce pass^k?
- The paper that proposed the metric, by Yao and colleagues in 2024, writes it out as "pass hat k". Other write-ups say "pass power k" or "pass to the k". The definition is the same in all of them: every one of k trials of a task succeeds.
Sources
- Mark Chen et al. (2021). Evaluating Large Language Models Trained on Code
- Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, Percy Liang (2019). SPoC: Search-based Pseudocode to Code
- Anthropic (2026). Demystifying evals for AI agents
- Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan (2026). Towards a Science of AI Agent Reliability