Home / Tools / Agent cost-per-task estimator

Free tool · runs in your browser · from Chapter 19

Agent cost-per-task estimator

Estimate what one agent run costs in tokens: the fixed prompt, the history it re-reads every step, retries and subagents. Free AI agent cost estimator.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The formula and the napkin numbers come from Chapter 19; the worker sum follows the chapter's description of subagents.

What does an AI agent cost per task?

An AI agent’s cost per task is the fixed prompt it sends on every step, multiplied by the number of steps, plus the growing history it re-reads on every step. That second term rises with the square of the run’s length, so a ten-step run costs more than twice what most engineers estimate, and a thirty-step run about five times as much.

The estimator computes that bill for your own numbers. Everything is in tokens first, because tokens are what the model reads and writes, and the shape of the bill does not depend on whose model you call. Add your own prices at the end if you want money.

Why does the bill compound instead of adding up?

The bill compounds because the model holds nothing between calls. Chapter 19 sets up the problem as a napkin exercise, with every number labeled an illustration: a standing equipment of about 3,000 tokens (system prompt, tool definitions, project instructions), 200 more for the task, and steps that each add about 1,000 tokens to the transcript, 300 written by the model and 700 returned by the tool.

The usual estimate is ten calls at the starting size, about 32,000 input tokens. The real figure is about 77,000, because “every step re-sends the entire transcript so far, because the model cannot remember what it cannot re-read.” Step one reads 3,200 tokens and step ten reads 12,200. In the book’s words, “a run’s input bill is the fixed equipment times the number of steps, plus the per-step growth times a term that rises with the square of the run’s length.”

Why an agent’s cost compounds rather than adds.
Figure 19.2 Why an agent’s cost compounds rather than adds. Every loop step re-sends the whole conversation so far, so each bar re-pays for a stable prefix plus all the history accumulated to that point; only the thin top slice (in accent) is genuinely new. The bars climb like a staircase, and a ten-step task costs far more than ten times a single message. Reuse this diagram

Written as a formula, with F for the fixed prefix, g for the growth per step and n for the number of steps:

Quantity Formula Napkin value (n = 10)
Input tokens F·n + g·n(n−1)/2 32,000 + 45,000 = 77,000
Output tokens (o + D)·n 3,000
Naive estimate F·n 32,000
History share of input g·n(n−1)/2 ÷ input 58%

At thirty steps the same formula gives about 531,000 input tokens, the “roughly half a million” the chapter quotes. The staircase chart in the tool draws each step’s read as a bar: the teal part is the prefix, the same every step, and the ochre part is the history, taller every step.

Which overheads belong in the estimate?

Three overheads belong in the estimate because they are billed whether or not anyone sees them. The first is output pricing: providers charge for “everything the model writes” at a separate rate, “typically priced several times higher per token.” The napkin run wrote 3,000 tokens and read 77,000, which is why the chapter concludes that “the agent’s economics are mostly input economics.”

The second is hidden reasoning. A model’s private deliberation “is billed as output whether or not anyone displays it,” and a step that writes 300 visible tokens “may deliberate for two thousand more.” Enter that as D, the reasoning tokens per step.

The third is retries. Malformed outputs, fallback prompts and corrective round trips all spend full-price tokens with no visible result. The chapter cites one cost-governance framework’s estimate of “ten to twenty percent of total consumption in poorly instrumented systems,” labeled an illustrative range; the tool’s retry slider multiplies the whole bill by that overhead.

How do subagents change the cost?

Subagents multiply the cost because each one is a full agent. The book is specific: every worker “is a complete agent paying the full freight of its own equipment, its own tool results, its own growing history re-read on every step, plus the tokens spent briefing it and merging its findings back.” The tool adds each worker’s own staircase to the lead’s, plus a brief on the way in and a merge on the way out.

The best-known measurement of the result is “roughly four times a chat’s tokens for one agent and fifteen for a multi-agent system,” from a production write-up on a research agent. The book quotes it as a reference point and so does the tool, under the result. It is one team’s measurement on one workload. Your own estimate, built from your own prefix and step counts, is the number to plan with.

A worked example shows how fast this adds up. Keep the napkin lead (10 steps, 77,000 input tokens) and give it three workers, each with a 2,000-token prefix, 8 steps of the same 1,000-token growth, a 1,500-token brief and a 1,000-token merge. Each worker reads 44,000 tokens over its run plus the brief, so the three add 136,500 input tokens to the lead’s 77,000: the team reads almost three times what the lone agent did, before retries.

The same write-up found that “token usage by itself explains 80% of the variance” in performance on a hard browsing benchmark. Read with the cost estimate, that cuts both ways: tokens are what capability is bought with, so the aim is not to spend fewer of them but to waste fewer. The chapter’s metric for this is token yield rate, “the proportion of consumed tokens that contributed to a valuable output.”

Which levers lower the bill?

The levers that lower the bill act on the terms of the formula. Each one shows up in the estimator when you change the matching input:

  • Shorter runs. Because input grows with the square of n, cutting steps saves more than proportionally. Splitting one long run into a chain of short ones, each starting from a compact summary, resets the history term.
  • A smaller growth per step. Trim tool results before they enter the transcript. A tool that returns 700 tokens where 150 would do pays that difference on every later step.
  • Prompt caching. The prefix is identical on every call, so providers that cache a byte-identical prefix bill it at a fraction of the normal rate. Tick the caching box and enter your provider’s fraction; the tool applies it to the prefix of every call. The prompt caching entry explains why the prefix must stay byte-identical.
  • Fewer, leaner workers. A worker with a small prefix and a short task costs far less than a second copy of the lead.
  • A budget. A hard ceiling on steps or tokens, enforced by the harness, turns a runaway run into a capped one. The book treats a budget as one of the three exits every agent loop needs.

How accurate is this estimate?

This estimate is accurate about the shape of the bill and approximate about its size. Real steps vary: some tool results are long, some short, and the model’s output differs from step to step. Use averages from your own traces if you have them; until then, the napkin numbers are a reasonable start for a tool-using agent, which is why the tool loads them by default.

What the estimate cannot tell you is whether the tokens were worth spending. That depends on what the agent produced, which is a question for evaluation rather than accounting. The eval sample-size calculator helps with the other half: how many runs you need to know whether a cheaper configuration is as good as the expensive one.

Questions readers ask

Why does the estimate grow with the square of the step count?
Because the model remembers nothing between calls. Every step re-sends the whole transcript so far, so step ten pays for the nine steps before it. The history term adds up to g·n(n−1)/2, which grows with the square of the run's length.
Why does the tool show tokens and not dollars by default?
Prices differ by provider and change every year, while the shape of the bill does not. Type your own input and output rates to get money; the tool never pre-fills a price.
How do subagents change the bill?
Each worker is a complete agent paying for its own prefix, its own growing history and its own tool results, plus the tokens spent briefing it and merging its answer back. The multi-agent bill is the lead's run plus every worker's run.
Is the 15× multi-agent figure a rule?
No. It is one team's point-in-time measurement, which the book quotes as a reference. Use it to sanity-check your own estimate, never as a multiplier.

Sources

  1. Anthropic (2025). How we built our multi-agent research system
  2. Finout (2026). Token Economics and TokenOps: The Definitive Guide to FinOps for Tokens
  3. Anthropic (2024). Building effective agents