How much does an AI agent cost to run? The cost of one task times the number of tasks, where one task is not “calls times the cost of a call.” Every call re-reads the growing history, so tokens rise faster than steps. This guide builds the figure from six terms you can measure, in tokens and ratios, and leaves the prices to you.
I wrote it for a CTO who has a working pilot and a question from finance. There is no price in this post, and that is deliberate: a printed price goes stale and the shape of the bill does not. You will leave with a model whose blanks take your own traces and your provider’s current rates.
How much does an AI agent cost to run? The short answer
An AI agent costs the tokens of one run, weighted by how your tasks split into short and long, multiplied by the runs you start, plus tool fees, infrastructure and review time. Three numbers decide most of it: the steps in a run, the tokens each step adds to the transcript, and the share of runs that go long.
Chapter 19 of the book (in the full book) makes the same choice about prices in its opening: “no prices. Prices in this field move too fast to print. The shape of the bill is what holds still”. The same chapter gives the reason to do the sum before the invoice arrives: “teams that skip the napkin discover their unit economics from the first month’s invoice, which is the most expensive way to learn arithmetic.”
The answer is hard to see from outside a run. One developer wrote in June 2026: “Before starting execution, I often don’t even know how many tokens I’ll need or how much it might cost.” The model below does not remove that uncertainty. It tells you where the uncertainty sits, which is what a budget needs.
What are the terms in the cost of one task?
The cost of one task has six terms: how many steps the run takes, the fixed prompt sent on every step, the history re-read on every step, the output the model writes, the retries and failed runs you still pay for, and the costs that are not tokens. The table names each one, its unit, and where the real figure comes from.
| Term | Symbol | Unit | Where the figure comes from |
|---|---|---|---|
| Steps (model calls) per task | n | calls | Count of model calls in a run’s trace, by task class |
| Fixed prompt per call: system prompt, tool definitions, instructions, task | F | tokens | Input tokens of the first call of a run |
| Growth per step: what the model wrote plus what the tool returned | g = o + r | tokens | Difference in input tokens between consecutive calls |
| Output per step: visible output o plus hidden reasoning D | o + D | tokens | Output tokens reported for each call, reasoning included |
| Retries inside a run, and runs that fail | ρ, s | percent | Calls flagged as repairs; run status at the end |
| Tool fees, infrastructure, human review | (none) | your currency; hours | Tool invoices, hosting, and the review queue’s own timing |
A token is the unit a model reads and writes, and the unit a provider bills. The first four terms are token counts. The fifth multiplies them. The sixth is added at the end, in other units, and it is the one a token calculator leaves out.
One term is absent on purpose: workers. A subagent is, in the chapter’s words, “a complete agent paying the full freight of its own equipment, its own tool results, its own growing history re-read on every step, plus the tokens spent briefing it and merging its findings back.” Model each worker as its own run with the same six terms and add them. The post on multi-agent system overhead does that sum.
Why does token cost grow faster than the step count?
Token cost grows faster than the step count because the model keeps nothing between calls, so each call carries the whole transcript again. Chapter 19, in the section “Where Tokens (and Money) Go,” puts it this way: “every step re-sends the entire transcript so far, because the model cannot remember what it cannot re-read.”
Call the fixed prompt F and the growth per step g. Step 1 reads F, step 2 reads F + g, and step n reads F + g·(n − 1). Summed over the run:
- Input tokens = F·n + g·n(n − 1)/2
- Output tokens = (o + D)·n
The first input term is what the usual estimate counts. The second is the history, and it rises with the square of n. The chapter states the same thing in prose: “a run’s input bill is the fixed equipment times the number of steps, plus the per-step growth times a term that rises with the square of the run’s length.”
Here is a worked example with illustrative numbers. An agent has a fixed prompt of 4,000 tokens; on each step the model writes 250 tokens and a tool returns 950, so g is 1,200. Its tasks fall into three classes by length.
| Task class | Steps n | Fixed part F·n | History re-read | Input tokens | Output tokens | History share of input |
|---|---|---|---|---|---|---|
| Short | 4 | 16,000 | 7,200 | 23,200 | 1,000 | 31.0% |
| Typical | 9 | 36,000 | 43,200 | 79,200 | 2,250 | 54.5% |
| Long | 24 | 96,000 | 331,200 | 427,200 | 6,000 | 77.5% |
The long run has six times the steps of the short one and 18.4 times its input tokens. The calls-times-prompt estimate for the long run is 96,000 tokens, and the real figure is 4.45 times that. The estimator below opens on the long row; move the step slider to 9 or 4 to reproduce the others.
With JavaScript on, the Agent cost-per-task estimator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Agent cost-per-task estimator on its own page to share a result by link.
When does the history overtake the fixed prompt?
The history is the larger part of the input once n exceeds 1 + 2F/g. That follows from setting the two terms equal, and it is this post’s own rearrangement of the chapter’s formula. For the example, 1 + 2 × 4,000 ÷ 1,200 is 7.7, so a run of eight steps or more spends more on re-reading than on the fixed prompt.
The rule tells you which number to chase. An agent with a large prompt and short runs is paying for F, and trimming the prompt is the work. An agent with long runs or bulky tool results is paying for g and n, and a smaller prompt will barely show. The explainer on why agent cost compounds animates the same staircase.
How do you turn tokens into one number without a price list?
Use one fresh input token as the unit and express every other rate as a ratio to it. Then a run costs its billed input tokens plus k times its output tokens, where k is your provider’s output price divided by its input price. Multiply the units by your input price per token on the last line, and only there.
The chapter says output is “typically priced several times higher per token,” so take k = 5 as an illustration and replace it with your own. The long run then comes to 427,200 + 5 × 6,000 = 457,200 units, of which output is 6.6%. Doubling k to 10 raises the total by 6.6%. That is why the chapter concludes that “the agent’s economics are mostly input economics.”
Hidden reasoning is the exception to check first. The chapter warns that “a model’s private deliberation is billed as output whether or not anyone displays it,” and that a step may deliberate for two thousand tokens more than it shows. Put D = 2,000 into the long run: output becomes 54,000 tokens, the total becomes 697,200 units, up 52.5%, and output is 38.7% of it.
A general check, again this post’s own: output matters when k × (o + D) is comparable to the average prompt billed per step. In the long run the average prompt is 17,800 tokens and k × o is 1,250, so output is small. With reasoning on, k × (o + D) is 11,250, and it is not.
What do retries and failed runs add?
Retries and failed runs add tokens that produce nothing, and the meter counts them like any others. The chapter’s taxi image is exact: “A run that fails after forty steps is billed for forty steps.” Two percentages carry this into the model, and they are different things.
The first is the retry overhead inside runs, ρ: repair prompts after malformed output, repeated tool calls, validation passes. One cost-governance write-up (Finout, 2026) says “retries, validation passes, and agent-to-agent communication can account for 10–20% of consumption in poorly instrumented systems, and in complex agent pipelines it can exceed 20%.” That is a vendor’s working range with no stated sample, so use it as a placeholder until your traces give you the real figure.
The second is the share of runs that end in a result someone accepts, s. A run that fails still spent its tokens, so the figure to budget on is cost per successful task:
- Units per run started = class-weighted units × (1 + ρ)
- Units per successful task = units per run started ÷ s
With an illustrative ρ of 15% and s of 85%, the multiplier on the clean figure is 1.15 ÷ 0.85 = 1.353. The division assumes failed runs cost what the average run costs. A failed run that ran to its cap costs more than that, so once traces exist, replace the formula with the direct measurement: all units spent, divided by accepted results.
Careful with double counting. If your token totals come from traces, the retries are already inside them, and ρ is zero in the model. ρ exists only for estimates made before there is anything to measure.
What does the model miss if it stops at tokens?
It misses tool fees, infrastructure and people, and for some agents these are the larger share. None of them follows the token formula, so they sit on their own lines.
Tool and infrastructure cost has a per-call part and a fixed part. The per-call part is whatever each tool invocation costs outside the model: a paid search or data interface, compute seconds in a sandbox, a retrieval query. It scales with steps, so count tool calls per run by class. The fixed part is hosting, queues, storage for traces and logs, and the evaluation runs that accompany every change; it is paid monthly whether one task arrives or a million.
Human review time is share of tasks reviewed × minutes per review × tasks. An illustration: 20,000 tasks a month, one in ten sampled, three minutes each, is 100 hours of someone’s month. One commenter on Hacker News wrote in April 2026 that “Human time spent redirecting AI coding agents towards better strategies and reviewing work, remains dramatically more expensive than the token cost for AI coding, for anything other than hobby work.” That is one person’s judgment about coding agents, and it is a reason to compute this line before deciding the token line is the one that matters.
To compare the two, convert an hour of reviewer time into units at your own rates and divide by the units per run. The result is how many runs one hour of review pays for. If the review line is the larger one, shaving the prompt is the wrong project. Whether the checking is worth what it costs is a separate sum, worked through in the post on AI agent ROI.
How do you get from one task to a month?
Multiply runs started by the class-weighted cost of a run, and never by the cost of the average-length run. Cost is curved in the step count, so the mean of the costs is higher than the cost at the mean, and the long runs carry the bill.
Continue the example with a task mix of 60% short, 30% typical and 10% long, and 20,000 runs started in a month, all illustrative. The weighted input per run is 80,400 tokens. The 10% of runs in the long class consume 53.1% of them. The average step count is 7.5, and the formula at 7.5 steps gives 59,250 tokens, which is 26.3% too low.
Nobody knows the long share in advance, so the monthly figure is a range. Hold everything else still and vary only that share:
| Scenario | Short / typical / long | Units per run started (k = 5, ρ = 15%) | Units per successful task (s = 85%) | Units per month, 20,000 runs |
|---|---|---|---|---|
| Low | 65% / 30% / 5% | 78,574 | 92,440 | 1.57 billion |
| Base | 60% / 30% / 10% | 103,241 | 121,460 | 2.06 billion |
| High | 50% / 30% / 20% | 152,576 | 179,501 | 3.05 billion |
The high case is 1.94 times the low one, and nothing changed except which tasks arrived. Multiply the last column by your input price per token, add the tool fees, the fixed infrastructure and the review hours, and that is the monthly figure with its range.
How wide should the range be?
Wider than a planner’s instinct suggests, on the evidence available. A 2026 study of coding agents (Bai et al.), which analyzed trajectories from eight models on a public benchmark, reports that “runs on the same task can differ by up to 30x in total tokens” and that “task difficulty rated by human experts only weakly aligns with actual token costs.” A second 2026 paper (Ouyang et al.) opens with the same observation: “token consumption can vary by over an order of magnitude across runs.”
Both results are about coding tasks, and I would not carry the multiples to a narrow support agent. The direction carries. Your sense of which tasks are hard is a weak guide to which runs are long, so set the classes from measured step counts and read the percentiles. Chapter 19’s advice for latency applies to cost unchanged: “Read the percentiles rather than the mean.”
Which terms dominate, and which can you ignore?
Steps and growth per step dominate; in a long run the fixed prompt and the price ratio are small, with two exceptions. The table changes one term at a time in the long run of the example, whose base is 457,200 units at k = 5 with no caching.
| Change to the long run | Units | Change |
|---|---|---|
| Half the steps (24 to 12) | 142,200 | −68.9% |
| Smaller tool results, so growth halves (r from 950 to 350) | 291,600 | −36.2% |
| Half the model’s output per step (o from 250 to 125) | 407,700 | −10.8% |
| Half the fixed prompt (4,000 to 2,000) | 409,200 | −10.5% |
| Output price ratio doubled (k from 5 to 10) | 487,200 | +6.6% |
| Hidden reasoning of 2,000 tokens per step | 697,200 | +52.5% |
| Everything already sent read from cache at one tenth the rate (best case) | 101,160 | −77.9% |
Read the output row closely. Halving what the model writes saves 10.8%, and only 3.3 points of that come from paying for fewer output tokens. The rest is history, because everything the model writes is re-read on every later step.
For a first estimate, then, spend your effort on three figures: the step count of the long class, the share of runs in it, and g. Leave k and the visible output rough. Estimate the fixed prompt well only if your runs are short enough to sit below the crossover. Come back to output if reasoning is on, and to everything if caching is on, because caching reorders the table.
Do these rules hold for two other agents?
They hold for two cases built differently, both illustrative. A document pipeline with F = 1,500, a model output of 400 and tool results of 3,000 tokens crosses over at 1 + 3,000 ÷ 3,400 = 1.9 steps. Over a six-step run it reads 60,000 input tokens, 85.0% of them history, and output is 16.7% of its 72,000 units. The work there is the size of the tool results.
A coding agent with F = 12,000, an output of 500 and results of 2,500 tokens crosses over at 9 steps. At 45 steps it reads 3,510,000 input tokens, 84.6% of them history, and output is 3.1% of the total. The work there is the length of the run.
What does prompt caching change?
Prompt caching changes the price of re-reading, which is the largest term, and so it changes which term is largest. Prompt caching bills input that matches the opening of an earlier call at a fraction of the fresh rate. The chapter describes the best case for a loop: “the entire history becomes one long cached prefix and the meter charges full price only for the tail.”
In that best case, the first call reads F fresh and each later call reads only its newest g fresh. For the long run that is 4,000 + 23 × 1,200 = 31,600 fresh tokens and 395,600 cached ones, 92.6% of the input. Write c for the cached rate as a fraction of the fresh rate, and take c = 0.10 as an illustration to be replaced by your provider’s figure. Billed input is 31,600 + 0.10 × 395,600 = 71,160 units, and the run totals 101,160 with output.
Two things follow. The run costs 77.9% less, and across the base mix the saving is 68.8%. Output is now 29.7% of the long run, so the terms the last section called small are small no longer.
Treat those figures as a ceiling. The chapter’s limits apply: caching “discounts input only, does nothing for output; it is a billing fact, only as durable as the provider’s terms,” and one changed token near the front of the prompt ends the match from that point. Some providers also charge more to write a cache entry than to read fresh input, which this example leaves at zero.
Measured reuse is lower than the ideal. A 2026 analysis of production traces from one coding assistant (Liu et al.), sampling 761 million model calls, reports serving-side cache hit rates “averaging 90% within a turn, but falling to 55% across turn boundaries.” That is the serving system’s cache and not a billing figure, but it is the same reuse a discount depends on. Put your measured cached share into the model. The prompt caching savings calculator does this arithmetic for any layout; the estimator above discounts the fixed prompt only.
The model, with blanks
Here is the whole model as one block to copy into a spreadsheet or a planning document. Every blank is a figure from your own traces or your provider’s current price page. Money appears on the last lines only.
AGENT RUN-COST MODEL Agent: ____________ Rates dated: ________
A. ONE CALL (tokens, averaged from traces)
F fixed prompt per call (system + tools + instructions + task) ______
o visible model output per step ______
D hidden reasoning per step (billed as output) ______
r tool result returned per step ______
g growth per step = o + r ______
Crossover: history exceeds the fixed part when n > 1 + 2F/g = ______
B. ONE RUN, PER TASK CLASS (classes set from measured step counts)
input = F·n + g·n(n−1)/2 output = (o + D)·n
class share of runs steps n input tokens output tokens
short ______ % ______ ____________ ____________
typical ______ % ______ ____________ ____________
long ______ % ______ ____________ ____________
workers: add each worker as its own run, plus its brief and its merge
C. RATIOS FROM YOUR PRICE PAGE (unit = one fresh input token)
k output price ÷ input price ______
c cached-read price ÷ input price (1 if no caching) ______
h measured share of input tokens read from cache ______ %
billed input = input × (1 − h) + input × h × c
units per run = billed input + k × output
D. OVERHEADS
ρ retry overhead inside runs (0 if section B came from traces) ______ %
s share of runs that end in an accepted result ______ %
units per run started = Σ (share × units per run) × (1 + ρ) = ______
units per successful task = units per run started ÷ s = ______
(better, once measured: all units spent ÷ accepted results)
E. NOT TOKENS
tool fees per run (calls × fee, by class) ______
fixed infrastructure per month (hosting, storage, eval runs) ______
review: share reviewed ____ % × minutes ____ × runs = hours/month ______
M. THE MONTH
V runs started per month ______
token units per month = V × units per run started = ______
money = units × input price per token
+ V × tool fees + fixed infrastructure + review hours × rate
RANGE: repeat sections D and M with the long share at its low and high measured values
low ____________ base ____________ high ____________
Assumption that moves it most: ______________________________________
How do you measure the real figures from traces?
Read them from the token counts your provider returns with every call, recorded per call and grouped by run. A trace is the record of one run as a sequence of timed steps. If each model call’s span carries its input, cached and output token counts, every blank in the template is a query. The chapter expects this handover: “real traces will retire the spreadsheet within a week.”
Pull these before the budget conversation, over at least a few hundred runs of real traffic.
- Model calls per run, as a distribution: median, 90th and 99th percentile
- Input tokens of the first call of a run, which is F
- The average difference in input tokens between consecutive calls, which is g
- Output tokens per call, with reasoning tokens counted separately where the provider reports them
- The share of input tokens billed as cached, per run and overall
- Tool result size in tokens, by tool, to find the tool that drives g
- Paid tool calls per run, by class
- Run status at the end: accepted, failed, stopped by a cap, abandoned
- Calls that were repairs or retries, as a share of all calls
- The share of total tokens spent by the most expensive tenth of runs
- All units spent divided by accepted results, which is cost per successful task
- Review minutes per reviewed task, timed with the people who review
The post on LLM tracing with OpenTelemetry covers how to record these on spans, and the post on how to monitor AI agents in production covers what to chart afterwards.
Which alerts fire before the invoice does?
Three alerts cover the ways a monthly figure goes wrong, and all three are in tokens. A budget on each run is the first: a ceiling on steps and tokens enforced by the harness, so the worst run has a known price. The chapter’s reading is that “a capped run has a bounded worst-case bill and an uncapped one does not.”
The second is a daily burn alert: tokens spent so far this month against the base scenario’s pace, with the high scenario as the line that pages someone. The third is drift in units per successful task by class, which catches a prompt that grew, a tool that started returning more, or a success rate that fell.
Watch tokens and not only money. The chapter warns that “a stable invoice can conceal token volume growing as fast as prices fall, so watch consumption, never just the bill.” A sudden jump in the long share also has two causes that are not cost problems. One is a stuck run, covered in the post on what to do when an AI agent keeps looping. The other is hostile input that sends the agent on extra work, one of the outcomes of indirect prompt injection attacks.
How do you bring the number down?
Work on the term your sensitivity table ranks first, and pair every change with an evaluation run. For the example that order is shorter runs, smaller tool results, caching, then the prompt. The chapter’s condition applies to all of them: “Cheaper is only better if the task still succeeds.”
The aim is not the smallest token count. One lab’s 2025 write-up of a research system found that “token usage by itself explains 80% of the variance” in performance on one browsing benchmark, which is one team’s analysis of one workload. Tokens buy capability, so the target is waste: retries, discarded work and context that is paid for and never used. The levers themselves, ranked, are the subject of the planned post on how to reduce AI agent costs.
Where does this model stop working?
The model is exact about arithmetic and rough about your agent, in four places. Real steps are not uniform: one tool returns several times what another does, and an average g hides it. A harness that compacts or summarizes history breaks the smooth staircase, which lowers cost and makes the formula an upper bound.
Three classes are a simplification of a continuous, long-tailed distribution. If your 99th-percentile run is far beyond your long class, add a fourth class for it or set the cap there. The study cited above also found that models asked to predict their own usage “systematically underestimate real token costs,” so do not let the agent fill in the blanks.
The model says nothing about value. It will price an agent that produces nothing useful as confidently as one that does. And the unit trick depends on your provider pricing by the token with stable ratios; under a flat subscription or a per-task contract, the tokens are still the vendor’s cost and the model becomes a way to check their quote, as the post on AI agent pricing models explains.
The figure to take to the budget meeting
When someone asks how much does an AI agent cost to run, answer with one figure and its range: units per successful task, measured, times the accepted results you expect, at today’s rates, with the long share named as the assumption that moves it. That is a number finance can hold you to and a number you can re-measure next month from the same traces.
If there are no traces yet, take the estimate and say what it is. Then record tokens on every call from the first day of the pilot, because the long runs you did not predict are the bill.
Chapter 19, “Cost, Latency, and Performance,” develops the napkin sum, the five cost levers and the trade against latency and accuracy (in the full book). The AI agent security and operations guide places this post beside its neighbors, or you can see the formats.
Questions readers ask
- How much does an AI agent cost to run per month?
- Monthly cost is the number of runs started, times the cost of one run weighted by task class, times one plus the retry overhead, plus tool fees, infrastructure and review time. Price it at your own current rates. State it as a range, because the share of tasks that run long moves the total more than any price does.
- Why is my AI agent bill higher than calls times price?
- Each call re-sends the whole transcript so far, so later calls are larger than early ones. Input tokens for n steps are the fixed prompt times n plus the growth per step times n(n−1)/2. In this post's illustrative 24-step run the true input is 4.45 times the calls-times-prompt estimate.
- Are output tokens or input tokens the bigger cost for an agent?
- Input, in a tool-using agent with no caching, because the history is re-read on every step. In the illustrative long run, output is 6.6% of the total even when an output token is priced at five input tokens. Hidden reasoning or prompt caching can raise that share to about 30% or more.
- Does prompt caching make an agent cheap?
- It makes re-reading cheap, and only that. Caching discounts input tokens that match the previous call's opening bytes; it does not discount output, and a single changed token near the front of the prompt removes the discount from that point on. Use your measured cached share, not the ideal one.
- How do I estimate AI agent costs before I have any traces?
- Guess the fixed prompt size, the tokens added per step and the step count for a short, a typical and a long task, and compute each class separately. Treat the result as a shape to be corrected. A 2026 study of coding agents found runs on the same task differing by up to 30 times in total tokens.
Sources
- Longju Bai et al. (2026). How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
- Chaoqian Ouyang et al. (2026). TokenCast: Forecasting Token Consumption During LLM Agent Execution
- Banruo Liu et al. (2026). Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale
- Finout (2026). Token Economics and TokenOps: The Definitive Guide to FinOps for Tokens
- Anthropic (2025). How we built our multi-agent research system
- tabbott (2026). Hacker News comment on the cost of coding agents against human time
- kirillAIsolo (2026). Hacker News comment on not knowing the cost before a run