Learn / Explainers / Why agent cost compounds: the token staircase

Why agent cost compounds: the token staircase

Animated explainer · 4:06 · from Chapter 19

Space plays and pauses · ← → step · F fullscreen

A four-minute animated explainer: every agent step re-sends the whole transcript, so a ten-step run reads 77,000 tokens, not 32,000. See the staircase.

Transcript

Every step of the animation, in the words it uses on screen. Select a line to jump to it.

From Chapter 19

Cost, Latency, and Performance

This explainer condenses a few pages of Chapter 19, which is in the full book. The chapter traces the ideas through a worked example; the preface and Part I are free to read online.

Questions

Why is an agent's bill quadratic and not linear in the number of steps?
Because the model holds nothing between calls, every step re-sends the whole transcript so far. Step one reads the fixed prefix F, step ten reads F plus nine steps of history, and the column adds up to Input = F·n + g·n(n−1)/2. The first term grows with the steps; the second grows with their square. On the book's napkin (F = 3,200, g = 1,000) that is 77,000 tokens at ten steps and about 531,000 at thirty. Try your own numbers in the agent cost-per-task estimator.
Does prompt caching make the staircase go away?
No. Caching bills a re-read prefix that matches byte for byte at a discount, and an agent's grow-only transcript is its ideal customer, so it lowers the price of each step. The staircase of tokens is still there: caching discounts input only, it depends on the provider's terms, and a cached context still occupies the desk and dilutes the model's attention. Trimming history at step three still saves at every later step. The cost estimator lets you turn caching on with your own discount.
Why does the book refuse to print prices?
Prices in this field move too fast to print, while the shape of the bill holds still: fixed prefix times steps, plus growth times a term rising with the square of the run's length. That is why the explainer counts tokens. To get money, type your current input and output rates into the agent cost-per-task estimator; it never pre-fills a price.
How do retries and subagents change the estimate?
Retries and corrective round trips spend full-price tokens with no visible result; one cost-governance framework puts them at ten to twenty percent of total consumption in poorly instrumented systems (an illustrative range). Each subagent is a complete agent with its own prefix, its own growing history and its own tool results, plus a brief and a merge. In the explainer's own illustration (Chapter 19 gives no team figures), three workers of eight steps on a 2,000-token prefix would add 136,500 input tokens to the napkin's 77,000. Model your own team in the cost estimator.