What is a prompt caching calculator?
A prompt caching calculator estimates how many of an agent run’s input tokens are read from a provider’s cache instead of being processed fresh, and what that does to the bill. This one works step by step through a run, draws the cached and fresh tokens at each call, and lets you place one changing token in the prompt to see how much of the saving it destroys.
Prompt caching is the third cost lever in Chapter 19 of the book, which defines it this way: “the provider stores the reading-state of your prompt’s front, and any later call whose opening tokens match byte for byte loads the stored state and starts reading where the match ends.” The calculator turns that sentence into arithmetic. Its companion, the agent cost-per-task estimator, sizes the whole bill; this page looks at one lever on it.
How does prompt caching work?
Prompt caching works because a model reads left to right, and what it computes for the front of a prompt cannot be changed by anything that comes later. Chapter 19 derives the mechanism from a property the book introduces early, that generation is autoregressive: “The internal state the model builds while reading a prefix depends on that prefix alone, so tokens appended later can never reach back and change it”.
That makes the state reusable. A provider that has computed it once can store it, and a later call that begins with the same tokens can pick up from there. Two things follow for whoever runs the agent. Cached tokens are charged at a fraction of the fresh rate, and responses start sooner, because the skipped reading is work the model would otherwise do before producing its first word.
The chapter is careful about the size of the discount, and so is the tool. “Cache hits are billed at a steep discount to fresh input (commonly a small fraction of the rate, though the figures vary by provider and by year)”. The factor in the calculator is yours to type. The 10% it starts with is an illustration, and the result is reported in token-equivalents: fresh tokens count as one, cached tokens as your factor.
Why is an agent loop the ideal case for caching?
An agent loop is the ideal case because each call’s prompt is the previous call’s prompt with a little more on the end, so almost everything the model reads has been read before. The book puts it this way: “every step’s prompt is the previous step’s prompt plus an appendix, so the entire history becomes one long cached prefix and the meter charges full price only for the tail.”
The numbers make the point better than the principle does. Chapter 19’s own worked run starts at 3,200 tokens (about 3,000 of standing equipment plus a 200-token task) and grows by 1,000 tokens a step, figures the chapter offers as illustration. Over ten steps the model reads about 77,000 input tokens, because every step re-reads everything before it. Load those numbers with the button above and the prompt caching calculator shows the split, which is the tool’s arithmetic on the chapter’s run:
| Tokens | Share of input | |
|---|---|---|
| Input with no cache | 77,000 | 100% |
| Read fresh (the first call, then each step’s appendix) | 12,200 | 16% |
| Read from cache | 64,800 | 84% |
| Billed, at an illustrative factor of 10% | 18,680 token-equivalents | 24% |
The fresh column is small because it grows in a straight line: the first call pays for the whole prefix, and each later call pays only for what the last step appended. The total input grows with the square of the run’s length. Caching removes most of the squared part of the bill, which is why long runs benefit more than short ones. Drag the step slider to 30 and the cached share reaches 94%.
What is the most common prompt caching mistake?
The most common mistake is putting something that changes on every call near the front of the prompt, which ends the match at that point and leaves everything after it to be read fresh. The chapter’s rule is six words long: “stable content first, variable content last.”
It also names the usual culprits. “Change one early token and the cache is dead from that point on, so the classic self-inflicted misses are all small things placed carelessly near the front—a timestamp in the system prompt, a per-user ID above the tool definitions, tool schemas serialized in whatever order the framework felt like today.”
The calculator has each of these as a choice. Pick the timestamp on the napkin run and the cached share drops from 84% to zero, because the very first line differs on every call. Pick the position above the tool definitions and only the 800-token system prompt still matches, which leaves 9% cached when the value there changes on every call. The tool then draws a second chart with the same run after the variable content has moved to the tail, and states the difference per run. On the napkin numbers, the move saves about 52,000 token-equivalents on every ten-step run.
The fix is usually a small edit. A timestamp or a request ID can go in the latest message, after the history. Tool schemas can be serialized in a fixed order. A per-user value is best placed after everything that all users share. The book’s advice on where to look is “Audit the front of your prompt the way you would audit a hot path.”
One caveat about the positions. The book names the three misses; where exactly each one sits in the prefix is the tool’s assumption, based on a layout of system prompt, then tool definitions, then instructions. The tool also models whatever sits at the chosen position as changing on every call, which is true of a timestamp, a request ID or an unstable schema order. A true per-user ID is gentler: it stays the same within one user’s run, so it costs only the first call of that run, which cannot reuse the tools and instructions that other users’ calls already warmed.
Does caching mean you can stop trimming context?
No. Caching changes what a token costs to re-read and changes nothing about the token being there. Chapter 19 is direct about this: “it does not repeal context discipline, because a cached context is cheaper to re-read yet still occupies the desk”. Then comes the line worth remembering: “Cheap clutter is still clutter.”
The result table keeps this in view. Whatever the cached share, the last step of the run still reads the full prompt, and the model still has to find what matters in it. The problems the book groups under context rot depend on how much is in the window, and a discount on the meter does not shrink the window. The context window budget planner is the tool for that side of the question.
The two levers also interact. What follows is the tool’s reasoning from the byte-for-byte rule, not something the chapter spells out, and the calculator does not simulate it. Summarizing or removing old history rewrites the middle of the prompt, which ends the cached match at the point of the edit. That can still be the right trade, since fewer tokens are read at all from then on. It does mean a compaction step has a one-time cost in fresh tokens, and that editing the history a little on every call is the worst of both.
What does the calculator not model?
The calculator does not model output tokens, cache expiry, minimum cacheable sizes or how a given provider decides what to store. It models one run, with one changing token, under the assumption that every cached prefix is still available when the next call arrives.
Output is the largest omission and it is deliberate. The mechanism, in the chapter’s words, “discounts input only, does nothing for output”. Agent bills are mostly input, so the lever is still large, and the cost-per-task estimator adds the output side.
Expiry matters more than it looks. Providers keep a cached prefix for a limited time, so an agent that pauses for a human approval, or a run that resumes the next morning, may find the cache gone and pay fresh for the first call back. Some providers also charge more than the fresh rate to write a prefix into the cache. The optional surcharge field covers that case, and it produces a result worth seeing once: with a surcharge and a prefix that never matches, caching costs more than no cache at all. Both expiry and the surcharge are outside the book’s text; treat them as the tool’s additions and check your provider’s current terms.
Checking the terms is also the chapter’s advice. It calls caching “a billing fact, only as durable as the provider’s terms”. The rule about ordering will outlast any particular rate, because it follows from how models read. The percentages will not. The guide to agent security and operations places this lever beside routing, context discipline and caps, and the full argument, with the other three levers, is in Chapter 19, Cost, Latency, and Performance (in the full book).
Questions readers ask
- Why does the prefix have to match byte for byte?
- Because the cache stores the state the model built while reading a specific sequence of tokens. That state depends on every token before it, so one changed token early in the prompt makes everything after it a different sequence. The provider can reuse the stored state only up to the first difference, and reads the rest fresh.
- Why does prompt caching not reduce output tokens?
- Because the cache saves reading, and output is writing. Each output token is generated new on every call. Chapter 19 lists this among the honest limits of the mechanism: it discounts input only and does nothing for output, which usually costs more per token than input.
- If the cache makes history cheap, can I stop trimming context?
- No. A cached token is cheaper to bill and exactly as present in the context window. The model still attends over all of it, so the quality problems of a crowded context are unchanged. The book's summary is that cheap clutter is still clutter.
- What is the most common reason a cache never hits?
- Something small that changes on every call sits near the front of the prompt. The book names three: a timestamp in the system prompt, a per-user ID above the tool definitions, and tool schemas serialized in a different order each time. Each one kills the cache from its position onward.
- Does the calculator know my provider's prices?
- No, on purpose. You type the factor your provider charges for cached input, and any surcharge for writing to the cache. The result is in tokens and token-equivalents. Rates differ by provider and change by year, and the shape of the saving does not.
Sources
- Yichao 'Peak' Ji (2025). Context Engineering for AI Agents: Lessons from Building Manus
- Anthropic (2026). Prompt caching (provider documentation, example)
- OpenAI (2026). Prompt caching (provider documentation, example)
- Woosuk Kwon et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention