Home / Blog / Security, reliability and cost / How to Reduce AI Agent Costs: Six Levers, Ranke…

Security, reliability and cost

How to Reduce AI Agent Costs: Six Levers, Ranked by Return

How to reduce AI agent costs: six levers priced on one illustrative workload, ranked by return per effort, each with its quality risk and proof. Rank yours.

By Enrique Gutiérrez · Published · 17 min read

How to reduce AI agent costs comes down to six levers: cache the stable prefix, cap every run, batch what nobody waits for, route easy work to a cheaper model tier, trim what is re-read on every step, and move bulk work into code. Rank them by the share of your bill each removes per day of effort, with quality held.

I wrote this for a CTO who has the invoice and one engineer to assign. It is the sequel to the post on how much an AI agent costs to run, which builds the cost model. This post uses that model’s terms and does not rebuild them. There are no prices here, only tokens and ratios, and every number in the worked example is illustrative.

How to reduce AI agent costs: the short answer

Price each lever alone on your own traces, divide the saving by the effort to ship it, apply the highest, measure, and rank again. Hold two things constant while you do: the share of runs that end in an accepted result, and the pass rate on a fixed evaluation set.

The levers are the book’s. Chapter 19 (in the full book), in the section “Cost Levers,” says: “There are five levers worth your attention, and I will take them in the order that pays most often.” It then adds batch processing as a habit that “rounds out the list rather than joining it.” This post counts the habit as the sixth.

The chapter’s order is model choice, context discipline, caching, the cap, and work done in code. The order below is different, because it comes from one illustrative workload and a rule the chapter leaves to the reader. Your order will differ again, and finding it is the reason to price the levers on your own traces.

What unit should a saving be stated in?

State every saving in units per accepted result: all units spent, divided by the results someone accepted. A unit is one fresh input token, with output and cached tokens converted by your provider’s ratios, as the cost model post explains. Per-call and per-run figures leave out the runs that failed.

The chapter’s taxi-meter image gives the reason: “A run that fails after forty steps is billed for forty steps.” One developer put the consequence this way in November 2025: “‘Cost per successful task’ is the only metric that matters in production.”

Acceptance sits in the denominator, so it is a cost lever too. In the example below, raising accepted results from 850 to 900 per 1,000 runs lowers the cost of each by 5.6% without saving a token.

The unit has a blind spot. A change that cuts the bill by a quarter still looks good per accepted result after acceptance has fallen a long way, as the tiering section shows. So acceptance and the eval pass rate are gates of their own, checked separately.

What does the example workload look like before any lever?

The example is 1,000 runs of one agent, costing 127,323,000 units, of which 850 end in an accepted result: 149,792 units each. It reuses the earlier post’s agent, with a fixed prompt F of 4,000 tokens, growth per step g of 1,200 (250 written, 950 returned by a tool), and output priced at five units per token.

Run class Steps n Runs Units per run Share of the bill
Short 4 600 28,200 13.3%
Typical 9 300 90,450 21.3%
Long 24 90 457,200 32.3%
Runaway, no cap, failed 80 10 4,212,000 33.1%

Ten runs in a thousand take a third of the bill. I added that class to the earlier post’s mix because an uncapped tail is what the cap is for; if your traces show no such runs, that lever is worth less to you than it is here.

Which lever returns the most on this workload?

Caching returns the most by a wide margin when each lever is applied alone, and the cap is nearly as good per day of effort. The table gives every lever the same fields. The buttons narrow it by the shape of your bill.

A history-heavy bill comes from runs long enough that re-reading exceeds the fixed prompt, which the cost model post puts at n above 1 + 2F/g. A prompt-heavy bill comes from short runs with a large fixed prompt. The other three tags are a runaway tail, an easy majority of tasks, and work nobody is waiting for.

Lever Term it cuts Bill removed, alone Effort (days) What it risks The proof it worked Pays most on
1. Cache the stable prefix The rate paid on re-read input (F and the history) 74.1%, a ceiling (3.86 times cheaper) 3 Nothing in the output; a silent miss if the prefix changes Cached share of input tokens per call, by run history-heavy, prompt-heavy
2. Cap steps, tokens and fan-out n in the worst runs 24.1% (1.32 times) 1 Stopping a long run that would have finished Runs stopped by the cap, read one by one; acceptance unchanged runaway tail
3. Batch or defer The rate, on work with no one waiting 15.0% (1.18 times) 3 Nothing in the output; turnaround time Share of units billed at the batch rate; deadline misses nobody waiting
4. Route to a cheaper model tier The rate, on runs the small tier can do 27.7% (1.38 times) 10 Lower acceptance, repair loops, review time Eval pass rate per tier; escalation rate easy majority
5. Trim what is re-read g, the growth per step 33.4% (1.50 times) 5 Dropping a detail a later step needed Tokens per tool result; eval pass rate history-heavy
6. Move bulk work into code and better tools n, and g for bulky results 30.0% (1.43 times) 10 A wrong script; a wider attack surface Steps per run by class; eval pass rate history-heavy, prompt-heavy

The effort column is my placeholder for a small team and is the first thing to replace. The parameters behind each row: cached reads billed at one tenth of fresh input; a cap of 40 steps; 30% of units at half rate; the short and typical classes run whole on a tier at one fifth of the rate; tool results trimmed from 950 to 350 tokens; and steps cut by about a third in the three healthy classes, from 4 to 3, 9 to 6 and 24 to 16.

What is the ranking rule?

Divide the share of the current bill a lever removes by the days it takes to ship, take the highest, ship it with an eval run, then recompute every remaining lever on the new bill. The rule is this post’s own. The chapter supplies the condition it runs under: “Cheaper is only better if the task still succeeds.”

Recomputing matters because the levers overlap. Here is the example with the rule applied six times.

Order Lever Bill left (units) Removes, of the bill at that point Removes, of the original bill Cumulative
1 Cache 33,024,600 74.1% 74.1 points 3.86 times cheaper
2 Cap 29,076,600 12.0% 3.1 points 4.38 times
3 Batch 24,715,110 15.0% 3.4 points 5.15 times
4 Tier 12,530,190 49.3% 9.6 points 10.16 times
5 Trim 9,203,970 26.5% 2.6 points 13.83 times
6 Fewer steps 6,506,580 29.3% 2.1 points 19.57 times

Trimming removes 33.4% of the bill alone and 2.6 points of the original when it comes fifth. Caching had already made re-reading cheap, so there was less left to trim away. Tiering moved the other way, from 27.7% alone to 49.3% of what remained, because caching shrank the long runs and left the short ones a larger share.

Do not read the last line as a forecast. It assumes the cache hits on every call, acceptance never moves, and every lever lands in full. The first three rows change nothing the agent returns in this example, where the cap stops only runs that had already failed; a cap set too low would. The last three can change it at any setting, and they are the ones an eval run will trim.

Does the rule hold on three other agents?

It gives a different order on each, which is the point of running it. Three single-class cases, all illustrative, with each lever applied alone. The model writes 400, 500 and 200 tokens per step in the three cases. Trimming cuts each tool result from 3,000 to 1,000, from 2,500 to 1,250 and from 300 to 150 tokens, and the steps go from 6 to 4, 45 to 30 and 3 to 2:

Agent F, g, n Cache Trim results A third fewer steps
Document pipeline 1,500, 3,400, 6 51.9% 41.7% 52.2%
Coding agent 12,000, 3,000, 45 83.6% 34.2% 52.0%
Support agent, large prompt 9,000, 500, 3 52.9% 1.4% 34.9%

The support agent is prompt-heavy: its runs end long before the crossover at 37 steps, so halving its tool results saves 1.4%. That matches the table’s tags, where trimming is marked for history-heavy bills only. Caching and fewer steps pay on all three.

Lever 1: is the prefix you pay for actually cached?

Caching bills input that matches the opening of an earlier call at a fraction of the fresh rate, so it cuts the price of re-reading without changing a token of output. Prompt caching returns the most in the example because re-reading is most of the bill: 92.6% of the long run’s input is a repeat of the call before.

The effort is an audit, and the chapter states its rule: “stable content first, variable content last.” Its list of self-inflicted misses is “a timestamp in the system prompt, a per-user ID above the tool definitions, tool schemas serialized in whatever order the framework felt like today.”

The calculator below opens on the long run with the third of those faults. Caching is on, and the tools are serialized differently on each call. It reports 23,000 of 427,200 input tokens cached and a saving of 5%; with the variable content moved to the tail, 395,600 tokens are cached and the input bill falls 83%.

With JavaScript on, the Prompt caching savings runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Prompt caching savings on its own page to share a result by link.

This failure is silent, because nothing errors. One developer described finding it in June 2026: “i noticed cache_read-tokens per turn were around 0 cause the prefix invalidated every turns cache.”

The 74.1% is a ceiling. A 2026 analysis of production traces from one coding assistant (Liu et al.) reports serving-side cache hit rates “averaging 90% within a turn, but falling to 55% across turn boundaries and drastically invalidated after events like model switches or context compaction.” That is a serving cache and not a billing figure, but a discount depends on the same reuse. The proof is the cached share of input tokens per call, read from your traces.

Lever 2: does every run have a cap?

A cap puts a ceiling on steps, tokens and spawned workers, so the worst run has a known price. A budget costs a day to add where the harness already counts steps. In the example it takes the ten runaway runs from 80 steps to 40 and removes 24.1% of the bill.

The chapter’s reading: “a capped run has a bounded worst-case bill and an uncapped one does not.” It adds a warning about fan-out: “A subagent that can spawn subagents is the invoice’s version of unbounded recursion.”

The risk is stopping a run that would have finished. Set the cap from the measured distribution of step counts, well above the longest accepted run, and read every run it stops for the first weeks. A cap that fires is also a diagnosis, and the post on what to do when an AI agent keeps looping covers the causes and the earlier exits.

Lever 3: what can wait?

Work with no one waiting can go to the discounted batch processing that many providers offer, or simply run later. The chapter’s wording is to “route it to the discounted batch processing most providers offer, where you trade turnaround time for a lower rate.”

The output is the same, so the only risk is time. The saving is the share of units that can wait times the discount, which is why it is 15.0% here and could be zero for an interactive product. Check your provider’s terms before assuming a batch rate and a cache discount combine; the example assumes they do. The proof is the share of units billed at the batch rate, with missed deadlines beside it.

Lever 4: which work can a cheaper model tier do?

Routing sends easy work to a smaller model and keeps the strong one for what needs it. The chapter’s default is to “start on the cheapest model that plausibly works, and promote a step, a role, or a whole route to a stronger tier only where the eval set” shows the cheap one failing.

Published results show the size of the possible win, each on its own benchmark. A 2025 study of strong-weak collaboration on repository-level coding tasks (Gandhi et al.) reports that its best strategy “achieves equivalent performance to the strong model while reducing the cost by 40%.” The 2023 cascading paper (Chen, Zaharia and Zou) reports matching the best single model “with up to 98% cost reduction,” a best case on single queries and not on agent runs. A 2024 analysis (Dekoninck et al.) identifies “good quality estimators as the critical factor.”

This lever carries the most quality risk of the six. In the example, tiering alone leaves the bill at 92,079,000 units, or 108,328 per accepted result. If the smaller tier costs 50 accepted results, the figure is 115,099, still 23.2% below the start. It stays below until acceptance falls to 615 of 1,000.

That is the unit’s blind spot, in numbers. Cost per accepted result would approve a change that lost more than a quarter of the accepted work. One practitioner wrote in August 2026: “The cheaper model isn’t actually cheaper if a human has to spend the savings babysitting it.”

Tiering can also undo lever 1. A commenter explained in June 2026: “Switching models invalidates the cache, meaning everything up to the point of the switchover is processed like a new, uncached input token.” With this post’s illustrative ratios, a 17,200-token history costs 1,720 units as a cached read on the strong tier and 3,440 as a fresh read on a tier at one fifth of the rate.

So route whole runs where you can, which is what the example does. The cascade cost estimator prices a screen placed in front of an expensive call, with its break-even escalation rate. The proof is the eval pass rate per tier and the escalation rate; the chapter warns that “a cascade where half the traffic escalates is a cheap tier that is failing its job.”

Lever 5: what is re-read on every step and never used?

Trimming cuts g, the tokens each step adds, and every token removed is saved once per remaining step. The chapter’s example: “A thousand-token result that survives a thirty-step run is billed thirty times.” One vendor’s guide, Finout’s in 2026, calls context “Often the highest-leverage layer” and says trimming “can reduce input consumption by 30–60%.” That is a working range with no stated sample.

There are two ways to trim, and they treat the cache differently. Shaping a tool result before it enters the transcript leaves the prefix intact. Rewriting earlier history ends the cached prefix from the rewrite point. The post on context compaction for AI agents builds the second and prices what a rewrite gives up, so start with the first.

The risk is dropping a detail a later step needed. The proof is tokens per tool result, by tool, beside the pass rate on an eval set that includes long runs. The chapter notes that this lever often improves the agent as well, and that caching does not replace it: “Cheap clutter is still clutter.”

Lever 6: which steps should be code?

Bulk and exact work belongs in a program the agent writes or calls, with the intermediate data kept out of the transcript. Tools that do more per call belong here too, since they shorten the run. The chapter puts it in one sentence: “If a step of your agent’s work could be a program, the cheapest tokens are the ones that write the program.”

It cuts n directly and g for bulky results, because “an intermediate value that never enters the transcript is never re-read.” In the example a third fewer steps removes 30.0% alone. It is applied last here: once the first lever is in, it returns the least per day at every later step, and at ten days it shares the longest estimate on the list with tiering.

Two risks come with it. A script can be wrong in ways a reader of the transcript no longer sees, so test its output. And an agent that runs code is a larger target. If that agent also reads untrusted content and can reach the network, check it against the lethal trifecta before shipping. The proof is steps per run by class and the eval pass rate.

Which levers also make the agent faster?

Four of the six also shorten the wait, one does nothing for it and one lengthens it on purpose. The chapter’s line is that “Most of the cost levers are latency levers, because tokens are both the currency and the workload.”

A cached prefix and a trimmed context both cut the reading the model does before its first word. Fewer steps shorten the chain of calls. A smaller model reads and writes faster per token, with one exception the chapter names: “The cascade saves money and adds worst-case waiting.” A cap bounds the worst case only. Batching sells the wait for a lower rate. So if the complaint is that the agent is slow as well as expensive, levers 1, 5 and 6 answer both.

The accuracy–cost–latency triangle.
Figure 19.5 The accuracy–cost–latency triangle. Each corner is one thing you might maximize, and each edge names the pairing you can actually hold at once—accurate and cheap, accurate and fast, or cheap and fast. The verdict in the middle is the whole chapter in two words: you generally get to pick two. Reuse this diagram

The figure is the chapter’s summary of what remains after the waste is gone: “accuracy, cost, and latency form a triangle, and you usually get to pick two.” Until then, the chapter says, “eliminating waste is not a trade.” Levers 1 to 3 are waste removal. Levers 4 to 6 can become trades, which is why they carry an eval run.

The checklist: rank, apply, prove

Work through this list against your own agent. It assumes each model call in a trace records its input, cached and output tokens, and that each run records how it ended.

  • Compute units per accepted result for the last few hundred runs: all units spent divided by accepted results
  • Split the bill by run class and find the share spent by the most expensive tenth of runs
  • Freeze an evaluation set and record its pass rate and its run-to-run noise before changing anything
  • Price each of the six levers alone on those traces, as a share of the current bill
  • Write an effort estimate in days beside each lever and divide
  • Read the cached share of input tokens per call; if it is near zero after the first call, find what changes near the front of the prompt
  • Move every variable value behind the stable content, and fix the order in which tools are serialized
  • Confirm every run has a cap on steps and tokens, and every worker a cap on spawning
  • Read each run the cap stopped, and raise the cap only on evidence
  • List the work with no one waiting, and check that your provider’s batch and cache terms combine
  • Route whole runs to a cheaper tier, never switching mid-run without pricing the lost cache
  • Trim tool results before they enter the transcript, starting with the tool that adds the most tokens
  • Move bulk or exact steps into code, test the code, and check the agent’s new reach
  • After each change, report the pair: units per accepted result and the eval pass rate, then re-rank

Where does this ranking break down?

The ranking is exact arithmetic on an invented workload, and it is rough in four places. The effort figures are placeholders. The caching row is a ceiling, and a provider that charges more to write a cache entry lowers it; the calculator above has a field for that.

The levers are modeled as independent switches, and real ones are not: a trimmed context can change how many steps a run takes. The example also holds acceptance at 850 throughout, which is the thing an eval run exists to check.

The aim is not the smallest bill. One lab’s 2025 write-up of a research system found that “token usage by itself explains 80% of the variance” in performance on one browsing benchmark, which is one team’s analysis of one workload. The chapter draws the conclusion: “the aim of everything that follows is yield, never abstinence.” If the review hours or the value of the task are the real question, the token bill is the wrong place to look, and the post on AI agent ROI takes that up.

The pair of numbers

The short version of how to reduce AI agent costs is to do the changes that cannot alter the output first, measure, and let the measurement choose the next one. Check that the cache hits and that every run has a cap this week. Then price the rest on your traces.

Whatever you ship, report it as the chapter asks: “The pair of numbers is the deliverable; either alone is a story.” Units per accepted result, and the pass rate on a set you did not change.

Chapter 19, “Cost, Latency, and Performance,” develops the five levers, the two meters and the triangle (in the full book). The AI agent security and operations guide places this post beside its neighbors, or you can see the formats.

Questions readers ask

What is the fastest way to reduce AI agent costs?
Check two things first, because both take days and neither changes what the agent returns: whether the stable front of the prompt is actually being read from cache on every call, and whether every run has a cap on steps and tokens. On this post's illustrative workload those two changes remove 77.2% of the bill together, as a ceiling.
Does using a cheaper model always reduce cost?
No. It lowers the rate per token and can lower the share of runs that end in an accepted result, add repair loops, or add review time. Switching models in the middle of a run also discards the cached prefix. Compare cost per accepted result and the eval pass rate before and after, and route whole runs where you can.
Why did prompt caching not lower my bill?
The usual cause is something near the front of the prompt that changes on every call, such as a timestamp, an identifier, or tool definitions serialized in a different order. Caching matches the prompt from its first byte, so one early change removes the discount from that point on. Read the cached share of input tokens per call from your traces.
In what order should I apply cost levers to an AI agent?
Price each lever alone on your traces as a share of the current bill, divide by the effort to ship it, and take the highest. Ship it with an eval run, measure again, and re-rank what is left. On this post's illustrative workload the order comes out as caching, caps, batching, model tiering, trimming, then fewer steps.
How do I know a cost cut did not hurt quality?
Run the same fixed evaluation set before and after the change and report two numbers together: units per accepted result and the pass rate. If the pass rate fell by more than its run-to-run noise, the change is a trade, and someone has to decide it on purpose. A lower invoice alone proves nothing about quality.

Sources

  1. Finout (2026). Token Economics and TokenOps: The Definitive Guide to FinOps for Tokens
  2. Banruo Liu et al. (2026). Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale
  3. Shubham Gandhi, Atharva Naik, Yiqing Xie, Carolyn Rose (2025). An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
  4. Lingjiao Chen, Matei Zaharia, James Zou (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
  5. Jasper Dekoninck, Maximilian Baader, Martin Vechev (2024). A Unified Approach to Routing and Cascading for LLMs
  6. Anthropic (2025). How we built our multi-agent research system
  7. leo_e (2025). Hacker News comment on cost per successful task
  8. Majromax (2026). Hacker News comment on model switching and the cache
  9. lynox (2026). Hacker News comment on a broken cache in an agent loop
  10. dennispi (2026). Hacker News comment on cost per accepted outcome