What is rate limiting for an AI agent?
Rate limiting for an AI agent is the work of keeping a fleet of agent runs under the limits a model provider sets, and deciding what happens to the tasks that do not fit. The simulator above does the capacity arithmetic from Chapter 20 of the book, then runs two hours of traffic against it twice, so you can see what a queue, a concurrency cap and an admission policy each do.
The chapter opens the subject in launch week, when traffic triples and the dashboard fills with refusals from the provider. Every call in the stack has been taught to retry a transient failure, so each refused call waits, retries and rejoins a request rate that was already too high. The chapter’s description of the result is “a wall of synchronized retries that keeps you over the limit indefinitely.” Its summary of why this matters only at scale: “At one run, provider limits are invisible. At a thousand, they are your primary failure mode”.
How many agent runs per minute can you serve?
You can serve as many runs per minute as the provider’s tokens per minute divided by the tokens one run uses. The chapter gives the division in one clause: “tokens per run, which your meters already report, divided into the provider’s tokens per minute gives your ceiling in runs per minute.”
The reason the token axis is the one that matters is the shape of an agent run. “Providers cap on two axes, requests per minute and tokens per minute, and agents hit the token axis first”. Each step re-reads the whole growing transcript, so “one agent run pushes tens of chat-requests’ worth of tokens through the meter”.
The first panel does that division, and repeats it on the request axis, which is the tool’s addition. It opens with a ten-step run of 77,000 tokens, the napkin figure from the book’s cost chapter that the agent cost-per-task estimator reproduces, against limits that are made-up round numbers. Replace them with your own. The chapter is clear that these inputs do not hold still: “limits are commercial terms, and they change”. Its advice is about timing: “do it before launch week, because the alternative is discovering your ceiling as an outage.”
The panel also reports what your concurrency cap allows. A cap of forty runs in flight, each four minutes long, lets through ten runs a minute. Multiplying the ceiling by the run length gives the cap that matches it. That piece of arithmetic is the tool’s.
What do the queue, the cap and backoff each do?
The queue absorbs bursts, the cap controls how much work is in flight, and backoff deals with the refusals that still get through. The chapter lays the three along the path of a request.
- The queue. “A queue in front of the model turns bursts into a steady drain: arrivals are the world’s schedule, the drain is yours, and the queue absorbs the difference.”
- The concurrency cap. “A concurrency cap meters how much work is in flight at once, so the provider sees a controlled stream instead of your raw traffic.”
- Backoff with jitter. “Backoff with jitter handles the refusals that slip through anyway.”
The simulator models the first two. It leaves the third out on purpose: starts are held under the ceiling, so the provider never refuses a call in this model. What a crowd of retries does when that is not true is the subject of the circuit breaker and retry simulator.
A queue alone only delays the problem. If tasks arrive faster than they drain for an hour, the queue grows for an hour. The chapter treats its depth as an instrument. When queueing dominates latency, it says, the problem is capacity, and “a queue filling faster than it drains is that sentence expressed as a number you can alarm on.” The strip chart in the second panel is that number drawn over time.
What is backpressure?
Backpressure is a system telling the work upstream of it to slow down or go away, at the entrance, before it has accepted more than it can finish. The chapter’s definition: “when downstream capacity saturates, the system pushes back upstream, slowing or refusing new work at the entrance instead of accepting everything and collapsing under what it accepted.”
Its picture is a restaurant’s host stand. “The kitchen never turns away a diner; the front door does, and everyone seated eats.” For a service, the equivalent is an admission policy on your own interface: “past a chosen queue depth, refuse new submissions politely, with your own retry-later signal.”
The second panel runs the same traffic under both policies. With the starting numbers, six tasks a minute triple for an hour against a ceiling of ten:
| Accept everything | Refuse past 60 queued | |
|---|---|---|
| Refused at the door | 0 | 420 |
| Accepted and abandoned | 380 | 0 |
| Deepest queue | 180 | 60 |
| Longest wait before a task starts | just over 10 minutes | 6 minutes |
The drain is the same either way, because the door adds no capacity; here 1,060 tasks start under the first policy and 1,020 under the second. The difference is in what the remaining callers are told. The chapter weighs the two outcomes: “Refused at the door costs the caller a retry. Accepted and abandoned costs them their result, their patience, and, because the ticket eventually said failed with nothing to show, their trust.”
The counts come from the tool’s model, which is simple: tasks are a smooth flow, every run takes the same time, and a caller leaves after a fixed wait. One consequence is arithmetic the chapter does not state. A full queue takes its depth divided by the drain rate to empty. Set the depth limit to 150 with the same numbers and a full queue takes fifteen minutes to drain, for callers who leave after ten: 250 accepted tasks are abandoned behind the door. The panel warns when the limit is deeper than patience allows.
Why give interactive and background work separate lanes?
Separate lanes exist because every workload draws on the same allowance, and work that a person is waiting for should not queue behind work nobody is watching. The chapter says that treating them identically “is a choice, and the wrong one.” Its rule: “interactive work takes first claim on the scarce minutes, background work drains in the slack”, and the most patient bulk work goes to a batch tier that the provider schedules into its own spare capacity.
Tick the lanes box and the simulation starts interactive tasks first. With the starting numbers their wait drops to nothing, and the whole cost of the rush lands on the background lane, which waits longer or is refused. The tool gives background work its own, longer patience, sixty minutes to begin with. That number and the strict interactive-first rule are the tool’s. The batch tier is not modeled.
The chapter extends the idea past two lanes: “Cap each lane, each tenant, and each agent” so that one runaway workload cannot starve the rest. This is the fleet-sized version of a per-run budget. The sentence to keep is “A limit you subdivide on purpose fails one lane at a time; a limit you share informally fails everyone at once.”
Why does a long task need a ticket?
A long task needs a ticket because the connection it arrived on will not stay open for as long as the work takes. The opening example of the chapter’s section on serving is a four-minute run behind a load balancer that closes quiet connections at around thirty seconds. The user sees an error, the agent keeps working with nobody to answer, and a retry starts a second run with the same side effects. Raising timeouts does not help, the chapter says, since doing so “merely relocates it”.
The fix is the async job pattern: accept the task, return a task id at once, and do the work on a background worker. The third panel is the ticket’s state machine, which the chapter gives as “pending, then working, then exactly one of completed, failed, or cancelled”, with the rule that “a terminal state, once reached, never changes.” Try to set a completed task back to working and the panel refuses.
Two more buttons submit the task again. Submission is a call that changes something, and the network can drop before the ticket reaches the client, so the chapter requires an idempotency key: “a repeated submission returns the original ticket instead of printing a second one.” Without the key the panel creates a second task, which is the failure the pattern was meant to remove. Idempotent tools and safe retries covers the same mechanism one level down.
The fourth panel applies the chapter’s rule for choosing a shape, whose numbers it calls illustrations. Under about ten seconds, stay synchronous. From ten seconds to an hour, use the queue with polling. Beyond an hour, or with a human approval in the middle, use durable execution. For completion the rule is to “treat polling as the source of truth and webhooks as an optimization.”
What does the simulator leave out?
The simulator leaves out randomness, refused calls and everything downstream of the model. Real arrivals are lumpy and real runs vary in length, so a real queue is deeper at the same average load than the smooth one drawn here. Treat the counts as the shape of each effect.
It also assumes a scheduler that holds starts exactly at the ceiling. A cap alone does not do that: set the cap above what the ceiling sustains and a real system would send the overflow to the provider and get refusals back. The first panel says when your cap is in that position.
Downstream limits are the chapter’s last layer and are not in the model. Tools have rate limits too, and the discipline is to absorb each one inside the tool that owns it. The chapter quotes a practitioner’s statement of the boundary, that the agent “should never need to know that a downstream API has a 100-calls-per-minute limit.”
Getting under the ceiling in the first place is partly a matter of shrinking the run. What caching does to the tokens a run is billed for is in the prompt caching savings calculator, though whether cached tokens count against a rate limit is a commercial term to check. Shipping a change once the plumbing holds is the rollout plan generator.
The chapter ends the section with the test for all of it: “Backpressure done well means failing at the door, in words, instead of in the kitchen, in silence.” The guide to agent security and operations places this beside reliability and cost, and the full chapter is Chapter 20, Deploying and Scaling (in the full book).
Questions readers ask
- Why do agents hit the token limit before the request limit?
- Because each step of an agent's run re-reads the whole transcript so far, and the transcript keeps growing. Chapter 20 says one agent run pushes tens of chat-requests' worth of tokens through the meter, so a modest fleet of agents can use up a token allowance that would serve a large chat product. The request count per run stays small while the token count climbs.
- Why is refusing at the door kinder than queueing everything?
- Because a refusal costs the caller one retry, and a task that is accepted and never started costs much more. The chapter lists what the second caller loses: the result, their patience and, since the ticket eventually said failed with nothing to show, their trust. Accepting everything during a rush does not serve many more tasks: the drain is the same, and most of the extra ones accepted time out before they start.
- Why is polling the source of truth?
- Because a webhook can fail without anyone noticing. Chapter 20 lists how: firewalls drop them, they fail silently when the receiving endpoint is down, they arrive out of order when retried, and they need the client to run a publicly reachable server. Polling is plain and dependable. The two combine well: the client polls until the webhook arrives and then stops.
- How deep should the queue be?
- The chapter gives no number; it says to refuse past a chosen queue depth. The simulator adds one piece of arithmetic of its own: a full queue takes its depth divided by the drain rate to empty, so a depth larger than the drain rate times the caller's patience accepts tasks that will be abandoned anyway. Start from that product and adjust on what your own queue meter shows.
- Does jitter fix a storm of 429 errors?
- No. Jitter spreads retries out so that refused calls do not all come back in the same second, and the chapter calls it treatment rather than cure. If the fleet is asking for more than the limit allows, spreading the requests leaves it over the limit. The cure the chapter names is admitting less work, which is what the queue, the cap and the admission policy do.
Sources
- Tian Pan (2026). Async Agent Workflows: Designing for Long-Running Tasks
- Emanuel Mallia (2026). Building Production AI Agents: Lessons from Real Deployments
- David Yanacek (2019). Using load shedding to avoid overload
- Mark Nottingham and Roy Fielding (2012). Additional HTTP Status Codes (RFC 6585)