What is an AI agent rollout plan?
An AI agent rollout plan is the sequence of steps by which a change to an agent’s behavior reaches users: a little exposure at a time, with a way to undo it at every step. The generator above asks what is changing and what the agent can do, then lays out the ladder from Chapter 20 of the book as a checklist, with every threshold left for you to fill in.
The chapter’s name for the sequence is the rollout ladder: “a sequence of exposures, each rung buying more confidence with a slightly larger slice of reality, with the way back down wired at every height.” The five rungs are an eval gate, a shadow deploy, a canary, a ramp, and all traffic.
Why is shipping a prompt change different from shipping code?
Shipping a prompt change is different because behavior fails quietly, late and inconsistently, and the usual routine of merging, deploying and watching a dashboard catches none of that. The chapter tells the story of a one-line prompt edit that passes its suite on a Tuesday and shows up as a rise in support tickets on Thursday, with nothing to connect the two.
It gives three reasons, “worth keeping apart”. The first: “Behavioral failures are diffuse: bad outputs, not crashes, so nothing pages.” The second is that feedback is delayed, arriving as “a support ticket days later, never a 500 at deploy time”. The third is that outputs vary from run to run, “so any single green run proves little”.
Each rung of the ladder answers one of those. A suite compared as a distribution answers the third. A shadow on real traffic and a canary with cohort meters answer the first two, by looking for bad outputs deliberately before users report them.
What does each rung buy, and what does it cost?
Each rung buys confidence on a larger share of reality and charges for it in spend, time or exposure. The table is the chapter’s ladder in brief.
| Rung | What happens | What it buys | What it costs |
|---|---|---|---|
| 1. Eval gate | The regression suite decides the merge | Protection against failures you have already seen | The suite’s run time and spend |
| 2. Shadow deploy | The candidate runs beside the incumbent on real traffic; users see only the incumbent | Regressions on inputs no test author would write | The shadowed traffic runs twice |
| 3. Canary | A small slice of real users gets the candidate | Evidence from real people, on a slice small enough to lose | A few users see it first |
| 4. Ramp | The slice widens in stages while the numbers hold | The same evidence at growing volume | Waiting at each stage |
| 5. All traffic | The candidate becomes the incumbent | The change, shipped | Nothing new, if the rungs below held |
The first rung is the regression gate from the evaluation chapter, and the chapter states its limit in a sentence: “the gate tests the failures you already imagined. Every rung above it exists because reality imagines better than you do.”
The second is where real inputs come in. The shadow means to “run the candidate beside the incumbent, on real production traffic, with users seeing only the incumbent’s answers.” The candidate’s answers go to a log, where a calibrated judge compares the two at volume. The chapter adds a cheaper first step: “a useful variant runs the candidate over yesterday’s recorded traffic before it touches production at all.” The generator asks whether you have recorded traffic and puts the replay first when you do.
The third is the first time people see the change. “Route a small slice of real traffic to the new version (a percent, sometimes a tenth of one for high-stakes systems, every figure illustrative), watch the two cohorts side by side, and widen the slice in stages only while the numbers hold.”
Which changes need a shadow deploy?
Changes with wide reach need a shadow deploy, and small edits can skip it. The chapter sizes the rung by its price, since shadowed traffic costs double, and draws the line itself: “Shadow the changes with reach (a model swap, a prompt restructuring, a new tool schema) and let small edits climb past this rung on the gate’s word alone.”
The generator follows that rule for the kinds of change the chapter names, and adds two readings of its own, each marked in the result:
- A new tool is treated as a change with reach. The chapter lists a new tool schema; a whole new tool changes what the agent can do at least as much, and the tool counts a changed schema the same way.
- A small edit to an agent that can do irreversible things gets the shadow marked optional instead of skipped. The chapter’s rule would let it pass; the tool notes that at that tier the doubled spend may be worth paying.
The shadow is the only rung the chapter sizes by the change. The canary, the ramp and the gate are required for every change the generator handles, however small, because the chapter’s own cautionary tale is a one-line prompt edit that went straight from a green suite to all traffic.
The tiers are the four from the book’s chapter on oversight, the same ones the consequence tier classifier sorts actions into. At the highest tier the tool also starts the canary at the chapter’s smaller figure, a tenth of a percent, which the chapter gives for “high-stakes systems” without tying it to a tier, and adds ramp stages. How many stages a ramp should have is not something the chapter says; the tool suggests two to four by tier and leaves the percentages blank.
What makes a canary more than a small gamble?
Two disciplines make it more than a gamble: keeping each user on one version, and rolling back automatically. The chapter puts it in those terms: “Two disciplines separate a canary from a small gamble.”
The first is consistent assignment: “a user who lands on the canary stays on it for the session, because assignment per request splices two behaviors into one conversation”. A user who watches the agent’s manner change mid-thread “has found a bug you created by measuring.”
The second is automated rollback, on thresholds set in advance. The chapter lists what to watch, per cohort: “error and refusal rates, latency percentiles, cost per request, output-length distribution at both extremes, and the human signals of thumbs and regenerations and abandonment”. Those are wired “so that a breach flips traffic back to the incumbent with no human in the path.”
The reason for taking the person out is plain: “rollout incidents surface at night, and nobody makes good judgment calls at 2 a.m.” The decision is made earlier, in daylight, and encoded.
The generator lists those five metric families with a blank beside each. It will not suggest a number, because the right one depends on your baseline. If you type your daily sessions, it does one piece of arithmetic: how many sessions the first slice sees in a day. Below twenty sessions a day in the slice, a cut-off that is the tool’s own, the plan warns that the canary will take a long time to say anything. The eval sample size calculator helps decide how many sessions are enough.
What machinery does the ladder depend on?
The ladder depends on two pieces of machinery, a feature flag and versioned artifacts, and the chapter names them “because teams skip them and then cannot afford the climb.”
The flag is the control plane. Which version serves which share of traffic is configuration, so “promotion and rollback are config flips (instant, reversible, no redeploy), and the ramp becomes a sequence of edits to a number.” Answer “no” to the flag question and the way back from the canary upward changes from a flip to a redeploy, with a warning that automated rollback needs the flip.
Versioning is the bookkeeping underneath. “An agent’s behavior is a compound of code, prompt, model choice and its settings, and tool schemas”, and the prompt defines behavior as much as any of the others. So “version the compound as one immutable artifact, and change version one by shipping version one-point-one, never by editing in place.”
The payoff comes during an incident. One artifact is named; a difference in behavior is a difference between two artifacts; and “rollback means pointing the flag at the previous one, which still exists, unchanged, because nothing could change it.” The plan includes the artifact as a block to fill in, one line for each of the four parts and one for the version id.
What does the plan leave to you?
The plan leaves every number and every judgment about your own product to you. It supplies the order, the reasons and the questions; the thresholds, the session counts and the ramp percentages are blanks.
It also cannot check that the machinery it assumes is real. A regression suite that does not cover the behavior you changed is a first rung in name only, and a judge that has not been compared with human verdicts will mislead a shadow comparison. The judge agreement calculator covers the second, and the pass@k calculator the question of how many repeated runs make a suite’s result believable.
Its own choices are labeled where they appear: treating a new tool or a changed schema as a change with reach, the optional shadow for a small edit at the highest tier, the smaller starting slice at that tier, the count of ramp stages, and the low-volume cut-off. Where any of them is more lenient than the chapter, follow the chapter.
The chapter closes by explaining why the whole ladder is made of meters. “An agent’s bad answer arrives plated exactly like a good one.” So “every rung of this ladder is built from instruments instead of instincts”. The guide to agent security and operations places rollout beside reliability and cost, and the full chapter, with serving, durable execution and rate limits before it, is Chapter 20, Deploying and Scaling (in the full book).
Questions readers ask
- Why does a green eval suite not clear a model swap for production?
- Because the suite only tests the failures someone has already thought of. A model swap changes behavior everywhere, including on inputs no test author imagined. Chapter 20 says the gate tests the failures you already imagined, and that every rung above it exists because reality imagines better than you do. The next rung, a shadow deploy, runs the candidate on real traffic to find those.
- Why must canary assignment be per session?
- Because assigning each request separately mixes two versions inside one conversation. The user sees the agent's manner change mid-thread, which is a defect the measurement itself created, and the metrics for the two cohorts stop meaning anything. The chapter's rule is that a user who lands on the canary stays on it for the session.
- What exactly is an immutable artifact for an agent?
- It is one version identifier that covers everything that defines the agent's behavior: the code, the prompt, the model choice and its settings, and the tool schemas. Changing any of them produces a new version; nothing is edited in place. Then an incident points at exactly one artifact, and rollback means pointing traffic at the previous one, which still exists unchanged.
- What is the difference between a shadow deploy and a canary?
- In a shadow deploy the candidate runs beside the current version on real traffic, and users see only the current version's answers; the candidate's go to a log for comparison. In a canary a small share of real users actually receive the candidate's answers. Shadow costs double the model spend and risks nothing in front of users. A canary risks a little, on a slice small enough to lose.
- Does the plan tell me what thresholds to use?
- No. It lists the kinds of metric the chapter says to watch per cohort and leaves every number blank. The right threshold depends on your product and your baseline, and a number supplied by a tool would be false confidence. Choose them before the rollout starts, while nobody is under pressure.
Sources
- Tian Pan (2026). Releasing AI Features Without Breaking Production: Shadow Mode, Canary Deployments, and A/B Testing for LLMs
- TrueFoundry (2026). Agent DevOps: CI/CD, Evals, and Canary Deployments
- Danilo Sato (2014). CanaryRelease
- Pete Hodgson (2017). Feature Toggles (aka Feature Flags)