To shadow deploy an AI agent, run the candidate beside the live version on real traffic, show users only the live version’s output, and compare the two in a log. Then canary: route a small share of sessions to the candidate, read both cohorts side by side, and widen only while the numbers hold.
Those two steps are the middle of a five-rung ladder from Chapter 20 of AI Agents, Engineered (in the full book). This post takes one change up all five. By the end you can write a rollout plan in which every rung names what it must prove, its meter, its exit criterion and its way back. The plan is sized to your own traffic and to what your agent’s tools can do.
I mark what is the book’s and what is mine throughout. The book gives the rungs, their order, the canary’s two disciplines and the machinery underneath. It gives no ramp percentages, durations, sample sizes or thresholds, and it is silent on how to shadow deploy an AI agent whose tools change things.
Why does an agent change need a rollout ladder at all?
An agent change needs a ladder because its failures are quiet, late and noisy, so the merge-deploy-watch routine that works for code has nothing to watch. The chapter’s rollout section opens with a scene it calls a composite of the field’s incident reports. A prompt edit merges on Tuesday with a green suite, and support tickets rise on Thursday with nothing to connect the two.
It then separates three reasons. “Behavioral failures are diffuse: bad outputs, not crashes, so nothing pages.” Second: “Feedback is delayed: a support ticket days later, never a 500 at deploy time, so cause and effect drift apart.” Third, outputs vary from run to run on the same input, so one green run proves little.
Infrastructure teams supplied the remedy, which the book calls a rollout ladder: “a sequence of exposures, each rung buying more confidence with a slightly larger slice of reality, with the way back down wired at every height.”
The practice is older than agents. Danilo Sato’s 2014 definition of a canary release describes “slowly rolling out the change to a small subset of users” before everybody gets it. Alec Warner and Štěpán Davidovič, in Google’s Site Reliability Workbook (2018), are more exact: “We define canarying as a partial and time-limited deployment of a change in a service and its evaluation.”
What sits underneath the ladder?
Two pieces of machinery sit underneath: a feature flag that decides which version serves which sessions, and one immutable artifact per version. Both are the book’s.
With the flag, “promotion and rollback are config flips (instant, reversible, no redeploy), and the ramp becomes a sequence of edits to a number.”
The artifact is the bookkeeping. An agent’s behavior comes from code, prompt, model choice and settings, and tool schemas together, so the chapter says to “version the compound as one immutable artifact, and change version one by shipping version one-point-one, never by editing in place.”
That discipline makes the way back real: “rollback means pointing the flag at the previous one, which still exists, unchanged, because nothing could change it.” If someone can edit the live prompt in a console, the previous version is only a memory.
What are the five rungs, and what must each one prove?
The five rungs are the regression gate, the shadow deploy, the canary, the ramp and all traffic, in that order. Each row below says what the rung must prove, which meter proves it, when you may leave it, and how you get back down. B marks a cell that comes from the book; P marks this post’s addition.
| Rung | What it must prove | Meter | Exit criterion | Way back |
|---|---|---|---|---|
| 1. Regression gate | B: the change does not reintroduce failures you have already seen and written down | B: the regression suite, compared as a distribution over repeated runs | B: an enforced pass-rate bar, and no failure on the safety cases (Chapter 16). P: the bar’s value is yours | B: the suite decides the merge; a failing change has not shipped |
| 2. Shadow deploy | B: no regressions on real inputs that no test author would have invented. P: for an agent that acts, that the candidate proposes the same actions | B: a calibrated judge comparing candidate and incumbent answers at volume. P: a diff of proposed tool calls | P: a session count fixed beforehand, a judged-agreement bar, and a person reading every divergent proposed write | P: stop the shadow; users only ever saw the incumbent, provided every write was intercepted |
| 3. Canary | B: the two cohorts’ numbers hold with real users. P: only for drops large enough for the slice to detect | B: five families read per cohort: error and refusal rates, latency percentiles, cost per request, output length at both extremes, human signals. P: plus the task-success signal used to size the slice | B: thresholds chosen beforehand; assignment per session. P: a session count computed from traffic and the smallest drop you must see | B: automated rollback: a breach points the flag back at the incumbent with no person in the path. P: a written rule for runs in flight |
| 4. Ramp | B: the numbers still hold as the slice widens in stages. P: each stage shows a smaller drop or a rarer event than the one before, or holds a wider slice through your feedback lag | B: the same meters, the same two cohorts | P: per stage, a share and a hold counted in sessions, never shorter than your feedback lag | B: the same flag and the same automatic rollback at every stage. P: the same in-flight rule |
| 5. All traffic | P: nothing new; the candidate becomes the incumbent and monitoring takes over | P: the canary’s thresholds kept as alerts, plus trend rules for drift (Chapter 15) | B: the previous artifact still exists, unchanged. P: how long you keep it deployable, written down | B: point the flag at the previous artifact |
One naming note. The book’s figure and the site’s rollout plan generator label rung 1 “the eval gate”; the chapter’s prose calls it the regression gate, and they are one rung.
What does the regression gate prove, and what is it blind to?
The regression gate proves that the change still passes the bank of graded tasks you have collected, and it is blind to everything nobody wrote a case for. The chapter puts the limit in one line: “the gate tests the failures you already imagined. Every rung above it exists because reality imagines better than you do.”
How to build that gate, with paired runs, flaky cases and a zero-tolerance rule for safety cases, is the subject of regression testing for LLM applications.
What does the shadow deploy prove?
The shadow deploy proves how the candidate behaves on real inputs before any user depends on it. The chapter’s definition: “run the candidate beside the incumbent, on real production traffic, with users seeing only the incumbent’s answers.”
A calibrated judge (an LLM-as-a-judge checked against human labels) compares the logged answers at volume. The chapter adds a cheaper first step: “a useful variant runs the candidate over yesterday’s recorded traffic before it touches production at all.”
Its price is that “the shadowed traffic runs twice”, which sets the scope rule: “Shadow the changes with reach (a model swap, a prompt restructuring, a new tool schema) and let small edits climb past this rung on the gate’s word alone.”
What the shadow cannot prove is my addition. No user ever reacts to the candidate, so thumbs, rephrasings and abandonment are missing. And for an agent that acts, no action ever lands.
Chapter 20 says this rung is “also called a dark launch”. Martin Fowler’s 2020 entry on dark launching defines it as “taking a new or changed back-end behavior and calling it from existing users without the users being able to tell it’s being called.”
What does the canary prove?
The canary proves that real users, on a slice small enough to lose, get results no worse than the incumbent’s. The book’s slice is “a percent, sometimes a tenth of one for high-stakes systems, every figure illustrative”, and it holds the candidate to two rules: “Two disciplines separate a canary from a small gamble.”
The first is consistent assignment. A user who lands on the canary stays there for the session, because “assignment per request splices two behaviors into one conversation”.
Automated rollback is the second, on thresholds chosen beforehand, “wired so that a breach flips traffic back to the incumbent with no human in the path.” The chapter’s reason is about people: “rollout incidents surface at night, and nobody makes good judgment calls at 2 a.m.”
Why two cohorts at the same moment? The SRE Workbook chapter answers: “Because time is one of the biggest sources of change in observed metrics, it is difficult to assess degradation of performance with before/after evaluation.” It also advises few meters (“perhaps no more than a dozen”) and one canary at a time.
What do the ramp and all traffic add?
The ramp repeats the canary’s comparison at growing volume, and all traffic ends the comparison. The book says only to “widen the slice in stages only while the numbers hold”, with no percentages and no count of stages.
My rule for the ramp: a stage earns its place when it can show something the stage below could not. That is either a smaller drop in the main signal or a rarer event. A stage that shows neither still limits how many users meet a slow failure. Its hold should be at least your feedback lag, the days a complaint takes to reach you.
At all traffic the candidate becomes the incumbent and the previous artifact stays where the flag can reach it.
How do you shadow deploy an AI agent that takes actions?
You shadow deploy an AI agent that takes actions by sorting its tools: reads run for real, and writes are either recorded without being executed or pointed at a copy. The book does not cover this, so the table and its reasoning are mine.
Google’s SRE authors saw it for ordinary services. Of traffic teeing, their name for shadowing, they write: “Traffic teeing also doesn’t adequately identify risk in stateful systems.”
Mirroring infrastructure assumes a call with no consequences. Two labeled examples, both retrieved on 7 October 2026, show it. The Istio service mesh documentation calls mirrored requests “fire and forget”, with responses discarded. The Amazon SageMaker AI documentation on shadow tests says “Only the responses of the production variant are returned to the calling application.”
An agent breaks that assumption. A Hacker News commenter asked in July 2026 whether teams spin up shadow environments to verify side effects. ananos, writing as an author of the manifesto under discussion, replied “you mostly cannot mirror domain state”, and added: “An external side effect, once committed, is the one thing no sandbox rewinds.”
| Option for a tool in shadow | What it can prove | What it cannot prove |
|---|---|---|
| Read-only tools run for real | The candidate’s retrieval, reasoning and answer on live data | Nothing is lost, but read load on those systems doubles and the candidate sees real data, so the same access rules apply |
| Writes recorded and not executed | Whether the candidate would have taken the same action, with the same arguments, as the incumbent | Anything after the write: the candidate continues from a stand-in result, and the world’s reaction never happens |
| Writes against a sandbox copy | The write’s mechanics: valid arguments, accepted by the system, in the right order | Consequences outside the copy; and the copy drifts from production from the moment it is made |
| Replay of recorded traffic | Behavior on yesterday’s real inputs, before production is touched (the book’s variant) | Today’s state: recorded reads may be stale, and every write must still be recorded or sandboxed |
Recording is my default for any write, and its limit deserves plain words. After the first recorded write, the shadow run is on a fork of reality. A stand-in told the candidate the refund went through, and whatever it does next rests on that. Later steps are weaker evidence.
Recorded writes have one advantage over judged answers. A proposed tool call is structured, so comparing it with the incumbent’s call is a diff of names and arguments that needs no judge. The taxonomy of AI agent failure modes gives names to the divergences you find.
One company has published this shape for its own system. Mike Radzewicz, describing a construction-accounting pipeline in September 2025, defines a shadow system as one that “records outputs without writing side effects to customer systems.”
To shadow deploy an AI agent is therefore to test its decisions and leave their consequences untested. The canary is the first rung where an action lands, which is why approval gates set by consequence stay on through it.
How big should the canary be, and how long should it run?
The canary should run until it has enough sessions to detect the smallest drop you have decided to catch, and its share of traffic follows from that count. The book prints no sample size, so this section is mine. A share picked first, with a duration in hours attached, says nothing about what the canary could have seen.
Start with three decisions: one main signal with a pass or fail per session, the incumbent’s rate on it, and the smallest drop that would make you roll back.
The arithmetic is the standard sample size for comparing two proportions. With p0 the incumbent’s rate, p1 the rate you must detect, and k control sessions for every canary session:
pooled = (p1 + k × p0) / (1 + k)
n = ( z_a × sqrt( pooled × (1 − pooled) × (1 + 1/k) )
+ z_b × sqrt( p1 × (1 − p1) + p0 × (1 − p0) / k ) )²
/ (p0 − p1)²
z_a = 1.96 (two-sided test, 5% false-alarm rate), z_b = 0.8416 (80% power); round n up
With equal groups (k = 1) this is the form in the site’s eval sample size calculator. As the control group grows, it approaches the one-sample formula in the NIST/SEMATECH e-Handbook, section 7.2.4.2. The assumptions are that sessions are independent and that each counts once. Repeat sessions from one user break the first, so count users where sessions cluster.
Take illustrative inputs: 5,000 sessions a day, an incumbent at 90%, and a fall to 85% as the smallest drop worth a rollback.
| Canary share | Canary sessions a day | Canary sessions needed | Days |
|---|---|---|---|
| 0.1% | 5 | 317 | 63.4 |
| 1% | 50 | 320 | 6.4 |
| 5% | 250 | 336 | 1.3 |
| 10% | 500 | 358 | 0.7 |
| 50% | 2,500 | 686 | 0.3 |
Two things follow. The session count barely moves with the share until the groups approach equal size, while the days move two-hundredfold. And a one-percent canary read after one day has 50 sessions; in my simulation of 40,000 such canaries, it flagged the five-point drop 24% of the time.
I checked the formula two ways. Simulated canaries of 320 sessions beside a control 99 times larger detected the drop in 79.7% of 40,000 trials, close to the 80% target. The NIST one-sample formula with the same two-sided z gives 316, and with NIST’s one-sided 5% test and continuity correction it gives 273.
Smaller drops cost far more. A fall from 90% to 87% needs 853 canary sessions at one percent, or 17.1 days. Moving a thumbs-down rate from 10.0% to 10.5% needs 57,764 sessions in each of two equal groups. Tian Pan’s 2026 essay on gradual rollout, which the chapter cites, lands in the same range for “a 5% minimum detectable effect”: “plan for tens of thousands of sessions per arm”.
The calculator below opens on 90 of 100 against 85 of 100. It reports a two-sided p of 0.29 and 686 tasks per version, the equal-groups figure in the table’s last row.
With JavaScript on, the Eval sample-size calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Eval sample-size calculator on its own page to share a result by link.
One more limit, for the failures that matter most. If 320 canary sessions show zero wrongful refunds, the one-sided 95% upper bound on that rate is 1 − 0.05(1/320), which is 0.93%. A 95% chance of seeing even one event that happens once in a thousand sessions takes 2,995 sessions.
What if your traffic is too low for a canary?
With low traffic you widen the first slice, accept a coarser drop, or lean on the shadow, and you write down which. At 300 sessions a day, a one-percent canary gets 3 sessions a day and needs 106.7 days to see the five-point drop.
A 25% slice needs 441 sessions, which is 5.9 days. A half-and-half split needs 686 per group, which is 4.6 days. Settling for a ten-point drop at half-and-half needs 199 per group, or 1.3 days.
Shadowing helps most here. Candidate and incumbent see the same inputs, so the comparison is paired, and paired comparisons need fewer cases; the regression-testing post linked above covers that arithmetic.
Feature flag, automated rollback, kill switch: which does what?
The feature flag routes new sessions to a version, the automated rollback changes the flag when a meter breaches, and the kill switch stops loops that are already running. The book keeps them in different chapters.
| Control | What it acts on | What trips it | Where the book defines it |
|---|---|---|---|
| Feature flag | Which artifact the next session gets | A person, or the rollback, editing configuration | Chapter 20 |
| Automated rollback | The flag | A cohort meter crossing a threshold chosen beforehand | Chapter 20 |
| Kill switch | Every loop and agent running now, on any version | A person, when harm is in progress | Chapter 13 |
The kill switch is in Chapter 13, Writing the Outer Loop (in the full book), which defines it as “one obvious, fast, tested way to stop every running loop and agent at once.” Chapter 20’s rollout section never mentions it, so a plan that says “canary with a kill switch” has put a real control on the wrong job.
What happens to runs that are already in flight?
A flag flip does nothing to a run that has already started, so the plan needs its own rule for those runs. This point is mine; the book is silent on it. Agent runs can be long, and many are accepted as background tasks that outlive the request.
I borrow this figure from the chapter’s section on serving shapes. In the asynchronous shape the work has left the request path, so routing changes never reach it, and candidate runs keep going after a rollback.
Each in-flight candidate run has three possible fates, and reversibility decides among them. A run that has written nothing can be stopped and restarted on the incumbent. A run that has written something can pause at its next checkpoint for a person to look at. Stopping it dead comes last, since a half-done sequence of writes is its own incident.
In a February 2026 Ask HN thread on shutting down misbehaving AI, the commenter zachdotai sorted by the same property: “If no side effects have occurred yet, steer. If the agent already made an irreversible call or wrote bad data, kill.” I differ on the second half, where I pause before I kill: after a breach, no candidate run makes a new write without a person. Checkpoints, retries and budgets inside a run are the ground of how to make AI agents more reliable.
What evidence promotes a change to the next rung?
A change moves up when the rung below has produced its evidence in writing, and the list below is that evidence in order. The items restate the ladder table as an AI agent production checklist for one release.
- Underneath: one version id covers code, prompt, model choice and settings, and tool schemas; nobody can edit a live version.
- Underneath: the previous artifact is deployed and the flag can point at it today.
- Regression gate: the suite passed at the enforced bar over repeated runs, with no safety-case failure.
- Shadow: every tool is sorted as a real read, a recorded write or a sandboxed write.
- Shadow: the session count, fixed beforehand, was reached; the judged comparison met its bar; a person read every divergent proposed write.
- Canary: assignment is per session, and one canary runs at a time.
- Canary: the smallest drop to catch, and the session count it needs at your traffic, were written down before the start.
- Canary: every meter has a threshold set beforehand, rollback is automatic, and the in-flight rule is written.
- Canary: the planned session count was reached with no breach.
- Ramp: each stage has a share and a hold in sessions, no shorter than your feedback lag, and ended with no breach.
- All traffic: the previous artifact stays deployable for a stated period, and the thresholds stay on as alerts with trend rules for drift.
- Kill switch: separate from the flag; where it is, who pulls it and when it was last tested are written down.
A signed copy of this list is the kind of record that an AI agent governance framework for a small company asks a team to keep for each release.
A worked example: a prompt change to a support agent
This example takes one invented change up all five rungs with every blank filled. The agent answers support questions and can issue refunds, and the change restructures its system prompt. Traffic is 5,000 sessions a day, judged task success is 90%, and the team will roll back on a fall to 85%. Every number is illustrative.
The generator below opens on this case: a prompt restructuring, riskiest action irreversible or high-stakes, 5,000 sessions a day, with a suite, a flag and recorded traffic.
With JavaScript on, the Rollout plan generator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Rollout plan generator on its own page to share a result by link.
It shows “5 of 5 rungs required”, since a prompt restructuring is one of the book’s changes with reach. The tool suggests four ramp stages for this tier and raises one warning: a tenth of a percent is 5 sessions a day, under its cut-off of 20.
At 5 sessions a day the canary needs 317 sessions, which is 63.4 days. So the plan departs from the tool’s starting slice and says why: one percent gives 320 sessions in 6.4 days.
Holds count fresh sessions.
| Stage | Share | Hold (sessions) | Days at 5,000 a day | What it can show that the last could not |
|---|---|---|---|---|
| Canary | 1% | 320 | 6.4 | A fall from 90% to 85% |
| 1 | 5% | 893 | 3.6 | A fall to 87% |
| 2 | 10% | 2,995 | 6.0 | At least one event that happens once in 1,000 sessions, with 95% probability |
| 3 | 25% | 2,500 | 2.0 | Nothing finer: stage 2’s 2,995 sessions already show a fall to 88%. It holds a quarter of users through a two-day feedback lag |
| 4 | 50% | 5,000 | 2.0 | Nothing finer; it holds half of users back for a two-day feedback lag |
The canary and ramp take 20.0 days, and the shadow adds 3.0. Each row shows its price: dropping stage 2 saves six days and gives up the look at one-in-a-thousand events and at a fall to 88%.
## Rollout plan: a prompt restructuring (ILLUSTRATIVE EXAMPLE)
Riskiest action: tier 4 (irreversible or high-stakes) · sessions a day: 5,000
Regression suite: yes · feature flag: yes · recorded traffic: yes
Main signal: judged task success, incumbent 90%. Smallest drop to catch: to 85%.
Test: two-sided 5%, 80% power, sessions assumed independent.
### The artifact (one immutable version; ship a new one, never edit in place)
- [ ] Code: ____ · Prompt: ____ · Model choice and its settings: ____ · Tool schemas: ____
- [ ] Version id (a hash over the four): ____ · previous artifact still deployed: ____
### 1. The eval gate (the regression gate): required
- [ ] Suite passes at or above the agreed bar: ____, over repeated runs
- [ ] No failure on the safety cases
- Way back: Do not merge. Nothing has shipped.
### 2. The shadow deploy: required
- [ ] First, replay recorded traffic through the candidate
- [ ] Tools (post's addition): order lookup, policy search = real reads;
customer reply, refund = recorded, never executed; ticket note = sandbox copy
- [ ] A calibrated judge compares candidate and incumbent answers; agreed bar: ____
- [ ] A person reads every proposed refund that differs from the incumbent's
- [ ] Shadow runs for 2,995 sessions (20% of traffic mirrored, 3.0 days):
enough for a 95% chance of one 1-in-1,000 proposal
- Way back: Stop the shadow. Users only ever saw the incumbent's answers.
### 3. The canary: required
Start at 1% (50 sessions a day). A tenth of a percent is 5 a day and would
need 63.4 days; the refund approval gate stays on, so 1% is accepted.
- [ ] Assignment is per session, not per request
- [ ] Rollback is automatic on a breach, with no person in the path
- [ ] Error and refusal rate, latency percentiles, cost per request, output
length at both extremes, human signals: each trips outside the
incumbent's own range over the last four weeks
- [ ] Task success: read once at 320 sessions; roll back at 3.3 points or more
under the incumbent (the test's cut-off; it catches a true 5-point fall
80% of the time)
- [ ] Canary runs for 320 sessions (6.4 days) before widening
- [ ] In flight (post's addition): after a breach, no candidate run makes a new
write without a person; runs with no writes restart on the incumbent
- Way back: Point the flag at the previous artifact: a config flip, no redeploy.
### 4. The ramp: required (4 stages, the tool's count)
- [ ] Stage 1: widen to 5% and hold for 893 sessions (3.6 days)
- [ ] Stage 2: widen to 10% and hold for 2,995 sessions (6.0 days)
- [ ] Stage 3: widen to 25% and hold for 2,500 sessions (2.0 days)
- [ ] Stage 4: widen to 50% and hold for 5,000 sessions (2.0 days)
- [ ] The same thresholds and the same automatic rollback at every stage
- Way back: the same flag flip.
### 5. All traffic: required
- [ ] The previous artifact is kept, unchanged, for ____
- [ ] The thresholds stay wired as monitoring, with trend rules for drift
- Way back: the same flag flip.
Kill switch (Chapter 13, separate from the flag): where ____ · who ____ · last tested ____
Headings and checkbox lines follow the tool’s own Markdown export, plus three additions: the tool-by-tool shadow treatment, the sample arithmetic and the in-flight rule.
Does the ladder hold for a low-traffic, read-only agent?
It holds, with a wider first slice and ramp stages that add exposure and no finer evidence. Take a model swap on a read-only reporting agent at 300 sessions a day, with no recorded traffic. The generator shows 5 of 5 rungs, two ramp stages and two warnings.
Its warnings are the missing recorded traffic and the volume, since a percent is 3 sessions a day. Every tool is a read, so the shadow runs them all for real. The canary starts at 25% and holds for 441 sessions (5.9 days), and stage 1 is half-and-half for 686 sessions per group (4.6 days).
Neither ramp stage shows anything finer. Stage 1 repeats the five-point look, and past half the incumbent becomes the smaller, noisier group, so I fill a 75% stage as a two-day hold for feedback lag. Reads in flight at a rollback finish under the candidate.
What does the rollout ladder not catch?
The ladder does not catch what changes after the top rung, and it does not catch what none of its meters measure.
Drift is the first. A release that climbed cleanly can degrade weeks later with no commit to blame, because users, the hosted model or the documents behind the agent changed. That belongs to Chapter 15, Observability and Debugging (in the full book), whose rule is short: “it is a trend, so alert on trends.” The signals to watch once you ship an AI agent to production are in the post on how to monitor AI agents in production.
The second is the unmeasured failure. A model provider’s postmortem of 17 September 2025 describes three infrastructure bugs that degraded response quality. Its process already included staged exposure: “Engineering teams perform spot checks and deploy to small ‘canary’ groups first.” One routing bug “initially affected 0.8% of requests”. The company’s verdict: “The evaluations we ran simply didn’t capture the degradation users were reporting”.
That incident happened inside a provider, upstream of any agent, and I use it for one point. A fault touching under one percent of requests is what a small slice with a coarse meter will pass. Chapter 20’s closing explains why instruments matter: “An agent’s bad answer arrives plated exactly like a good one.”
Where does this approach stop working?
The approach weakens when traffic is very low (covered above), when no single signal can be scored, and when a first deployment has no incumbent.
My arithmetic assumes one pass-or-fail signal per session. A judged score adds the judge’s own error to sampling noise, so the computed counts are a floor. Looking at the canary every morning also inflates false alarms compared with one read at the planned count.
A first deployment has no incumbent to stand beside. The shadow then compares the agent with the person who does the task today: the agent proposes, the person acts, and agreement is the meter.
Before a change ships, write down what each rung must prove, how many sessions that takes at your traffic, and what happens to runs in flight when the flag flips back.
The ladder itself is the last section of Chapter 20, Deploying and Scaling (in the full book), and the glossary entries for each rung are free to read. The guide to agent security and operations collects the related posts and tools, and you can see the formats.
Questions readers ask
- What is a shadow deployment for an AI agent?
- A shadow deployment runs the candidate version of an agent on real production traffic beside the current version, while users receive only the current version's output. The candidate's output goes to a log, where a calibrated judge or a diff of proposed tool calls compares the two. For an agent that acts, the candidate's writes must be recorded and never executed.
- What is the difference between a shadow deploy and a canary release for AI agents?
- In a shadow deploy no user sees the candidate, so it tests real inputs without real consequences or user reactions. In a canary release a small share of real sessions is served by the candidate, so it tests consequences and reactions, but only differences large enough for that share to detect in the time you give it.
- Is a dark launch the same as a shadow deploy?
- The book treats the two as synonyms. Martin Fowler's 2020 definition of dark launching is calling new back-end behavior from existing users without the users being able to tell, which matches. He also notes that people now use the term for canaries and other partial releases, so say which one you mean in a plan.
- How long should a canary run for an AI agent?
- Until it has seen the number of sessions needed to detect the smallest drop you have decided to catch. With illustrative inputs, detecting a fall from 90% to 85% task success takes about 320 canary sessions against a large control group: 6.4 days at 50 canary sessions a day, and about 63 days at 5 a day.
- Which changes to an AI agent can skip the shadow deploy?
- Small prompt edits can, on the regression gate's result alone. Chapter 20 of the book says to shadow the changes with reach, naming a model swap, a prompt restructuring and a new tool schema. Every change still meets real users on a canary slice first.
Sources
- Alec Warner, Štěpán Davidovič, with Alex Hidalgo, Betsy Beyer, Kyle Smith and Matt Duftler (2018). Canarying Releases, in The Site Reliability Workbook, ch. 16
- Danilo Sato (2014). CanaryRelease (25 June 2014)
- Martin Fowler (2020). DarkLaunching (29 April 2020)
- NIST/SEMATECH (2012). e-Handbook of Statistical Methods, section 7.2.4.2, Sample sizes required
- Tian Pan (2026). Releasing AI Features Without Breaking Production: Shadow Mode, Canary Deployments, and A/B Testing for LLMs (9 April 2026)
- Anthropic (2025). A postmortem of three recent issues (a model provider's postmortem, 17 September 2025)
- Mike Radzewicz (2025). No Blind Spots: How Our Shadow Pipeline Hardens AI Agents in Adaptive (a company's account of its own system, 11 September 2025)
- Istio documentation (2026). Mirroring (one example of a service mesh's traffic-mirroring docs; retrieved 7 October 2026)
- Amazon SageMaker AI documentation (2026). Shadow tests (one example of a cloud ML platform's shadow-testing docs; retrieved 7 October 2026)
- _ananos_, Hacker News (2026). Comment on shadow environments for agents, in the thread The Sandboxing Manifesto for Agentic Execution (21 July 2026)
- nordic_lion, Hacker News (2026). Ask HN: How do you shut down misbehaving AI in production? (13 February 2026)
- zachdotai, Hacker News (2026). Comment on steering versus killing a running agent (February 2026)