Home / Blog / Agents at work / Why an AI Agent Works in Demo and Fails in Production

Agents at work

Why an AI Agent Works in Demo and Fails in Production

Search "AI agent works in demo fails in production" and you get taxonomies. One case shows what a demo measures and nine questions to ask before a date.

By Enrique Gutiérrez · Published · 18 min read

People type “AI agent works in demo fails in production” into a search box after a launch slips or an adoption chart goes flat. The short answer: a demo measures an input somebody chose, run once, with a person watching. Production measures every input users send, on every run, with nobody watching. The agent didn’t change. The measurement did.

I wrote this for a founder or product manager who has just watched a good demo and been asked for a date, or who shipped on a good demo and can’t explain the quiet that followed. It walks one documented public case and one illustrative case with numbers, then hands over nine questions to ask before a date goes into a plan. The engineering that fixes each problem lives in other posts, and I link to them where they apply.

What does “AI agent works in demo fails in production” mean?

It means two measurements were treated as one. A demo and a production deployment differ in three ways that have nothing to do with the model: which inputs are tried, how many times each is run, and whether a person is looking at the result. Each difference hides a kind of failure.

The opening of Chapter 20, “Deploying and Scaling,” (in the full book) puts the first and third in one image: “Cooking dinner for a friend who stands in your kitchen is a demo: one dish, one diner, and if the sauce breaks you both watch it happen.” The table is this post’s own arrangement of the idea.

What a demo measures What production measures What stays hidden in the demo
Input A few requests somebody chose, phrased cleanly, against data that was prepared The whole mix users send, including the unclear, the double and the out-of-scope, against live systems How often real requests look like the chosen ones, and what the agent does with the rest
Runs One run per input Every run, for months, across changes to the prompt, the model and the tools How often the same input gives a different result
Observer The builder and the audience, reading each result as it appears Nobody reads each output; errors surface through complaints, if they surface Who finds out when it is wrong, and what the finding costs

A practitioner on Hacker News described the distance in one sentence in November 2025: “Agentic AI is easy to get to a functionally complete state, but going from functionally complete to reliable is where most teams struggle” (ianmcgraw, who identified himself as an engineer at an evaluation vendor). “Functionally complete” is what a demo shows.

One figure I chose to leave out. Articles on this topic repeat that some large percentage of agent deployments fail. One such article gave its number on the word of “one panelist,” with an explanation attached, and a commenter who quoted the passage objected to “an opinion, with no evidence, as some kind of axiom” (hn_throwaway_99, October 2025); his target was the explanation, and the number rests on the same footing. I found no failure-rate figure whose denominator I could read, so this post carries none.

What happened in one documented public case?

A large restaurant chain tested automated voice ordering at its drive-thrus from 2021, reached more than 100 restaurants, and ended the test in 2024 without widening it. The trade press and a business network reported it from a memo to franchisees, so the facts below are theirs.

Restaurant Business reported in June 2024 that McDonald’s was ending the automated order taking test it had run with IBM and would “remove the technology from the more than 100 restaurants that have been using it,” and that the company was “ending this test without any sort of expansion.” The same article said there had been “questions about whether that technology is ready for prime time, amid concerns about order accuracy.”

CNBC’s report added one detail on the cause: “Two sources familiar with the technology told CNBC that among its challenges, it had issues interpreting different accents and dialects, which affected order accuracy.” The chain declined to comment on accuracy. Its memo said “there have been successes to date,” and Restaurant Business quoted the company saying its work with the partner had “given us the confidence that a voice-ordering solution for drive-thru will be part of our restaurants’ future.”

Read against the table, the reported cause sits in the first row. Accents and dialects are the input mix: the part of reality no prepared set of orders contains until real customers arrive at the speaker. Nothing public says what the system was tried on before the restaurants. What the record shows is a staged test meeting real customers in a hundred sites and ending there, before the other restaurants depended on it.

Three limits on this case. Neither company published an accuracy figure, and the percentages that circulate for this test come from sources I did not verify, so I don’t repeat them. The cause rests on two unnamed sources. The system’s design is not public either, and the test began in 2021, so I don’t call it an agent in this post’s sense; I use it for the shape of the gap and make no claim about how it was built.

The observer row has its own public record. In a 2024 tribunal decision, an airline was ordered to pay a customer C$650.88 after its website chatbot gave him wrong information about a bereavement fare; the tribunal member wrote: “It should be obvious to Air Canada that it is responsible for all the information on its website” (The Guardian, February 2024). The person who found the error was the customer.

What do published measurements say about one run versus many?

Two dated benchmark papers measured what the middle row of the table predicts: scores fall when an agent is run repeatedly or over several turns. Both are snapshots of the systems of their year, and the direction is what to keep.

The authors of τ-bench (Yao et al., 2024) ran agents on simulated customer-service tasks several times each and reported that the agents of the day were “quite inconsistent (pass8 <25% in retail).” Here passk is the share of tasks an agent gets right on all of k attempts, so the figure says fewer than a quarter of retail tasks were solved eight times out of eight. The companion post on pass@k versus passk explains the two metrics.

A 2025 benchmark of business tasks reported that “leading LLM agents achieve only around 58% single-turn success on CRMArena-Pro, with performance dropping significantly to approximately 35% in multi-turn settings” (Huang et al., 2025). A demo is closer to the single-turn setting: one clear request and one answer. A real conversation is closer to the second.

What does the gap look like with numbers?

The case below is an illustrative composite: I invented the company and every number in it to show how the three differences turn into figures a planning meeting can use. Nothing in it is a measurement of a real product.

What did the demo show?

A software company builds an agent for billing support: refunds, plan changes, invoice questions. The builder picks five past tickets, runs each once in front of the leadership team, and all five come out right. Someone asks when it can ship, and a date six weeks out goes into the plan.

What is five out of five evidence of? As arithmetic, the 95% Wilson interval for five passes in five trials runs from 56.6% to 100%. So even if the five tickets had been drawn at random, a true pass rate just above one in two would still be compatible with the demo. They were not drawn at random, which means the interval describes tickets like the five that were picked and nothing else.

What did 200 real tickets show?

The agent launches on all billing tickets. Nothing crashes. Ten weeks later reopened tickets are up, and the support team has started redoing the agent’s tickets without telling anyone.

The post-mortem does what the demo skipped. A support lead takes 200 unedited tickets from the month before launch, runs each through the agent once, and grades the outcome against what the billing system should show. The result is 156 passes, which is 78%. The 95% Wilson interval is 71.8% to 83.2%.

With JavaScript on, the Eval sample-size calculator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Eval sample-size calculator on its own page to share a result by link.

At a volume of 1,400 billing tickets a month, a 22% failure rate is about 308 wrong tickets a month, or roughly 15 per working day over 21 working days. The interval makes that a range of about 235 to 395 a month. No dashboard showed it, because a wrong answer is a completed ticket.

Customers see a different number again. If tickets failed independently, which is a simplifying assumption, a customer who opens five billing tickets in a year would get all five handled correctly with probability 0.78⁵, about 28.9%. The pass@k calculator does that sum for any rate and count.

Where did the 44 failures come from?

The support lead runs each of the 44 failures a second time and sorts them by cause. The four groups map onto the three differences, and each one names the question, from the list further down, that would have surfaced it before the date was promised.

Cause (illustrative) Failures Share of the 44 Share of the 200 Difference Question that surfaces it
Request unlike the demo’s: two issues in one ticket, missing details, unclear intent 17 38.6% 8.5% Input 1 and 3
Wrong or missing data: a stale policy page, an account record absent from the customer system 12 27.3% 6.0% Input 5
Passes when run again: the same ticket succeeds on a second attempt 8 18.2% 4.0% Runs 2
Said done, wasn’t: the reply confirms a refund and the billing system shows none 7 15.9% 3.5% Observer 6

The second row is the one the book warns about by name. In the section “Business and Operational Agents” of Chapter 23, Research and Business Agents (in the full book), the sentence reads: “it surprises people arriving from the demos: the hard part is the plumbing.” A demo runs against prepared data, and the stale policy page is not in it.

The last row is the smallest and the worst. At 3.5% of 1,400 tickets, about 49 customers a month are told a refund was issued when it wasn’t. Naming what each of these is called, and how to read it in a run’s record, is the job of the field taxonomy of AI agent failure modes.

Why did nobody notice for ten weeks?

Because the failure was a quality problem with no crash, and the people who saw it responded by working around it. Chapter 20’s section “The Rollout Ladder” gives the mechanism: “Behavioral failures are diffuse: bad outputs, not crashes, so nothing pages. Feedback is delayed: a support ticket days later, never a 500 at deploy time, so cause and effect drift apart.”

Another practitioner put the same thing more bluntly: “Most agent failures are silent. Most failures occur in components that showed zero issues during testing” (Mesterniz, Hacker News, December 2025). His “most” is one person’s experience and I can’t put a number on it.

The support team’s reaction has a name too. The section “Trust Calibration” in Chapter 24, Agent UX and Human Trust (in the full book) describes two ways reliance goes wrong. Over-trust is relying on a system more than it deserves. Under-trust is the mirror image, and the chapter says what it produces: “the agent’s good work is re-done by hand, the savings evaporate, and eventually the system is abandoned.” A demo creates the first; weeks of unexplained errors create the second.

The target of trust calibration.
Figure 24.1 The target of trust calibration. Reliance should track the agent’s actual reliability—the diagonal, in accent. Above it lies over-trust, where reliance outruns reliability and mistakes ship; below it, under-trust, where good work is redone by hand until someone switches the agent off. The goal is to match trust to reliability, not to maximize it. Reuse this diagram

In the illustrative case the team moved from the over-trust side of that figure to the under-trust side with no incident in between. That is what quiet rejection looks like from the inside: the agent keeps running and the people around it stop relying on it. The goal the chapter sets is appropriate reliance, which needs evidence about reliability that neither the demo nor the launch supplied.

What is the verification gap, and why does a demo hide it?

The verification gap is the book’s name for “the distance between how convincing research output looks and how cheaply it can be checked.” Chapter 23 defines it for research agents in the section “The Verification Gap” and adds: “The gap is a property of the domain.” Applying it to every kind of agent output is this post’s extension.

A demo hides the gap because the audience closes it for free. Five people who know the billing policy read five replies, and the checking costs nothing anyone counts. In production the same check has to be paid for on every ticket or skipped. The seven “said done, wasn’t” tickets are what skipping looks like: the reply was, in Chapter 23’s phrase, “confident, well-formatted, wrong.”

The “Trust Calibration” section of Chapter 24 opens with the reason a reader accepts such a reply. It lists the properties of a calm, specific, competent-sounding status message and concludes: “Every one of those properties raises your willingness to click approve, and none of them is evidence.”

So the cost of checking belongs in the business case. Chapter 23’s section “Framing Return on Investment Against Risk” says it flatly: “An agent whose output costs as much to verify as to produce by hand has a return of approximately nothing, however impressive the demo.”

Here is that sentence as illustrative arithmetic on the 200 tickets. Suppose a person handles a ticket in 6 minutes and reviews an agent-handled one in 3. People alone take 20 hours. With the agent, reviewing all 200 takes 10 hours and redoing the 44 failures takes 4.4 hours, for 14.4 hours in total. The saving is 5.6 hours, 28% of the original work. A plan that assumed the agent removed all 20 hours was pricing the demo.

Review of every ticket is the expensive end, and it is where a team that has lost trust ends up. The cheaper designs check the outcome in the system of record with code, and send a sample to a person. Which of those is possible depends on the task, which is why the gap is a property of the domain.

What should you ask before promising a date?

Ask nine questions, grouped by the three differences, and treat any that has no answer as the next piece of work. The list is this post’s own; each question names the row of the first table it probes.

  • 1. Who chose the demo’s inputs, and what share of real requests look like them? (Input)
  • 2. How many times was each input run, and did the result ever change? (Runs)
  • 3. What is the pass rate on an unedited sample of real requests, with its interval, and who decided what counts as a pass? (Input)
  • 4. What caused each failure in that sample? (Input, Runs)
  • 5. Which live systems and documents did the demo fake, prepare or skip, and who owns their quality? (Input)
  • 6. When the agent is wrong in production, who finds out, by what means, and how long after? (Observer)
  • 7. What does checking one output cost, and is that cost in the business case? (Observer)
  • 8. What is the worst thing the agent can do with nobody watching, and what limits it? (Observer)
  • 9. What is the way back, and which numbers, written down now, widen the rollout or stop it? (Input, Observer)

In the illustrative case, question 1 had the answer “the builder, and nobody knows.” Questions 2 through 9 had no answer on the day the date was promised.

That suggests a rule for the date itself, which is mine and simple on purpose. From a demo, promise a measurement date: the day questions 3, 4 and 7 will have answers from a real sample. State the launch as a condition on those answers, with the thresholds written before the sample is graded. The thresholds are yours to set, because they depend on what a wrong output costs in your product.

The rule changes what a demo is for. A good demo justifies paying for the measurement. Chapter 23 gives the matching advice on where to launch first: internal and low-stakes, because there “the audience reports failures instead of churning.” It also names the opposite order: “Run the sequence in reverse (customer-facing first, because that is where the demo impressed) and you are betting an unproven governance stack on your most visible workflow.”

The product decisions that sit around this one are the subject of AI agents for product managers, and if the agent in the demo belongs to a vendor, the same questions apply to the build-versus-buy decision for AI agents. A buyer can go further with the fifteen reliability questions for an enterprise vendor.

Which checks close the gap?

One check sits behind each question, and none of them needs a new model. The table states each in plain terms with what it produces and what it costs; the buttons filter by the difference a check closes. The costs are my judgment except where the book is named.

Question The check, in plain terms What you get What it costs Closes
1 Pull an unedited sample of past requests and mark which ones resemble the demo’s The share of real work the demo represented An afternoon with the request history input
2 Run each sampled request several times A rate per request, in place of a single pass or fail Several times the run cost of the sample runs
3 Have someone who knows the right answer grade the sample against a written definition of “pass” A pass rate with an interval; the start of an eval set A domain expert’s days; the definition is the hard part input
4 Sort every failure into a cause A fix list ordered by size Hours per few dozen failures input, runs
5 List every system and document the agent reads or writes in production, with an owner for each The plumbing work the demo skipped Integration and content cleanup, which can exceed the agent work input
6 Compare what the agent said it did with what the system of record shows, in code where possible, and send a sample of live runs to a person Wrong-and-unnoticed outputs, found by you before the customer Integration work, plus reviewer hours every week observer
7 Time a reviewer on a sample and put the figure in the business case The real saving, after checking An hour of timing; an honest spreadsheet observer
8 Cap what the agent can do unattended and route irreversible actions to a person A worst case with a known size Engineering per action; some waiting observer
9 Run it beside the people first with its output held back, then on a small slice, with stop conditions written in advance Performance on live traffic before anyone depends on it The held-back traffic runs twice, as Chapter 20 notes input, observer

Rows 3 and 9 are rungs of what Chapter 20 calls the rollout ladder. Row 3 is its first rung, and the chapter marks the limit of that rung in two sentences: “the gate tests the failures you already imagined. Every rung above it exists because reality imagines better than you do.” Row 9 is the shadow deploy followed by the canary, a small slice of real traffic with a way back.

The rollout ladder.
Figure 20.5 The rollout ladder. A change climbs an ascending staircase of gates—an eval gate in CI, then a shadow on real traffic, a small canary behind a feature flag, a ramp, and finally all traffic. The promotion path is the accented arrow; the dashed rollback arrow is wired back to the base before you need it, always available to slide a change down again. Reuse this diagram

The how-to for each row lives elsewhere. Sample size for row 3 is worked out in how many eval examples you need. Row 6 after launch is how to monitor AI agents in production, and its founder-level version is the planned post on how you know if an AI agent is working. Row 8 is approval gates by consequence and AI agent guardrails.

The repairs that follow from row 4 are in how to make AI agents more reliable. The rollout plan generator drafts row 9. Why a longer task fails more than a short one is the subject of why agent errors compound and of a short animated explainer.

How do two other agents read against the nine questions?

Two different agents, walked through the same list, end at the same rule with different questions deciding. Both are sketches with no numbers.

An internal agent that answers employees’ questions about company policy. Question 5 decides it: the agent reads the policy documents as they are, including the stale page and the two that contradict each other. Question 6 has a kinder answer than in the billing case, since colleagues report a wrong answer. Question 8 is small, because the agent only writes text to an employee. The measurement date is close and the launch condition is mostly about document cleanup.

An agent that writes market-research briefs for a planning team. Questions 6 and 7 decide it. Nothing in the system of record can confirm a claim about a market, so row 6 has no code version, and the check is a person opening the cited sources. Question 8 has no cap either: Chapter 23 argues that a research agent’s output “keeps a human in the loop at any level of measured skill.” The launch condition is a reviewing step whose cost, from question 7, still leaves a saving.

In both sketches a demo would have looked the same as it did for billing: a clean question and a fluent answer. What the list changes is which missing answer the team goes to find first.

Where is this argument weak?

It is weak in four places. First, the main worked case is invented. Its proportions are there to show the method; a real sample could split very differently, and the split is the thing to find out.

Second, 78% is not a verdict. A product that drafts replies for a person to send can be useful at that rate, because the person is the check and question 7 prices it. A product that issues refunds unattended is a different matter at the same number.

Third, a real sample ages. The requests change, the documents change, and a prompt or model change resets what the sample told you. The nine questions are worth asking again at every change that reaches users.

Fourth, the public case supports less than I would like. It shows a wide test that was not widened and a reported cause in the input mix. It doesn’t show the numbers, and I found no public post-mortem of an agent launch that gives the demo, the sample and the failure causes together.

The sentence to bring to the planning meeting

If you searched “AI agent works in demo fails in production” because it describes your project: the agent didn’t get worse between the two. The demo answered “can it do this once, on this input, while we watch?” and the plan needed “how often, on what users send, when we don’t?” Whether the saving survives the cost of checking is the subject of the planned post on AI agent ROI.

A demo is evidence that the task is possible. Give a date for the measurement, and let the measurement give the date for the launch.

Chapter 20, “Deploying and Scaling,” develops the rollout ladder and the composite incident this post borrows its mechanism from (in the full book); Chapter 23, Research and Business Agents (in the full book) covers the verification gap and where to deploy first, and Chapter 24, Agent UX and Human Trust (in the full book) covers trust calibration. The Agents at work guide collects the related posts and tools, and you can see the formats.

Questions readers ask

Why does an AI agent work in a demo and fail in production?
A demo and production measure different things. The demo measures an input that somebody chose, run once, while a person watches the result. Production measures the full mix of inputs users send, on every run, with nobody checking each output. An agent can be unchanged between the two and still score very differently.
How many demo runs are enough to trust an AI agent?
No number of demo runs on chosen inputs answers the question, because the inputs were chosen. As a matter of arithmetic, five passes out of five allow a true pass rate as low as 56.6% at 95% confidence (Wilson interval). A sample of unedited real inputs, graded by someone who knows the right answer, is the evidence a launch decision needs.
What is the verification gap?
The book's Chapter 23 defines it for research agents as the distance between how convincing research output looks and how cheaply it can be checked, and calls it a property of the domain. This post applies the same idea to any agent: before launch, ask who checks each output in production and what that check costs.
What should a founder ask before promising a ship date for an AI agent?
Nine questions in three groups. About the input: who chose the demo's inputs, what the pass rate is on unedited real ones, what caused each failure, and which systems the demo faked. About runs: how many times each input was run. About the observer: who finds out when it is wrong, what a check costs, what caps the worst action, and what the way back is.
Is a successful demo worth anything?
Yes. A demo shows the task is possible with the current model, tools and data on at least one input, which is real information and the right basis for funding a measurement. It is evidence of possibility. It says little about how often the agent succeeds on the inputs users will send.

Sources

  1. Jonathan Maze, Restaurant Business (2024). McDonald's is ending its drive-thru AI test
  2. Kate Rogers, CNBC (2024). McDonald's to end AI drive-thru test with IBM
  3. Shunyu Yao et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
  4. Kung-Hsiang Huang et al. (2025). CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
  5. Leyland Cecco, The Guardian (2024). Air Canada ordered to pay customer who was misled by airline's chatbot
  6. ianmcgraw (2025). Hacker News comment on moving agents from functionally complete to reliable
  7. Mesterniz (2025). Hacker News comment on silent agent failures
  8. hn_throwaway_99 (2025). Hacker News comment on an article's panelist-sourced failure figure