How do you know if an AI agent is working? Usage, uptime and the agent’s own “done” can’t tell you, because all three stay the same while it fails. You know from evidence the agent did not write: a weekly read of a random sample, the system of record, human corrections, reversals downstream, and a planted task.
I wrote this for a founder or product manager who is accountable for an agent and does not read its code. A silent failure is a run that ends with no error and a wrong or missing result. This post argues that silence is the normal way an agent fails, names six kinds, and gives five checks you can ask for and read yourself.
The engineer’s version of the same question, with alert rules and a dashboard spec, is the post on how to monitor AI agents in production. The numbers in my worked example are illustrative, and the grouping into six kinds and five checks is mine, built on Chapter 15 and Chapter 16 of AI Agents, Engineered (both in the full book).
Why does a failing agent look like it is working?
A failing agent looks like a working one because the four things people watch all describe the run, and none describes the result. The agent returns fluent text, it reports success, no error is raised, and the counts on the dashboard record a run as handled whether or not it was handled well.
Chapter 15 opens its section “Production Monitoring” with the scene. Every dashboard is green while the agent answers billing questions from last quarter’s refund policy: “No exception was thrown. No status code lied; the runs really did complete.” A few pages earlier the same chapter gives the two views side by side: “Your monitoring saw a healthy service. Your user saw the product fail.”
Fluency is the first reason. Chapter 23 (in the full book), in the section “The Verification Gap,” says fluency and confidence are produced with equal ease whether the content is right or wrong, “so the failure arrives dressed identically to success.”
The fourth reason is the one a founder can act on. A product announcement from December 2025, which asks “how do you know if it’s actually doing what it’s supposed to do?”, offers three health metrics as one vendor’s example: agent error rate, interaction latency and escalation rate. Think about a wrong answer delivered quickly with no handoff. It moves none of the three.
A support-QA vendor’s blog made the sharper point in July 2026: “The resolution flag, if it exists, was set by the agent itself, which believed it succeeded.” A completion rate built on that flag is the agent’s opinion of itself, counted.
What does the evidence say about false “done”?
Published measurements show that an agent claiming success while the world disagrees is common on benchmarks, and I found no measured rate for production agents in general. Read the three figures below for their shape, with their scope attached.
Laksh Advani’s 2026 preprint opens: “LLM agents can fail silently by asserting task completion when the environment state shows otherwise.” Across 9,876 recorded runs on one benchmark and 1,879 on a second, its abstract reports that such false success made up 45% to 48% of failures in some settings, 3% in another, and 75.8% among coding-agent runs that made an explicit status claim. The first two are shares of failures, the third is a share of the runs it names, and all three come from one author’s study.
Cao, Driouich and Thomas (2026) looked at the other side, the runs a benchmark had credited as successes. Their abstract reports that “27-78% of benchmark reported successes are corrupt successes concealing violations across interaction and integrity”. The task was scored complete and a rule was broken on the way.
Practitioners describe the same thing. One wrote in October 2025 of an assistant that reported an email as sent: “the email was never sent. The attachment was wrong. Nobody told me.” Another in March 2026: “agents report success but system state tells a different story.” A sandbox vendor’s own unaudited count (February 2026) put misreported endings at “about 12% of sessions with error states.”
What kinds of silent failure are there?
I count six kinds of silent failure, and they differ in which check can see them. The examples come from one imaginary agent that reschedules home-repair appointments, so that you can compare them. The buttons narrow the table to the failures one check catches.
| Kind | Plain example | Why nothing on the dashboard moves | Caught by |
|---|---|---|---|
| 1. False done | The agent tells the customer the visit is moved to Thursday; the booking still says Tuesday | The run finished and the agent marked it done | read, record, corrections, downstream |
| 2. Confident wrong answer | The agent quotes last year’s cancellation fee | A reply was sent; no field records whether it was true | read, downstream |
| 3. Right outcome, broken rule | The visit is moved, by double-booking a technician | The outcome the customer asked for exists | read, corrections |
| 4. Kept a case it should have handed over | A customer reports a gas smell and gets a polite reschedule | A handoff that never happened leaves no event to count | read, downstream |
| 5. Slow drift | The share of wrong fees creeps up for a month after a policy page is reorganized | Each day looks like the day before | read, corrections, downstream |
| 6. Work that never happened | The intake stops feeding the agent; it runs on nothing and reports no problems | Zero requests fail when zero requests arrive | planted, downstream |
Kind four is the Kaizo page’s “unsafe non-escalation,” where “no escalation event is ever logged”; the other names and the grouping are mine. Kind six comes from an operator’s notes on four incidents (September 2026), whose first row reads: exit code 0, scheduler reported success, and “The upstream source had been switched off. Seven weeks.” The same notes observe that a pipeline that processes zero records looks, by those measures, like one processing thousands.
These are not the engineer’s categories. The post on AI agent failure modes sorts failures by what went wrong inside the run. This table sorts them by what a person outside the run can see.
Why is the agent’s own report not evidence?
The agent’s report is not evidence because the agent writes it with the same machinery that produced the mistake, so it can say done whether or not anything happened. Chapter 16’s section “Building an Eval Set” puts it in one sentence: “An agent that says “Done! The meeting is scheduled” has produced a sentence, and sentences are cheap; query the calendar, run the generated code, read back the database row.”
Chapter 15 gives the rule for reading any run. “In a trace, the tool calls and results are ground truth about what happened; the narration is testimony about why. Read the testimony; verify against the record.” A trace is the book’s word for the complete record of one run.
Asking a second model to read the transcript does not fix this on its own. Advani’s abstract reports that model judges “rely on surface completion proxies”, confident closing language among them, instead of checking what changed. The post on agent trajectory evaluation covers that result and how engineers grade a run against the state of the world.
So here is the test I apply to every check a team proposes: would this check still pass if the agent were wrong in this way? If yes, it is not a check on that failure. The operator’s notes quoted above say the same: “A check is only useful to the extent that it is independent of the failure mode it is meant to detect.”
Why does the domain decide how hard this is?
The domain decides because some work comes with a cheap independent check and some does not. The book calls the difference the verification gap: “the distance between how convincing research output looks and how cheaply it can be checked. The gap is a property of the domain.”
A rescheduling agent sits on the easy side. The booking system is an oracle, a source of truth outside the model, and comparing it with the agent’s claim takes one lookup. A research agent that writes a market summary has no such record, so checking one output can take as long as writing it. Chapter 23 defines the gap for research agents; applying it to every agent is my extension, as it is in the post on why an AI agent works in a demo and fails in production.
How do you know if an AI agent is working? Five checks you can read
You know by running five checks that do not depend on the agent’s account of itself. Each sees some of the six kinds and is blind to others, so the table lists both. You can read all five without engineering help, though the second and fifth need an engineer to set up once.
| Check | What you ask for | Catches | Misses |
|---|---|---|---|
| Weekly read | A random sample of finished runs, read against a written rule | Kinds 1 to 5 | Kind 6; anything rarer than the sample can show |
| System of record | For every run the agent marked done, does the record match what it told the customer? | Kind 1 | Kinds 2 and 4; tasks where doing nothing leaves the record correct |
| Corrections and redos | How many of the agent’s results did a person edit or redo within a day? | Kinds 1, 3 and 5, where a person sees the result | Whatever nobody looks at; it also falls when people stop checking |
| Complaints and reversals | Repeat contacts, refunds, reopened cases, cancellations within a set window | Kinds 1, 2, 4, 5 and 6, late | Kind 3; customers who leave without telling you |
| Planted task | A fixed request with a known right answer, sent on a schedule | Kind 6 and outright breakage | Partial failure; the planted case is usually an easy one |
What is the weekly read, and what rule does it need?
The weekly read is a person who knows the work reading a random sample of finished runs and marking each pass or bad against a rule written down beforehand. It is the only check of the five that can see a failure nobody has named yet.
Chapter 15 defends the habit: “metrics tell you that quality dipped, and only the trace shows you the new failure you had no name for yet”. Its efficient version reads the worst ten runs by cost or step count, which is right for finding bugs. To estimate a rate you need a random draw, because a worst-first sample overstates it. Husain and Shankar’s evals FAQ gives a second reason to keep one: “Keep some random traces in every batch. This gives you a chance to find failure modes that your current signals do not describe.”
The rule has to be one that two readers would apply the same way. Chapter 16’s standard: “if two domain experts, given your criterion and an output, can disagree in good faith about whether it passed, sharpen the criterion, not the agent.” For the rescheduling agent, a run passes when the booking matches what the customer was told, every stated fact matches current policy, no scheduling rule was broken, and anything on the handover list was handed over.
What do the other four checks add?
The other four cover what a sample cannot: every run, the runs nobody sampled, and the runs that never happened. The system of record check is code comparing each claimed result with the booking, the ledger or the sent folder. It covers every run and it sees only what the record holds.
Corrections and redos come free wherever people work beside the agent, and Chapter 16 names the customer’s version: “A user who immediately rephrases the same request has graded the first answer, for free.” Read this number with care. A falling correction rate can mean a better agent or people who have stopped checking.
Complaints and reversals catch nearly everything, and they are the slowest and most expensive check, since a customer has already paid for the failure. Compare them with the rate before the agent took over the work. A planted task is a fixed request with a known answer, sent daily. Some teams call it a canary; the book keeps that word for a rollout step, so I use planted.
How many runs should you read?
Read 50 random finished runs a week where volume allows, and 100 when a decision is close, because a clean sample only bounds the failure rate from above. The figures below are 95% Wilson score intervals, a standard way to put a range around a proportion.
| Runs read, none bad | The failure rate could still be as high as |
|---|---|
| 20 | 16.1% |
| 50 | 7.1% |
| 100 | 3.7% |
| 300 | 1.3% |
Twenty clean runs feel reassuring and prove little. If the failure you fear happens once in fifty runs, a read of twenty finds none about two times in three (0.98 to the power 20 is 0.67). Husain and Shankar’s minimum setup is to “Spend 30 minutes manually reviewing 20-50 LLM outputs whenever you make significant changes.” A read of that size finds the kinds of failure; it is too small to hold a rate to a limit.
If the agent finishes fewer than about 50 runs a week, read all of them. The arithmetic behind these counts, and what changes when you compare two versions, is in how many eval examples you need.
What number should make you stop the agent?
Stop the agent when the evidence says its failure rate is above a limit you wrote down before launch, or when one run does something on your never list. Stopping here means moving the agent back to drafting for a person to approve, one step down the autonomy slider, until the cause is found.
The rule has three parts, and it is mine:
- Never-events stop at once. Write the short list of results that one occurrence makes intolerable, such as one customer’s address sent to another. No sample size applies.
- At or above the stop count, stop. The whole 95% interval for the failure rate is above your limit.
- At or below the clear count, continue. The whole interval is below your limit. Anything between the two counts means the sample cannot decide: read another batch and add the counts.
| Runs read | Limit 5% | Limit 10% | Limit 20% |
|---|---|---|---|
| 50 | stop at 6 or more bad; never clear | stop at 10 or more; clear at 0 | stop at 16 or more; clear at 4 or fewer |
| 100 | stop at 10 or more; clear at 0 | stop at 16 or more; clear at 4 or fewer | stop at 28 or more; clear at 12 or fewer |
“Never clear” is the honest cell. Fifty runs cannot show that a failure rate is under 5%, since even a perfect sample leaves room for 7.1%. A tight limit needs a bigger sample or a check that covers every run.
The limit itself is a business judgment about what a failure costs and who bears it. Chapter 16 warns against judging by the average alone, calling it “exactly the wrong summary for a system whose product is its worst runs”, which is why the never list sits above the counts.
Worked example: a week that looked fine
In this illustrative week the agent’s own number, the record and a human reader give three different answers, and only the third triggers the stop rule. The agent reschedules repair visits. The team wrote a 10% limit before launch, reasoning that a bad reschedule costs a phone call and a wasted visit.
The dashboard shows 1,200 requests, no errors, full uptime, and 1,164 runs marked done by the agent: 97.0%. Chapter 15 has a name for totals like requests handled, vanity metrics, which grow “while saying nothing about whether anyone was served well.”
The five checks say this:
- System of record. Of the 1,164 runs marked done, the booking disagrees with what the customer was told in 70, or 6.0%. Confirmed outcomes are 1,094 of the 1,200 requests: 91.2%.
- Corrections. Dispatchers edited 85 of the 1,164 bookings within a day, or 7.3%.
- Downstream. 132 of the 1,200 customers made contact again within 48 hours, 11.0%, against 5% before the agent.
- Planted task. It passed on 7 of 7 days. The planted request is a simple one, and simple requests were fine.
- Weekly read. Of 50 random runs marked done, 6 were bad: three false dones, two wrong fees, one gas-smell message not handed over.
Six bad in 50 is 12.0%, with an interval from 5.6% to 23.8%. Against a 10% limit the table says stop at 10 and clear at 0, so six decides nothing. The team reads 50 more the next week and finds 11 bad.
Added together, that is 17 bad in 100. The stop count for 100 runs at a 10% limit is 16. The calculator below opens on the same numbers stated as passes, 83 of 100, and gives a pass rate between 74.5% and 89.1%; the failure rate is therefore between 10.9% and 25.5%, all of it above the limit.
With JavaScript on, the Eval sample-size calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Eval sample-size calculator on its own page to share a result by link.
The agent goes back to drafting for dispatchers to approve. Notice what each number would have told a founder alone: 97.0% from the agent, 91.2% from the record, 83% from a reader. The record check was accurate and incomplete, because wrong fees and missed handovers leave the booking correct.
What goes on the weekly read sheet?
The sheet fits on one page and records the rule, the sample, the counts and the decision, so that next week’s reader applies the same standard. Bracketed values are yours.
WEEKLY READ: [agent name] · week of [date] · reader: [name, knows the work, did not build it]
WRITTEN BEFORE LAUNCH
Failure limit: [10]% of finished runs · set on [date] by [name]
Never list (one occurrence stops the agent): [ ... ]
A run passes when: (1) the record matches what the user was told
(2) every stated fact matches current policy
(3) no rule on the rule list was broken
(4) anything on the handover list was handed over
THIS WEEK'S SAMPLE
Finished runs: [n] · read: [50] chosen at random by [method], not by the agent's builder
Bad runs: [ ] of [50] false done [ ] · wrong answer [ ] · broken rule [ ]
not handed over [ ] · other, new kind [ ]
Never-events seen: [ ]
THE OTHER FOUR CHECKS
Agent says done: [ ] of [ ] runs
Record confirms: [ ] of [ ] runs marked done
Edited or redone in a day: [ ] of [ ] last week: [ ]
Repeat contact / reversal: [ ] of [ ] before the agent: [ ]%
Planted task: passed [ ] of [ ] sends
DECISION (from the count table for this limit and sample size)
[ ] Never-event seen -> stop now
[ ] Bad count at or above stop count -> stop: agent drafts, a person approves
[ ] Bad count at or below clear count -> continue
[ ] In between -> read [50] more next week and add the counts
Each bad run filed as a test case with: [engineer]
The last line is where this sheet meets engineering. Each bad run becomes a case in the test set, and the gate described in regression testing for LLM applications keeps the same failure from shipping again. Chapter 16 is honest about that gate too: “a green suite is evidence, never proof.”
Where does this approach break?
It breaks first where reading is expensive. Fifty reschedules at two minutes each are 100 minutes a week; fifty research reports could be a week of work, so in a domain with a wide verification gap the sample shrinks and a person should review each output before anything irreversible depends on it.
The counts assume random, independent runs and an unchanged agent. Runs from one customer’s long conversation count for less than their number, and adding two weeks together is only fair if no prompt, model or policy changed between them. If one did, start the count again.
A sample cannot see rare events. A failure that happens once in a thousand runs will not appear in 50, which is why the never list needs a check that covers every run, written by an engineer. And the reader matters: someone who built the agent, or who does not know the work, will pass runs a dispatcher would fail.
The evidence has limits too. The benchmark figures above measure particular agents on particular tasks in 2026. A vendor’s survey of 1,340 practitioners at the end of 2025 found that “Nearly 89% of respondents have implemented observability for their agents, outpacing evals adoption at 52%” (LangChain), which fits the argument here and comes from one company’s audience.
The question to take to Monday’s meeting
How do you know if an AI agent is working? Ask which of your current numbers the agent could get wrong without the number changing. Then ask for the five that would change: the read, the record, the corrections, the reversals and the planted task, with a limit and a never list written before anyone looks.
The monitoring and trace-reading material behind this post is in Chapter 15, Observability and Debugging (in the full book), and the grading rules are in Chapter 16, Evaluating Agents (in the full book). The guide to evaluating and observing agents collects the related posts and tools, and you can see the formats.
Questions readers ask
- How do you know if an AI agent is working?
- Compare what the agent says it did with evidence the agent did not write. Read a random sample of finished runs each week against a written rule, check outcomes in the system of record, count how often people correct or redo its work, watch complaints and reversals downstream, and run a planted task with a known answer.
- What is a silent failure in an AI agent?
- A silent failure is a run that ends without an error and with a wrong or missing result. The agent returns a fluent reply, often reports success, and every count on the dashboard records the run as handled. The failure shows up later, in the system of record or with a customer.
- Can you trust an AI agent when it says a task is done?
- Not as evidence. The message that says done is text the agent produced, and it can be written whether or not the action happened. Treat it as a claim and confirm it somewhere the agent does not control: the calendar, the database row, the sent folder, the ledger.
- How many agent runs should you review by hand each week?
- Read 50 random finished runs a week if the volume allows, and 100 when the decision is close. With 50 clean runs the failure rate could still be as high as 7.1% at 95% confidence, and with 100 clean runs 3.7%. Below about 50 runs a week, read all of them.
- When should you stop or roll back an AI agent?
- Stop it when one run does something on your written never list, or when the bad runs in your sample put the whole confidence interval above the failure rate you agreed to tolerate. With a 10% limit that is 10 bad runs in 50, or 16 in 100. Stopping means returning the agent to drafting for a person to approve.
Sources
- Laksh Advani (2026). From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents (single-author preprint; abstract read)
- Hongliu Cao, Ilias Driouich, Eoin Thomas (2026). Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation (preprint; abstract read)
- LangChain (survey run 18 Nov to 2 Dec 2025). State of Agent Engineering (a vendor's survey of its own audience)
- Hamel Husain, Shreya Shankar (2026). AI Evals: Everything You Need to Know (evals FAQ; read October 2026)
- Kaizo (2026). Non-Escalation Policy for AI Agents: Silent Failures (a vendor blog; one example of a support-QA vendor's taxonomy)
- Moe Basi, Salesforce (2025). Agent monitoring (a product announcement; one example of a vendor's agent health metrics)
- PegasidsAI (posted to Hacker News by maxadams) (2026). Silent Failures (repository README; one operator's incident notes)
- pupibott, Hacker News (2025). Show HN post on an assistant that reported an email sent (October 2025)
- SkiFreeWin3, Hacker News (2026). Show HN post on verifying what agents did against system state (March 2026)
- nkov47as, Hacker News (2026). Hacker News comment reporting a sandbox vendor's own session logs (February 2026)
- skhatter, Hacker News (2026). Hacker News comment on silent failures in multi-step agent workflows (April 2026)