Agent observability is the practice of recording each agent run as a trace, a tree of spans for every model call, tool call and handoff, and watching those runs as a population. A trace says what happened. It holds no statement of what should have happened, so correctness needs a second thing: a judgment.
I wrote this for an ML engineer who already has tracing, or is about to buy it, and has been asked whether the agent is any good. Most pages on agent observability list what to capture. This one draws the line between a record and a verdict: which questions a trace answers alone, which it cannot, and the three steps that turn traces into judgments.
One number shows the gap. In a public survey one tooling vendor ran from 18 November to 2 December 2025 (1,340 responses), “89% of organizations have implemented some form of observability for their agents,” while “Just over half of organization (52.4%) report running offline evaluations on test sets” (LangChain, State of Agent Engineering). The sample is self-selected and the vendor sells both kinds of tool, so read it as a description of those respondents: 36.6 points more of them report observability than report offline evaluations on test sets. The same page puts those doing no evaluation at all at 29.5%.
What is agent observability, and what does it leave out?
Agent observability is the ability to answer questions about a run from what the run recorded, without running it again. It leaves out the standard the run is measured against. The general definition comes from the tracing standard’s own primer (undated; read October 2026): “Observability lets you understand a system from the outside by letting you ask questions about that system without knowing its inner workings.”
For an agent the recorded unit is the trace, built from spans: one per model call, one per tool call, a subtree per subagent. What each span must carry is the subject of the post on LLM tracing with OpenTelemetry, and I won’t repeat the field list here.
One vendor’s comparison (Braintrust Team, June 2026) sorts three neighboring words by the question each answers. Monitoring asks “Is the system healthy?” Observability asks “What is happening inside the system?” Debugging asks “Why did this specific request fail, and how do I fix it?”
All three are questions about what is. A fourth question is missing from that list, and it is the one a product owner asks first: was this run right? That question belongs to evaluation. Chapter 15 of AI Agents, Engineered, “Observability and Debugging,” (in the full book) draws the same boundary in its section “Production Monitoring”: “Monitoring tells you the agent is running; whether it is correct is an evaluation question”.
What can a trace answer, and what can it not?
A trace answers questions about what crossed the boundary between your code, the model and the tools in one run. It cannot answer questions about the world after the run, about what should have happened, about how often, or about what inside the model caused an output. The table sorts fourteen common questions into those five kinds of evidence; the sorting is mine.
To place a question that isn’t in the table, ask five things of it in order, and note that one question can need more than one kind.
- Is it about what was sent, returned, counted or timed in this run? The trace alone answers it.
- Is it about the state of some system after the run, or about what the user needed beyond what they typed? It needs the system of record, or the user.
- Does it contain “should,” “correct,” “right,” “allowed” or “good”? It needs a judgment: an expectation and a judge.
- Is it about how often, or about what would happen under a change? It needs many runs.
- Is it about what inside the model caused an output? That was not recorded, by any tracer.
| Question about a run | What the trace holds | What is missing | Evidence needed |
|---|---|---|---|
| What did the model see when it made this call? | The rendered prompt and the tools offered | Nothing, if content was captured | trace alone |
| Which tool was called, with which arguments, and what came back? | The tool span | Nothing | trace alone |
| Where did the time and the tokens go? | Timings and token counts per span | Nothing | trace alone |
| Did the model’s input differ between the good run and the bad run? | Prompt hash, prompt version, responding model, on both traces | Nothing, given both traces | trace alone |
| At which step do two runs first disagree? | Both span trees | Nothing; which of the two is wrong is the next row | trace alone |
| Which step was the first wrong one? | Every step, in order | What a right step looks like here | trace plus judgment |
| Was the final answer correct? | The answer text | A reference, a rubric or a check | trace plus judgment |
| Should this tool have been called at all? | The call, and the principal if you recorded it | The policy for this user and task | trace plus judgment |
| Was a required step skipped? | Only the steps that happened | The list of steps that must happen | trace plus judgment |
| Did the action take effect? | The tool’s reported status | The state of the system the tool acted on | trace plus system of record |
| Did the user get what they needed? | The request as typed, and any feedback signal | The user | trace plus system of record |
| How often does this failure happen? | One sample | The same check over a population of traces | many runs |
| Would a different prompt or tool description fix it? | The recorded input to re-ask from, if content was captured | Repeated re-runs with the change applied | many runs |
| What inside the model caused this output? | The model’s narration, if it wrote any | The mechanism | not recorded |
The first five rows are why tracing is worth its cost: without them, none of the rest can be asked. The other nine are why a full trace store doesn’t tell anyone whether the agent is good.
How does one incident split across the table?
It splits into six questions with five different kinds of evidence. This is a worked example with illustrative numbers. A customer writes that a support agent told them a refund was done, and no money arrived.
Was the refund tool called? Trace alone. Step 7 is issue_refund with the right order and amount, and the result reads status: queued.
Did the refund take effect? System of record. The payment ledger shows the queued refund was rejected the next morning by a fraud rule. Nothing in the trace could show that, because it happened after the run ended.
Should the agent have said “refunded”? Judgment. The expectation, once someone writes it down, is that the final message claims completion only when the tool’s status is settled. A few lines of code can be the judge.
Which step was the first wrong one? Judgment again. Step 7 did what was asked. Step 8, the final message, is the first step that breaks the expectation.
How often? Many runs. The same check over last week’s traces finds 14 of 212 refund runs (6.6%) with a queued status and a completion claim.
Why did the model say it? Not recorded. Its narration at step 8 says the refund “was processed successfully,” which restates the error and explains nothing. A hypothesis (the tool result never says the refund is incomplete) can be tested by replay, covered below.
The post on agent trajectory evaluation covers how to write the end-state check that the second question calls for.
Why can’t a trace tell you why the agent did it?
Because the only “why” in a trace is the model’s own account, and that account is generated text. Chapter 15, in “Traces and Spans,” separates the two kinds of content: “In a trace, the tool calls and results are ground truth about what happened; the narration is testimony about why. Read the testimony; verify against the record.”
There is a measurement of how incomplete that testimony can be. In an April 2025 study, one lab gave models a hint toward an answer, confirmed the models used it, and then checked whether their written reasoning said so. Of the two reasoning models tested, one “mentioned the hint 25% of the time” and the other “39% of the time,” averaged across hint types (Anthropic, 2025). Those are two models on one constructed task, so the figures don’t transfer to your agent; the finding that does is that narration can leave out the cause.
Replay doesn’t close this gap either. One practitioner’s account of record-and-replay lists it among the limits: “It doesn’t explain why the model said what it said” (Tian Pan, April 2026). The book says the same: “Replay shows you what happened, never why the model said it.”
Pages that rank for agent observability can suggest otherwise. One frames the layer agents add as the question “Why did it take that action?” A reader on Hacker News described what the tools deliver: “most tools record what happened (tool X was called, output was Y), but not why the agent deviated from the plan” (zippolyon, March 2026).
What a trace does support is a cause stated as a hypothesis about inputs. “The result at step 4 was cut short, and the model answered from the half it received” can be checked against the record and tested by changing that input.
What does a judgment need that a trace lacks?
A judgment needs three things: an expectation written down before or apart from the run, a judge that compares the run with it, and a verdict stored against the run. This three-part framing is mine. A trace supplies the evidence the judge reads and none of the three.
| Part | What it is | Forms it takes |
|---|---|---|
| Expectation | What right looks like for this input | An end state to assert, a forbidden call, a required step, a reference answer, a rubric |
| Judge | What compares the run with the expectation | Code, a calibrated model, a person |
| Verdict | The result, attached to the run | Pass or fail, the first wrong step, a failure category |
The judge is where the neighboring posts go deep: whether an LLM judge is reliable for the model option, and AI agent evaluation metrics for what to count once verdicts exist.
How do traces become judgments?
Traces become judgments in three steps, and they are the part of agent observability that no tracer does for you: someone chooses which runs to read, writes a verdict and a note on each, and promotes the ones worth keeping into the eval set. Chapter 15 calls the last step the habit its author most hopes a reader takes from that part of the book, and it describes a failing trace as “the seed of an eval case: the input, the observed wrong behavior, and your judgment of what right looks like.”
Which traces should you read?
Read two batches: the worst runs by a signal you already have, and a random sample. The book’s version of the first, from “Production Monitoring,” is to “sort the week’s sessions by a signal that correlates with trouble (cost, step count, a low judge score) and read the worst ten from the top.”
Sorting by a signal finds the failures that signal describes. The random batch is for the rest, and its size decides how rare a failure it can find. If a failure occurs in a fraction p of runs and you read n traces drawn at random, the chance that the batch holds at least one instance is 1 − (1 − p)n, assuming independent draws. The table is my arithmetic.
| Failure occurs in | 10 traces read | 30 traces read | 100 traces read |
|---|---|---|---|
| 1% of runs | 9.6% | 26.0% | 63.4% |
| 2% of runs | 18.3% | 45.5% | 86.7% |
| 5% of runs | 40.1% | 78.5% | 99.4% |
| 10% of runs | 65.1% | 95.8% | over 99.9% |
The other direction matters as much. If you read n random traces and see a failure in none, the one-sided 95% upper bound on its rate is about 3 ÷ n: about 10% after 30 traces (9.5% exactly), about 3% after 100. Thirty clean traces do not show that a failure is rare.
The evals FAQ by Hamel Husain and Shreya Shankar (modified September 2026) gives a stopping rule in place of a fixed count: “Continue until new traces stop revealing failure modes or changing existing ones.” Its floor is “We recommend reviewing at least 100 traces.”
What do you write on a trace?
Write a pass or fail against a stated expectation, the first wrong step, and a free-text note on what you saw; categories come later, from the notes. The same FAQ describes the first pass as open coding, in which annotators “review and write open-ended notes about traces, noting any issues,” and it advises “noting the first failure observed in a trace, as upstream errors can cause downstream issues.” The book’s method is the same: “read the trace forward to the first step where something is wrong.”
Expect the expectation to move while you annotate. A 2024 study of people grading model outputs named the effect criteria drift: “users need criteria to grade outputs, but grading outputs helps users define criteria” (Shankar and colleagues, 2024). So write the expectation on each note, and re-read your early notes after the first thirty.
Once notes exist, group them into named failure categories. The six shapes in the agent bug bestiary are a starting set. A research group built a published taxonomy this way: its authors analyzed 150 traces with expert annotators, reached an inter-annotator agreement of kappa = 0.88, and arrived at 14 failure modes for multi-agent systems (Cemri and colleagues, 2025).
Agreement between two readers is the check that a category means something. If two of you label the same twenty traces and disagree often, fix the definitions before counting anything.
The note below is the unit of work. Its sections follow the five kinds of evidence, so that a fact read from the trace is never mixed with a verdict or a guess.
TRACE REVIEW NOTE
Run id: [...] Reviewer: [...] Date: [...]
Selected by: [random sample | worst by <signal> | user flag | incident]
1. Facts read from the trace (no opinions here)
- Task as the user typed it: [...]
- Steps: [n]; tools called, in order: [...]
- Final output, in one line: [...]
- Anything cut, empty or errored in a tool result: [step, what]
2. Checked outside the trace
- System of record: [what was looked up, and what it shows] or "not checked"
- User: [feedback, follow-up, complaint] or "none"
3. Judgment
- Expectation used: [the end state, rule or rubric line, written out]
- Judge: [code check | calibrated model judge | person: name]
- Verdict: [pass | fail | cannot tell: what is missing]
- First wrong step: [step number, and what a right step would have been]
4. Open note (what you saw, in your own words)
[...]
5. Category (fill in after the first 30 notes exist)
[name from the team's failure list, or "new: ..."]
6. Cause
- Hypothesis about an input: [...]. The model's narration is not evidence.
- Test: [re-ask step n with <one change>, k times] -> [result] or "untested"
7. How often
[unknown until the check runs over many traces | k of n runs in <window>]
8. Promote to the eval set? [yes | no | golden trace]
- The failure, or the correct behavior, is one the team will guard: [y/n]
- The expectation can be written as a check: [y/n]
- The input and tool results can be rebuilt with no live dependency: [y/n]
- Personal data removed or replaced: [y/n]
- Split: [dev | regression | held out]
9. Replay level this run supports: [read | re-ask one call | re-run harness]
Missing for the next level: [...]
Here is the refund incident on the note. Section 1 holds eight steps, issue_refund at step 7 returning status: queued, and a final message saying the refund is done. Section 2 holds the ledger entry that shows it rejected.
In section 3, the expectation is “claim completion only on a settled status,” the judge is a code check, the verdict is fail, and the first wrong step is 8. Section 6 holds the hypothesis about the tool result’s wording, and section 7 holds 14 of 212.
When does a trace become an eval case?
Promote a trace when three conditions hold, which are mine: the behavior is one you intend to guard, the expectation can be written as a check, and the input and tool results can be rebuilt with no live dependency. A trace that fails the second condition stays a note until someone can say what right looks like. The note adds a fourth line, personal data removed, which is a condition for storing the case and not for choosing it. The book’s line for the ones that pass is “a flagged trace is an eval case that has not been written down yet.”
Promotion changes what the trace is. The eval case keeps the input, the recorded tool results and the check, and it drops the transcript as a reference. A case that asserts the old run’s exact sequence fails every valid alternative path, which the trajectory post explains.
Promote some passes too. One essay on replay recommends that teams “capture production runs as golden traces, then replay them against new code versions” (Tian Pan, April 2026). Each promoted case then needs a split so that the cases you tune against are kept apart from the ones you report; the three-set splitter assigns them. A fuller method is planned for a post on building a golden dataset from real traces.
A lab’s guide to agent evals makes the reverse dependency explicit: the eval needs the transcripts too. “You won’t know if your graders are working well unless you read the transcripts and grades from many trials” (Anthropic, January 2026).
What must agent observability record so a run can be replayed?
It must record every value that crossed the nondeterministic boundary, in full, and how much you record sets which of three levels of replay you get. The book’s phrase for the target is “every crossing of the nondeterministic boundary”; the three levels are my framing of its two techniques plus plain reading.
| Level | What you can do | What must have been recorded | What it cannot tell you |
|---|---|---|---|
| 1. Read | Follow the run step by step | The span tree; arguments and results; the rendered prompt or a reference to it | Anything about a second attempt |
| 2. Re-ask one call | Send one step’s exact input to the model again, change one thing, compare | The full rendered prompt as content, the tool definitions offered, the sampling parameters, the model that responded | What the rest of the run would then do; each re-ask is a fresh sample |
| 3. Re-run the harness | Run your own code again with every model and tool response served from the recording | Every model response and tool result in full, with errors and timeouts; the clock, random seeds and generated identifiers; configuration read; the code version | What the model would say to a changed prompt |
Level 2 is the book’s first move: pull the rendered prompt of the failing step and “replay just that one call against the pinned model.” It is also the feature one engineer on Hacker News said three observability vendors had failed to give their team: to log a completion and “re-run the exact same completion” (seany62, October 2025). A hash of the prompt is not enough for it; the content has to exist somewhere you can read.
In the refund incident, level 2 means re-asking step 8 ten times, then ten more with the tool result reworded to say the refund is not complete. Suppose the claim appears in 7 of the first 10 and in 0 of the second 10 (illustrative). That supports the hypothesis, and zero in ten still allows a rate of up to 25.9% by the same one-sided bound, computed exactly. When a re-ask doesn’t reproduce the failure at all, the book counts that as information: “If it does not reproduce, that is a finding too.”
Level 3 has one property that decides whether it can be trusted. If the code asks for something outside the recording, the engine “fails loudly rather than silently falling through to a live system” (Tian Pan, in the essay cited above). The same essay names the precondition: “It can’t replay what it didn’t record.”
A recorded response is an answer to the recorded prompt. So, on my reading, level 3 tests changes to your code (parsers, context assembly, stop logic), and a change to a prompt or a tool description sends you back to level 2 or to a live run, repeated enough times to count.
The first three items give you level 2, and items four to ten add level 3. The last is the condition for being allowed to keep any of it. The book’s rule for the whole list: “what you fail to capture is precisely what you will not be able to debug.” Step-by-step use of both techniques on a failing run is the subject of a planned post on how to debug an AI agent.
Where does this approach break?
Agent observability turned into judgments this way breaks on cost, on volume, on privacy, and on what it leaves unguarded.
Reading and annotating are slow. The authors of the evals FAQ report spending “60-80% of our development time on error analysis and evaluation” in their projects.
Some teams decline that trade. In the Hacker News thread quoted above, the original poster did not want “datasets” or “scores,” and a reply argued that outputs “can’t really be automatically scored” and that “It’s practical just to do it manually.” I’d answer that a person re-running a completion and deciding it looks wrong is a judgment too. It has an expectation and a judge, and the verdict was never stored.
Random sampling needs volume. With 40 runs a week, read all of them, and treat the table of probabilities as an argument for patience.
Judging needs content, and content is where users’ data lives. A trace store that keeps only hashes can show that an input changed and cannot support levels 2 or 3. The tracing post covers the storage options.
An eval set grown only from observed failures guards against the past. It says nothing about inputs no user has sent yet, so it supplements cases written from the task’s requirements.
Seeing a failure does not stop it. A check that blocks or caps a run belongs to the harness, which the post on how to make AI agents more reliable takes in order: retries, budgets, recovery. Watching the population for drift is the job of the post on how to monitor AI agents in production.
The one thing to keep
Agent observability gives you records, and a record becomes useful when someone attaches an expectation and a verdict to it. Before the next vendor call or the next sprint, take ten traces, five from the worst of the week and five at random, and fill in the note for each. Count how many sections you could complete from the trace alone. The rest is your evaluation backlog.
Chapter 15 ends on an image: “The traces are the flight recorder.” The flight plan is Chapter 16, Evaluating Agents (in the full book), on evaluating agents; the record-and-replay material is in Chapter 15, Observability and Debugging (in the full book). The guide to evaluating and observing agents collects the related posts and tools, and you can see the formats.
Questions readers ask
- What is agent observability?
- Agent observability is the practice of recording each agent run as a trace (a tree of spans for every model call, tool call and handoff) and watching the population of runs over time, so that a wrong answer can be read step by step without running the agent again. It records what happened. Whether that was correct is decided by evaluation.
- What is the difference between observability and evaluation for AI agents?
- Observability produces records: what the model saw, which tools were called, what came back, what it cost. Evaluation produces verdicts: a written expectation, a judge and a pass or fail stored against the run. Evaluation depends on observability for its evidence, and observability with no evaluation can only report that runs completed.
- Can a trace tell me why the agent made a decision?
- It can show what the model was given and what it wrote about its own reasoning. The first is a record. The second is the model's account, which can omit the real cause. Treat a cause as a hypothesis and test it by re-asking the same model call with one input changed, several times.
- How many traces should I review?
- It depends on how rare a failure you want to meet. With independent random draws, 30 traces give a 78.5% chance of seeing a failure that occurs in 5% of runs, and 100 traces give a 63.4% chance for one that occurs in 1%. The evals FAQ by Husain and Shankar recommends at least 100 and stopping when new traces stop showing new failure modes.
- What do I need to record to replay an agent run?
- To re-ask one model call: the full rendered prompt, the tool definitions offered, the sampling parameters and the model that responded. To re-run your own code against the recording: every model response and every tool result in full, including errors and timeouts, plus the clock, random seeds, generated identifiers and the code version.
Sources
- LangChain (survey run 18 Nov to 2 Dec 2025). State of Agent Engineering
- Hamel Husain and Shreya Shankar (2026). AI Evals: Everything You Need to Know (FAQ)
- Anthropic (2026). Demystifying evals for AI agents
- Anthropic (2025). Reasoning models don't always say what they think
- Tian Pan (2026). Deterministic Replay: How to Debug AI Agents That Never Run the Same Way Twice
- Braintrust Team (2026). 7 best tools for debugging AI agents in production
- OpenTelemetry (undated; read October 2026). Observability primer
- Shreya Shankar et al. (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
- Mert Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail?
- seany62 (Hacker News) (2025). Ask HN: Good LLM Observability Platforms?
- PuppyGraph (undated; read October 2026). What Is Agent Observability? How Does It Work?
- zippolyon (Hacker News) (2026). Hacker News comment on what agent logs record