What is a trace waterfall?
A trace waterfall is one agent run drawn as a stack of horizontal bars, one bar per unit of work. Where a bar starts is when the work began, its length is how long it took, and its indent records what called it. The viewer above draws that picture from a JSON trace and then checks the run against what Chapter 15 of the book says a trace should carry.
The two nouns come from distributed tracing. “A span is one unit of work: a single operation with a start time, an end time, and a bag of structured attributes describing what it did.” A trace is the whole run: “A trace is a tree of spans, linked parent to child, where the child links record who called whom.” For agents, the chapter passes on a mapping from a trace-first debugging guide: “Every LLM call is a span. Every tool call is a span. Every sub-agent is a sub-tree.”
The figure is the first sample the tool loads. Three model calls, a web search and a database call hang under one run, and the database call errs. The chapter’s caption says why the picture is worth drawing: “The bug is not the final answer but the first wrong span”.
What should every span record?
Every model-call span should record the rendered prompt and its hash, the tokens in and out, the latency, the cost, the model requested and the model that answered, and the finish reason. Every tool span should record the tool’s name, its full arguments, its full result and the result’s size, its latency, the principal it ran for, and a structured error if there was one. The viewer shows this inventory for the span you select and marks each entry the span lacks. Sorting the entries by span type is the tool’s arrangement; the chapter gives one list.
The chapter’s list reads: “the tool’s name, its full arguments, and its full result; token counts in and out; latency; cost; the model’s pinned identity, both the one you requested and the one that answered; the reason generation stopped, in the model’s own report; and the error, if any”. Three more entries come from elsewhere in the chapter. The prompt version sits beside the prompt hash in the checks for a run that degraded. The result’s size is the detection layer for cut results. The principal, meaning “which principal asked for it and what scope authorized it”, is the attribute no generic convention records for you.
The rendered prompt gets the most attention, and the guide the chapter quotes puts the reason in three sentences: “The rendered prompt is what the model actually saw. The template plus variables is what you intended. Bugs live in the difference.” The hash is the cheap companion. The chapter’s advice is to “store a hash of the rendered prompt as a span attribute, so when the same query behaves differently on two days, one comparison tells you whether the model’s input actually changed”.
One span sits outside the list. The chapter asks you to “attach a session-level summary at the end of each run, carrying the totals—steps taken, cost, final status”, so that sorting a thousand runs by cost is one query. The last section of the tool counts all of this: for each attribute, how many of the spans that should carry it do. The counts are the tool’s, and there is no pass mark, because the chapter gives none. What it gives is a warning: “what you fail to capture is precisely what you will not be able to debug”.
What does the viewer flag?
The viewer flags seven shapes that a program can find without understanding the task. Each rule is the tool’s own reading of a clue Chapter 15 describes, and each threshold is the tool’s.
| Flag | The tool’s rule | The chapter’s clue |
|---|---|---|
| Error span | The span records an error or an error status | The figure’s first wrong span |
| Repeated tool call | The same tool, 3 or more times in a row; calls one model call makes together count once | The stuck loop |
| Result looks cut | A result that starts like JSON and does not parse | Silent truncation |
| Result size jump | A result more than 3 times the median size for that tool in the run | Result sizes as a span attribute |
| Argument with no source | A text argument of 5 characters or more found in neither the request nor an earlier result | Hallucinated tool arguments |
| Stopped for another reason | A finish reason that is neither a normal stop nor a tool request | The finish-reason distribution |
| At the step cap | The step count reaches a cap the trace states | The wrong stop condition |
An eighth entry, model mismatch, is shown as context and never counts as a flag. A request made by alias always has a responding model with a different name, so the difference alone proves nothing. It matters when you compare two runs.
The repeated-call rule follows the chapter’s own query for a stuck loop: “group tool spans by name within a session and flag long consecutive runs”. The flag sits on the first call of the run, because “the cause is almost always visible in the first repeated span”. The number three is the tool’s choice. The chapter uses the same number for a different purpose, the point where the harness should intervene: “after the third identical attempt inject a message that says so and asks the model to change course or report the obstacle”.
The cut-result rule is narrow on purpose. The class the chapter most wants you to fear has no error to find. In the words of an essay it quotes, “There is no error, just a clean tool call followed by a clean assistant reply”. A result that begins as JSON and stops mid-array is the one version of that a parser can see. The same shape appears when the trace store capped the result as it saved it, so compare the stored size before you blame the transport. The size rule covers the chapter’s wider advice to “record every tool result’s size as a span attribute, alert when a tool’s result distribution jumps, and fix oversized results at the source by paginating”. Inside one run it needs three or more results from the same tool, so on short traces it stays silent.
The argument rule is the chapter’s mechanical test for an invented value: “an argument’s value should be sourced from the user’s text or from a prior span’s result, so diff it against both, and a value with no source was invented”. The tool’s version is much weaker than a person’s diff. It looks for the exact text of each argument, so the commonest false hit is a value the model reworded or reformatted, followed by a fixed option of the tool, such as a sort order, and a value from an earlier turn. It only runs when the trace includes the request text. Treat it as a pointer to an argument worth reading.
The viewer marks the earliest flagged span in accent, passing over an error on a parent when a span beneath it is flagged too. That is the tool’s ordering, and it is a candidate. The chapter’s method asks for a judgment the tool cannot make: “read the trace forward to the first step where something is wrong—not the last step, where the wrongness became visible—because that first step is the bug, and everything after it is consequence”.
How do I compare a good run with a bad one?
Load the bad run, then choose or paste a second run to compare. The viewer pairs the spans of the two runs in order and points at the first pair that differs in type, name, arguments, result or prompt hash. Then it reports three checks in the order Chapter 15 gives them. The checks run even when every pair matches, because a changed model or prompt version can leave the shape of a run untouched.
The move is the chapter’s treatment for a run that worked last month and fails this month with no code change: “Take one good May trace and one bad June trace with the same input shape and diff them; the disciplined version of the move is finding the first step where the two disagree.” The three checks follow: “the responding model against the requested one (an alias rolled forward); the prompt version (a teammate shipped an edit); the prompt hash (same template, but a referenced variable now renders different content, which is to say the corpus moved under you)”.
The fourth sample shows the third case. Both runs ask the same question with the same model and the same prompt version. The search tool returns a different policy text in the bad run, so the first difference is at the tool span, and the prompt hash of the next model call differs too. The pairing by position is the tool’s alignment and it is crude: if one run takes an extra step early, every later pair will differ. The first difference is still the place to start reading.
A diff of two traces is the small-scale view of drift. The figure shows the large-scale one, a quality score sliding across a population of runs where no single run looks broken.
What JSON does the viewer read?
The viewer reads an object with a spans array, or a bare array of spans. Each span is an object with the keys below. Every key is optional. Without times the bars are drawn in order with equal width; everything else feeds the panel, the flags and the counts.
| Key | Meaning |
|---|---|
id, parent |
The span’s id and its parent’s id (no parent for the root) |
type |
run, model, tool, agent or summary |
name |
The label on the row |
start_ms, end_ms |
Start and end in milliseconds, from any origin (duration_ms can stand in for the end) |
prompt, prompt_hash, prompt_version |
The rendered prompt, its hash, and the prompt’s version |
model_requested, model_responded |
The two model identities |
tokens_in, tokens_out, cost, finish_reason |
The model call’s measurements, cost in your own unit |
tool, arguments, result, result_bytes |
The tool call and the stored size of its result |
principal |
On whose authority the tool call ran |
error |
Text, or an object with code, message and retryable |
steps, final_status |
Totals on the summary span |
Two keys on the outer object are optional: input, the user’s request, which the argument rule needs, and step_cap, which the step-cap rule needs.
The chapter notes that the industry has been converging on an open, vendor-neutral telemetry standard with agreed names for the generative-AI attributes. The viewer accepts that standard’s export shape and its attribute keys for the operation, the two models, the token counts, the finish reason and the tool call as aliases for the keys above. This is best effort. The chapter warns that those conventions are still in development and that names can change between releases, so the viewer lists the span keys it did not read, up to a dozen, and guesses at none of them.
Size limits are the tool’s: about 2 MB of JSON and 2,000 spans, which is one run and not a day of traffic.
Why is content hidden by default?
Content is hidden by default for your own traces because rendered prompts and tool results are where user data lives. The viewer shows sizes in place of text until you clear the checkbox, and the Markdown export carries names, sizes, hashes and counts only.
This follows the chapter’s rule for trace stores in general: “treat content capture as a deliberate, redactable, off-by-default data flow, with retention limits and access rules, rather than a free debugging convenience”. The same reasoning is why a pasted trace is kept out of the page address. Other tools on this site put their whole state in the link so it can be shared. Here the link holds only which samples are open, which span is selected and the checkbox.
What can the viewer not tell you?
The viewer cannot tell you whether the answer was right, and it cannot make the failure happen again. It reads one recorded run. Correctness is an evaluation question, and reproduction needs a recording you can replay.
The chapter’s cheap first move on reproduction uses what the trace already holds: pull the rendered prompt of the failing step and replay that one call against the pinned model. If the trace does not hold the rendered prompt, the schema counts in the tool have already told you what to add. The guide the chapter quotes prices the stakes: “The span schema you set up upfront determines whether the next debug session takes 5 minutes or 5 hours.”
Four limits are specific to this page. The flags are per run, so anything that only shows across many runs, such as a slow change in finish reasons, is out of reach; the chapter’s line that “A drift toward “ran out of tokens” or toward content filtering is an early, cheap symptom that something upstream changed shape” is about a dashboard, and the per-span flag here is a weaker cousin. The rules read structure and never meaning, so lost or poisoned context is found by you, reading the rendered prompt in the panel. The schema has no field for the model’s own output, so two runs that make the same calls and give different final answers show no difference. And the sample traces are invented for the page; their timings follow the figure and are illustrative.
Once a flag has sent you to a span, the agent bug bestiary has the symptom, cause and fix for each of the chapter’s six shapes. If the span is a tool call with a vague error, run the definition through the tool contract linter. For repeated calls, the circuit breaker simulator shows what a retry policy does to a failing dependency, and the agent cost per task estimator shows what the extra steps cost. The agent verifiability scorecard asks whether your traces are complete as one of its questions, and the multi-agent failure explorer covers failures at the seams between agents, which appear in a trace as ordinary parent and child edges. The full argument is in Chapter 15, and the agent evaluation guide picks up where a fixed trace becomes a test.
Questions readers ask
- Why store the rendered prompt and not the template?
- Because the rendered prompt is the text the model actually received, and the template is only what you meant to send. A retrieval step that added the wrong document, a truncation that dropped the one turn that mattered, or a variable that rendered empty are all invisible in the template and plain in the rendered text. Chapter 15 adds a cheap companion: store a hash of the rendered prompt on the span.
- What is the first thing to look for in a bad trace?
- The first step where something is wrong, read forward from the start. Chapter 15 says that step is the bug and everything after it is consequence, so the last step, where the damage became visible, is the wrong place to begin. The viewer marks the earliest span that one of its rules flagged, as a candidate for you to open.
- Why record both the requested and the responding model?
- Because hosted models sit behind aliases, and aliases roll forward. If you record only the name you asked for, an upgrade leaves no mark in your data. With both on every span, a change in the responding model is one comparison away, and it is the first of the three things Chapter 15 tells you to check when a run that used to work stops working.
- Is my trace sent anywhere?
- No. The page has no server side and makes no network requests. A trace you paste or open is parsed in the browser tab and held in memory. It is not written to the page address, so a copied link never carries it, and it is gone when you reload. Prompt, argument and result text is hidden by default for your own traces, and the Markdown export never includes it.
- Does a trace with no flags mean the run was correct?
- No. The rules catch shapes a program can check: a repeated call, an error, a result that does not parse. A wrong answer built on a result that looks fine passes every rule. Chapter 15 treats tool calls and results as ground truth about what happened and the model's narration as testimony about why, and reading the two against each other is still your job.
Sources
- Respan (2026). AI Agent Debugging
- Gabriel Anhaia (2026). Tool-Result Truncation: The Silent Bug That Makes Agents Lie
- OpenTelemetry (2026). Semantic conventions for generative AI systems