Home / Blog / Evaluating and observing agents / LLM Tracing With OpenTelemetry: What an Agent S…

Evaluating and observing agents

LLM Tracing With OpenTelemetry: What an Agent Span Should Carry

LLM tracing with OpenTelemetry starts from a span contract: what each agent span must carry and never record, mapped to two moving conventions. Copy it.

By Enrique Gutiérrez · Published · 25 min read

LLM tracing with OpenTelemetry means recording each agent run as a tree of spans (one per model call, one per tool call, a subtree per subagent) and deciding what every span must carry. The tree rests on a stable standard. As of October 2026 the AI attribute names are all marked Development, so own the contract and map outward.

I wrote this for a backend engineer who already runs distributed tracing and has been told the agent must appear in it. Most guides to LLM tracing with OpenTelemetry end at an install step. This one ends at a contract: what each span must carry and what it must never carry. A dated mapping to two public conventions comes with it, so you instrument once and keep your traces when the tooling moves.

What does a trace have to let you answer?

A trace has to answer the questions you will ask about a bad run without running it again. Chapter 15 of AI Agents, Engineered, “Observability and Debugging” (in the full book), makes that the whole test: “Now the part that decides whether your traces are worth having, which is what each span carries.”

The book’s free glossary defines a trace as “The complete, structured record of one run—every model call, tool call, and result, as a tree of spans carrying timings, token counts, and costs.” A span is “One unit of work in a trace: a single operation with a start time, an end time, and structured attributes.” The questions below come from the book; the pairing with fields is mine.

Question you will ask Fields that answer it Chapter
“what exactly did the model see at the moment it made this call?” reference to the rendered prompt; the tools offered 15
Did the model’s input change between the good day and the bad day? hash of the rendered prompt; prompt version 15
Which model answered? requested model; responding model 15
Where did the tokens go? input and output tokens on each model call 15, 19
Who called whom? parent links 15
What did the tool receive and return, and did all of it reach the model? tool name, arguments, result, sizes, structured error 15, 18
On whose authority was the action taken, and on what? principal and scope; target id for a write 15, 17
What crossed between two agents? brief and return on the handoff 11
Which of yesterday’s runs were expensive or failed? the session summary 15

The first question is the one most logging fails. The book’s reason is a small economy: storing the prompt template and its variables in place of the rendered prompt, “the final assembled text that actually went to the model”. A variable that rendered empty is invisible in the template and plain in the rendered text.

This post covers one run’s schema. The wider practice, agent observability, adds monitoring and drift across many runs.

Which span kinds does an agent run need?

An agent run needs five span kinds: the run, the model call, the tool call, the subagent handoff and the session summary. Chapter 15 takes the mapping from a trace-first debugging guide (Chen, 2026): “Every LLM call is a span. Every tool call is a span. Every sub-agent is a sub-tree.”

The tree has a fixed shape. “The root span is the user’s request.” Each pass through the loop hangs a model-call span beneath it, and each tool span hangs off “the model call that asked for it”. When an orchestrator delegates, the subagent’s whole run becomes a subtree under the handoff.

One agent run drawn as a span waterfall.
Figure 15.2 One agent run drawn as a span waterfall. Each bar is a span: where it starts is when it happened, how long it runs is how long it took, and the indentation of its row records who called whom—the tool spans nest under the model call that asked for them, and everything nests under the root run. Every span carries the same record, the attributes bracketed over the first model call: the rendered prompt, the tokens, the cost, the latency. The bug is not the final answer but the first wrong span—here the long database call that erred, in accent—and a waterfall makes it the thing your eye lands on. The timings and ordering are illustrative. Reuse this diagram

The fifth kind looks redundant, and the book asks for it anyway: “attach a session-level summary at the end of each run, carrying the totals—steps taken, cost, final status”. With it, “show me yesterday’s runs sorted by cost” is one query.

Retrieval steps, guardrail checks and human approvals get no kind of their own in the book, and I treat each as tool-call-shaped.

What should each agent span carry, and what must it never record?

Each span carries the fields that answer one of the nine questions, and no credential ever lands on any of them. The table is my synthesis of the chapters in its last column. A dagger (†) marks a field the book does not ask for, and storing content by reference is this post’s choice as well.

Span kind Required Optional Never record Why Chapter
Run (root) trace id†; run id†; start, end; status; agent name and version†; convention name and pinned version† session id†; task type†; tenant as an id†; clock reads and seeds, for replay the user’s name or email where a pseudonymous id answers the question groups every span of one request and states which vocabulary the trace speaks 15
Model call parent link; step index†; requested model; responding model; prompt version; hash of the rendered prompt; reference to the rendered prompt; hash of the tool definitions offered†; input and output tokens; finish reason; start, end; the decision†; structured error, if any sampling parameters; cached and reasoning token counts as sub-totals†; reference to the output; time to first token† the rendered prompt inline, unless the content policy allows it for this environment; a template presented as the prompt; request headers answers what the model saw, whether its input changed, which model answered and where the tokens went 15
Tool call (child of the model call that asked) tool name; call id†; arguments, by reference or redacted; result, by reference; bytes returned, bytes delivered to the model†, bytes stored†; truncated flag†; status; structured error (code, message, retryable); start, end; principal and scope; read or write†; for a write, a pseudonymous target id† hash of the arguments†; attempt number†; idempotency key for a write†; tool version† credentials or session tokens passed as arguments; whole documents inline when a reference and a size will do; personal data unredacted by default ground truth of what the agent did, to what, and whether the whole result reached the model 15, 17, 18
Handoff (parent of the subagent’s subtree) agent name†; the brief and the return, by reference; status; the subagent’s stop reason†; start, end subtree token roll-ups under their own names† a copy of the parent’s full context; roll-ups under the per-call token names the seam between agents becomes an ordinary parent-child edge 11, 15
Summary (one per run, written last) steps of the lead agent’s loop; model calls across all agents†; tool calls†; run token totals under their own names; final status; stop reason†; cost estimate; price-table version† evaluation score; user feedback flag totals under the per-call token names sorting a thousand runs by cost or outcome is one query 15, 19

The book’s own inventory is shorter and worth quoting, because the table mostly sorts it by span kind: “the tool’s name, its full arguments, and its full result; token counts in and out; latency; cost; the model’s pinned identity, both the one you requested and the one that answered; the reason generation stopped, in the model’s own report; and the error, if any, in the structured shape” that Chapter 5 teaches. Elsewhere the book adds the prompt hash, the result’s size and the authority behind an action.

Three additions came from testing other failures. A hash of the arguments shows identical repeated calls without reading content. The handoff’s stop reason says whether a subagent ended on an answer, a cap or a budget. The target id says what a write touched, where principal and scope say only who acted.

Three more fields deserve a reason each. The hash is of the rendered prompt, so that “one comparison tells you whether the model’s input actually changed”. The decision field is mine: the book reads the decision from the tool spans beneath a model call, and I add a field so that finding every step that answered without calling a tool is one query. Tool definitions get their own hash because they often travel beside the messages; a 2026 issue on one agent framework reports them missing from the record.

What does the contract look like as text?

It looks like a short block a team can commit beside the tracer, with the same required, optional and never entries as the table. Field names here are placeholders that you rename once and then keep. Its bracketed header lines hold the two decisions the contract forces: the convention version emitted and where content lives.

SPAN CONTRACT: [agent or feature] · owner: [name] · reviewed: [date]
Convention emitted: [name] at [pinned version or commit] · upgrades are migrations
Content policy: [by reference | inline in pre-production | not stored]
  readers: [role] · retention: [days] · redaction: [rule or none]

every span
  trace_id, span_id, parent_span_id     # who called whom
  kind: run | model_call | tool_call | handoff | summary
  start_time, end_time                  # latency = end - start
  status: ok | error
  error: {code, message, retryable}     # only when status = error

run
  run_id, agent.name, agent.version
  convention.name, convention.version
  optional: session_id, task_type, tenant_id, clock_reads, seeds

model_call
  step                                  # within its agent
  model.requested, model.responded
  prompt.version                        # from source control
  prompt.hash, prompt.ref               # hash of the RENDERED prompt; pointer to it
  tools.hash                            # tool definitions offered
  tokens.input, tokens.output           # this call only; cached counts are sub-totals
  finish_reason                         # in the model's own report
  decision: tool_calls[names] | handoff | final_answer | ask_human
  optional: params, tokens.cached, tokens.reasoning, output.ref, time_to_first_token

tool_call                               # child of the model_call that asked
  tool.name, tool.call_id
  args.ref, result.ref                  # content per policy
  result.bytes_returned                 # as the dispatcher received it
  result.bytes_delivered                # as the model received it
  result.bytes_stored                   # as the trace kept it
  result.truncated                      # tool said so, or delivered < returned
  principal, scope                      # whose authority, which permission
  side_effect: read | write
  target.id                             # writes only; pseudonymous
  optional: args.hash, attempt, idempotency_key, tool.version

handoff
  agent.name, brief.ref, return.ref
  stop_reason                           # answer | step cap | budget | error
  optional: subtree.tokens.input, subtree.tokens.output

summary
  steps                                 # the lead agent's loop only
  model_calls, tool_calls               # across all agents
  run.tokens.input, run.tokens.output   # sums of model_call spans, own names
  cost.estimate, price_table.version
  final_status, stop_reason
  optional: eval.score, user.feedback

never on a span
  credentials, API keys, session tokens, request headers
  a template presented as the prompt
  whole documents inline, or a copy of the parent's context on a handoff
  raw personal data where a reference or a pseudonymous id answers the question
  a token total under the same name as the per-call field

What does one traced run look like?

One traced run looks like a small tree in which every number can be recomputed from the model-call spans. The example is illustrative: I invented the agent, the token counts, the byte sizes and the prices, and nothing in it is a measurement. A support agent is asked to move a delivery, reads the order, hands a policy question to a subagent and answers.

run  move-delivery  status=ok  agent=support@v31  convention=[name]@[pin]
├─ model_call step=1  tokens in=3,200 out=500  finish=tool_calls  decision=call get_order
│  │           model M -> M-0412  prompt v14  prompt.hash=9f2c  tools.hash=a1e0
│  └─ tool_call get_order  ok  read  principal=user:u_771  scope=orders:read
│               returned=1,930 B  delivered=1,930 B  stored=1,930 B  truncated=false
├─ model_call step=2  tokens in=4,200 out=500  finish=tool_calls  decision=handoff
│  └─ handoff policy_check  ok  stop_reason=final_answer  subtree.tokens in=5,000 out=1,000
│     ├─ model_call step=1  tokens in=2,000 out=500  finish=tool_calls  decision=call get_policy
│     │  └─ tool_call get_policy  ok  read  principal=service:agent  scope=policy:read
│     │               returned=2,060 B  delivered=2,060 B  stored=2,060 B  truncated=false
│     └─ model_call step=2  tokens in=3,000 out=500  finish=stop  decision=final_answer
├─ model_call step=3  tokens in=5,200 out=500  finish=stop  decision=final_answer
└─ summary  steps=3  model_calls=5  tool_calls=2  run.tokens in=17,600 out=2,500
            cost.estimate=0.0602 units  price_table=invented-v1
            final_status=success  stop_reason=final_answer

Timings, ids, content references and optional fields are left out, and the model and hash fields appear on the first call only. Neither tool call writes, so no target id appears. The summary’s steps=3 counts the lead’s loop, and model_calls=5 adds the subagent’s two.

Each call reads its agent’s fixed prefix plus everything written so far. The lead’s second call therefore reads 3,200 + 500 + 500 = 4,200 tokens (prefix, first output, tool result), and its third reads 5,200. The subagent starts from its own 2,000-token prefix, which holds the brief, and returns its 500-token answer to the lead.

What is the token rule?

The rule is one sentence: token counts are measured on model-call spans only, and any other span that shows tokens shows a sum of the model calls beneath it under a different field name. The lead’s calls read 3,200 + 4,200 + 5,200 = 12,600 input tokens and the subagent’s read 2,000 + 3,000 = 5,000, so the run total is 17,600. Output is 5 × 500 = 2,500.

Both roll-ups, the handoff’s (5,000 in, 1,000 out) and the summary’s (17,600 in, 2,500 out), repeat numbers that already exist on model calls. Suppose all three used the per-call name. A query that sums the attribute would return 17,600 + 5,000 + 17,600 = 40,200 input tokens, about 2.3 times the real figure.

The cost line uses invented prices of 2 units per million input tokens and 10 units per million output tokens. That gives 17,600 × 2 / 1,000,000 = 0.0352 and 2,500 × 10 / 1,000,000 = 0.0250, which sum to 0.0602 units.

With JavaScript on, the Agent cost-per-task estimator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Agent cost-per-task estimator on its own page to share a result by link.

The estimator opens on this run and shows 17,600 tokens in and 2,500 out. It reports the lead at 12,600 input tokens over 3 steps, against 9,600 for prefix times steps, and one worker at 5,000 in and 1,000 out. Brief and merge are zero because the brief sits inside the worker’s prefix.

Is the contract enough to find a wrong tool call at step 6?

It is now, and the first draft of the contract was not. I tested it on an invented failing run: every span reports ok, and at step 6 the agent calls a tool for changing an address on an order that has already shipped. Reading only required fields, the trace gives the cause in four moves.

  1. The model call at step 6. The decision is a call to the address tool. The responding model and tools.hash are the same at every step, and the definitions behind the hash include the redirect tool the agent should have called.

  2. Forward from step 1. Every tool span is ok. At step 4 a shipment lookup shows 11,840 bytes returned and 4,096 delivered, with truncated=true.

  3. The corroboration. Input tokens at step 5 grew by what a 4,096-byte result adds, far short of the full one.

  4. The proof. The stored copy is whole at 11,840 bytes, and the shipped status sits past byte 4,096. The rendered prompt at step 6 ends mid-array.

The first three moves read structure only, and the fourth needs the content store. The draft failed in three places. It had one size and one flag, which could not separate a cut that reached the model from a cap on the stored copy. It also lacked a step index and left the tool definitions optional.

All three are required now. The shapes themselves, truncation included, are sorted in AI agent failure modes, and each trace question in the agent bug bestiary reads a field from this contract.

One gap remains. A cut made upstream of the dispatcher, by a tool that does not declare it, leaves returned equal to delivered. Only the per-tool size distribution across runs shows it, which is why the book says to “record every tool result’s size as a span attribute”.

How does LLM tracing with OpenTelemetry map onto the contract?

LLM tracing with OpenTelemetry covers the contract in two layers of very different maturity: a stable span model and a vocabulary for AI fields that is still in development. Chapter 15 names the project only in a footnote, “named here as the category’s labeled example”, and describes its generative-AI conventions as “marked as still in development rather than stable”. The rest of this section is my reading of the primary pages on 6 October 2026.

The lower layer is settled. The Tracing API specification is headed “Status: Stable, except where otherwise specified”, and its span has a parent, timestamps, attributes and a status. Trace identity crosses processes through Trace Context, a W3C Recommendation since 23 November 2021.

What status do the AI attribute names have?

The names had Development status everywhere I looked on that date. Each generative-AI document in the conventions repository is headed “Status: Development”. The attribute registry page carried 127 Development badges and no Stable one, by a count made for this post. The only Stable attributes in its span tables (error.type, server.address and server.port) are borrowed from the core conventions.

This year the ground moved as well. The former pages read “This page has moved and is no longer maintained in this repository”. The new repository was created on 5 May 2026, had no tagged release, and the “Schema URL” section of its README read only “TODO”. Eight change notes labeled breaking were queued for its first release, one of them renaming a cache-token attribute.

An October 2025 issue on one agent framework reports an exporter built on “some deprecated semantic conventions” for message content. Change reaches the layer underneath as well. In March 2026 the project announced that it “is deprecating the Span Event API”, with new code writing events as logs tied to the span (Molkova, Pająk and Stalnaker, 2026). I did not check which released libraries emit the current names, so none of this describes adoption.

What does the mapping look like today?

It looks like the table below, which is a snapshot: names as printed in each specification on 6 October 2026, paired by me. The second column is the OpenTelemetry generative-AI conventions. The third is OpenInference, an independent convention that describes itself as “a semantic convention specification for AI application observability, built on OpenTelemetry” and states that “Every OpenInference trace is a valid OTLP trace”. Its pages state no stability level that I found, and I read its attribute names without reading every definition.

Contract field OpenTelemetry GenAI (Development) OpenInference
Span kind gen_ai.operation.name openinference.span.kind
Requested model gen_ai.request.model llm.request.model_name
Responding model gen_ai.response.model llm.response.model_name
Input tokens gen_ai.usage.input_tokens llm.token_count.prompt
Output tokens gen_ai.usage.output_tokens llm.token_count.completion
Cached input sub-total gen_ai.usage.cache_read.input_tokens llm.token_count.prompt_details.cache_read
Finish reason gen_ai.response.finish_reasons llm.finish_reason
Prompt version gen_ai.prompt.version llm.prompt_template.version
Rendered prompt, content gen_ai.input.messages (Opt-In) llm.input_messages
Output, content gen_ai.output.messages (Opt-In) llm.output_messages
Tools offered gen_ai.tool.definitions (Opt-In) llm.tools
Tool name gen_ai.tool.name tool.name
Tool call id gen_ai.tool.call.id tool_call.id; tool.id on the result side
Tool arguments, result gen_ai.tool.call.arguments, gen_ai.tool.call.result (Opt-In) tool_call.function.arguments inside the model’s output message; no name assigned on the tool span that I found
Session gen_ai.conversation.id session.id
Cost none found llm.cost.total
Prompt hash none found none found
Result sizes, truncated flag none found none found
Principal and scope none found user.id only
Session summary span none found none found

The two differ in design as well as spelling. OpenInference says its kind attribute “is required for all OpenInference spans”, and its content switches default to recording inputs and outputs unless hidden. The first convention classifies by operation name and marks every content attribute Opt-In.

The bottom four rows are the fields this post leans on hardest, and Chapter 15 predicts one of them: “Notice also what no generic convention will record for you: whose authority an action was taken on.” A convention names what a model client can see, and these four live in your harness.

Should the rendered prompt be stored in the trace?

Store a hash and a reference on the span, and put the content in a separate store with its own readers and its own expiry. The rendered prompt is the only complete answer to what the model saw, and the book warns that prompts and tool results “are where your users’ data lives”, in a store that “was not designed as a regulated data repository”.

The book’s instruction is to “treat content capture as a deliberate, redactable, off-by-default data flow, with retention limits and access rules, rather than a free debugging convenience”. The first convention agrees in its own words: “OpenTelemetry instrumentations SHOULD NOT capture them by default, but SHOULD provide an option for users to opt in.” The second defaults the other way, so check what your tracer does before assuming.

Option What the span holds What it costs
Reference plus hash hash, size and a pointer; content in a second store a second store to run and secure; a join during debugging; after expiry the trace says the input changed and cannot say how
Redaction content with sensitive values masked the redactor misses some values; the stored text no longer equals what the model saw, so take the hash first
Sampled content content for errored and flagged runs plus a fraction of the rest the run you need may be outside the sample; the keep decision needs the whole run
Inline everywhere full content as attributes personal data in an operations store; size limits; suits pre-production
Nothing hashes and sizes only the trace shows that something changed and never what

The convention recommends the first option “in production environments where telemetry volume is a concern or sensitive data needs to be handled securely”. It gives the pattern no standard shape yet, and the section ends “TODO: document a common approach to record references to externally stored content.”

Two legal texts pull in opposite directions, and this paragraph is context, with no legal advice in it. The GDPR’s data-minimisation principle limits personal data to what is “adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed”. The EU AI Act requires that high-risk systems “technically allow for the automatic recording of events (logs) over the lifetime of the system”.

How big does a trace get?

Size depends mostly on whether prompts are stored inline, because every model call re-sends the history. In the worked trace the five calls read 17,600 input tokens, and only 8,200 of them were new to their agent (3,200 + 1,000 + 1,000 for the lead, 2,000 + 1,000 for the subagent). Inline capture stores 2.1 times the new material, and the ratio grows with every step.

A 2024 issue describes the effect: “The same message ends up getting logged on its own many times, each time a new request with the same history is resubmitted”. Storing each rendered prompt once, keyed by its hash, removes the repetition.

On sampling, the book’s rule is to “sample the successes down to a fraction while keeping every errored run”. That decision needs the finished run, and the reference tail-sampling component “keeps spans in memory while it waits to make a sampling decision.” Chapter 15 quotes a survey of debugging practice saying most agent failures still return a successful status. So I would also keep runs by summary fields: a step or model-call count above the healthy range, a high cost, a low score.

How do you attribute cost without double counting?

Compute cost from the model-call spans and a price table you own, and store the price-table version beside the result. A stored cost disagrees with a recomputed one after any price change.

On the date I read them, the first convention had no attribute with cost or price in its name. A proposal to add one had been open since May 2025 with 29 comments. The second convention defines one.

Three errors recur in the issue trackers, and the token rule blocks the first:

  • A total on a parent under the per-call name. One issue (2025) states it exactly: queries that add up the values “will get a result that’s 2x too big”.
  • Cached tokens added to an input figure that already includes them. A 2025 report on one observability backend shows a displayed input total of 35,720 where the real usage was 17,903, of which 17,817 were cached.
  • Two libraries disagreeing on the same call. A September 2026 issue reports two instrumentation packages emitting different input counts for one cached call. The first convention’s text settles it on paper: the input count “SHOULD include all types of input tokens, including cached tokens.”

Treat cached and reasoning counts as sub-totals inside input and output, and never add them. One of the queued changes removes cache counts from the convention’s internal agent span because they “aggregate across models and inference calls, which makes them misleading”.

The book wants per-call tokens for diagnosis before billing: “Token counts per span are what make context bloat visible”. The shape of the bill is a separate subject, covered in how much an AI agent costs to run and in the explainer on why agent cost compounds. The contract’s job is to make the bill computable from a trace, and Chapter 19 expects that to happen fast: “real traces will retire the spreadsheet within a week”.

How do you keep the contract stable while conventions move?

Keep your own field names as the contract, translate to a public convention at the export boundary, and pin the version you translate to. The book gives the rule in two clauses: “pin the version you emit, and treat upgrades as deliberate migrations”. For the few attributes an alert depends on, it says to “keep copies under a namespace you control”.

Teams doing LLM tracing with OpenTelemetry face two risks, and practitioners weigh them differently. One agency guide (Metacto, 2026) argues that “the cost of switching from a vendor-specific schema later is much higher than the cost of tracking a stabilizing spec now”. Another author (Ganglani, 2026) takes the other side: “A small, forward-compatible custom schema gets you 80% of the value with 20% of the churn.” The first fears lock-in to a vendor, and the second fears churn in a vocabulary.

A mapping takes both seriously, through four habits:

  • One translation table. Contract field on the left, convention name on the right, with the date and the version it was read from.
  • Dashboards and alerts read your names. The field guide Chapter 15 cites (SentryML, 2026) is blunt: “do not build a paging alert on top of an attribute name the spec might rename.”
  • A fixture trace in the test suite. One recorded run, checked against the required fields after every upgrade of a tracing library.
  • Additions come from incidents. The book’s account of how a schema grows: “the fix is a new span attribute, which is how trace schemas actually mature—one unanswerable question at a time”.

How do you audit one existing trace against the contract?

Pull one real run with at least one tool call and tick only what the stored trace shows. A line you cannot verify from the record counts as a fail, and each fail becomes a ticket. The twelve lines work the same whether the spans came from LLM tracing with OpenTelemetry libraries or from a tracer you wrote.

  • Every span has a parent link, the tree has one root, and each tool span hangs under the model call that asked for it. Where to look: the parent ids in the raw span records.
  • Every model call shows the requested and the responding model as two fields. Where to look: one model-call span.
  • Every model call has a prompt version and a hash of the rendered prompt. Where to look: two calls with different inputs must show different hashes.
  • The exact text the model saw at one step can be retrieved, tool definitions included. Where to look: follow the reference from a late step.
  • Input and output tokens sit on each model call, and no other span reuses those names. Where to look: sum the attribute across the trace and compare with the provider’s count.
  • Every model call records a finish reason. Where to look: the last model call.
  • Every tool span has a name, arguments, a result and a status, with a structured error on failure. Where to look: one failed tool call.
  • Tool results show the size returned and the size delivered to the model. Where to look: the largest result in the run.
  • Tool spans record the principal and the scope, mark read or write, and name the target of a write by a pseudonymous id. Where to look: one call that writes.
  • Every handoff records the brief, the return and why the subagent stopped. Where to look: one delegation, or mark the line not applicable.
  • The run ends with a summary carrying steps, model calls, token totals, final status, stop reason and a cost with its price-table version. Where to look: the last span.
  • The content policy and the pinned convention version are written down, and no credential appears in any attribute. Where to look: the contract block, then search the trace for a known key prefix.

A trace that passes supplies what agent trajectory evaluation needs to grade a recorded run: tool names, arguments, results, step count and final status. A judge’s score can ride on the summary span as an optional field. Whether that score deserves trust is a separate question, taken up in is LLM-as-a-judge reliable.

Where does a span contract stop helping?

A span contract stops helping at the point where the record itself is wrong, or where the question is why the model chose as it did. A tracer that drops a span produces a clean-looking trace.

The contract is a synthesis from one book and two specifications. I know of no study showing that teams with these fields debug faster, and the worked trace is constructed. The mapping is a reading of an untagged branch on one day, so re-read the primary pages before you copy a name.

A span records authority after the fact. One commenter wrote in March 2026 that tracing “just shows you the execution path, not the authority boundary”. A principal on a tool span tells you afterwards who was acting; a permission check has to stop the call.

A complete trace still needs a reader. The TRAIL benchmark of 148 annotated agent traces (Deshpande and colleagues, 2025) reports that the best model it tested scored 11% on the benchmark. The reading method itself belongs to how to debug an AI agent from its traces.

Regulated settings add retention and access duties this post leaves out. A coding agent with no backend can keep the same contract in a local file of span records.

The one thing to keep

LLM tracing with OpenTelemetry gives you a stable tree and, as of this writing, a moving dictionary, so the durable artifact is the contract you write between them. Chapter 15 states the stakes: “what you fail to capture is precisely what you will not be able to debug, and you rarely get to choose, in advance, which incident you will need”. Run the checklist on one trace this week and commit the block.

The glossary entries for trace and span are free to read, as are the Preface and Chapters 1 and 2. Chapter 15, “Observability and Debugging” is in the full book, along with Chapters 5, 11, 17, 18 and 19, which supply the error, handoff, authority and cost rows. The guide to evaluating agents collects the neighboring posts and tools, or you can see the formats.

Questions readers ask

Are the OpenTelemetry GenAI semantic conventions stable?
As read on 6 October 2026, no. Every generative-AI convention document was headed Status: Development, the attribute registry page showed no Stable badge, there was no tagged release of the new repository, and breaking changes were queued. Emit the conventions pinned to a version and keep your own field names for anything an alert depends on.
What should an LLM span record?
The requested and the responding model, the prompt version, a hash of and a reference to the rendered prompt, a hash of the tool definitions offered, input and output tokens, the finish reason, start and end times, and what the model decided to do next. Each tool call gets its own child span with arguments, result, sizes and status.
Should I store prompts and completions in traces?
Store a hash and a reference on the span, and keep the content in a separate store with its own readers and its own expiry. The OpenTelemetry generative-AI convention read for this post tells instrumentations not to capture content by default. The same convention reserves inline content for cases such as pre-production environments.
Why do my token totals come out doubled?
Usually a total was copied onto a parent span under the same attribute name as the per-call value, so a query that sums the attribute counts it twice. The other common cause is cached tokens added to an input figure that already includes them. Sum model-call spans only, and give every roll-up its own field name.
What is the difference between the OpenTelemetry GenAI conventions and OpenInference?
Both are attribute vocabularies laid over the same span model. As read on 6 October 2026 they use different names, opposite defaults for capturing content, and only OpenInference defines a cost attribute. I found no prompt hash, tool result size or principal behind a tool call in either.

Sources

  1. OpenTelemetry Authors (2026). Semantic conventions for generative AI systems (repository main at commit cb10b70; Status: Development; read 2026-10-06)
  2. OpenTelemetry Authors (2026). Moved: Generative AI semantic conventions (stub pages, read 2026-10-06)
  3. OpenTelemetry Authors (2026). OpenTelemetry specification: Tracing API (Status: Stable, except where otherwise specified)
  4. OpenTelemetry Authors (2026). Tail Sampling Processor (README; stability: beta for traces)
  5. Liudmila Molkova, Robert Pająk, Trask Stalnaker (2026). Deprecating Span Events API
  6. W3C (2021). Trace Context (W3C Recommendation, 23 November 2021)
  7. OpenInference contributors (2026). OpenInference specification (README, semantic conventions, configuration; read 2026-10-06)
  8. Frank Chen (2026). AI Agent Debugging: A Trace-First Guide to Finding the Bug Fast
  9. SentryML Editorial (2026). OpenTelemetry GenAI Semantic Conventions: A 2026 Guide
  10. Kunal Ganglani (2026). OpenTelemetry Instrumentation for AI Agents
  11. Metacto (2026). LLM Tracing in Production: What to Capture and Why
  12. Deshpande, Gangal, Mehta, Krishnan, Kannappan, Qian (2025). TRAIL: Trace Reasoning and Agentic Issue Localization
  13. European Parliament and Council (2016). Regulation (EU) 2016/679 (General Data Protection Regulation), Article 5
  14. European Parliament and Council (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 12
  15. alexmojaki (GitHub) (2025). Totals recorded on an outer span under the same attribute name (GitHub issue #19)
  16. adriangb (GitHub) (2025). Proposal for cost attributes on generative-AI spans (GitHub issue #287)
  17. flagbug (GitHub) (2025). Cached input tokens counted twice in a displayed total (GitHub issue #10592)
  18. stephentoub (GitHub) (2024). Message history logged again on every request (GitHub issue #1621)
  19. roy-tong (GitHub) (2026). Two instrumentation packages emitting different input token counts for one cached call (GitHub issue #4449)
  20. zacbrannelly (GitHub) (2026). Tool definitions missing from the recorded prompt (GitHub issue #789)
  21. cephalization (GitHub) (2025). An exporter emitting deprecated generative-AI attribute names (GitHub issue #8829)
  22. chirdeeps (Hacker News) (2026). Hacker News comment on tracing and authority boundaries between agents