Home / Blog / Evaluating and observing agents / Building a Golden Dataset for LLM Evaluation Fr…

Evaluating and observing agents

Building a Golden Dataset for LLM Evaluation From Real Traces

Build a golden dataset for LLM evaluation from production traces: what to sample, how to label, how to split and when to retire it. Copy the case record.

By Enrique Gutiérrez · Published · 19 min read

A golden dataset for LLM evaluation is a versioned set of cases built from real runs. Each case holds an input, a fixed world to run against, a reviewed statement of what right looks like, a grader, and a record of where the case came from. What makes it golden is the review and the bookkeeping.

I wrote this for an ML engineer who has a trace store and no eval set. The advice you have already met is one sentence long: collect real traces and curate from them. One engineer on Hacker News put it that way in March 2026: “collecting real traces, and then hand-curating a golden dataset from those traces.” This tutorial is the procedure behind the word “curating,” in seven steps, with one worked example carried from 120 traces to 49 cases.

Two things are left to other posts on purpose. How many cases you need is a statistics question, answered in how many eval examples you need. What the gate does with the set once it exists is in the post on regression testing for LLM applications.

What is a golden dataset for LLM evaluation?

A golden dataset for LLM evaluation is an eval set whose expectations a person has reviewed and whose history is written down. Chapter 16 of AI Agents, Engineered, “Evaluating Agents,” (in the full book) defines the eval set, in “Building an Eval Set,” as “a curated collection of tasks, each an input plus some way to grade the output: a reference answer, a checkable condition, or a rubric describing what good looks like.”

Two neighboring terms get confused with it. A golden trace is a different object: Chapter 15 defines golden traces as “recorded runs whose behavior you consider correct,” replayed against new code. A golden trace is a recording you replay; a golden case is a task you run fresh and grade.

And “golden” does not mean a stored perfect answer. For an agent the expectation is often an end state or a property, because many different outputs are correct.

The reason to build from traces and not from imagination is in the same section of Chapter 16: “A suite grown this way tracks what your agent actually gets wrong; a suite written in an afternoon of imagination tracks what you feared in advance.”

The seven steps, in order:

  1. Sample traces from five sources, with a quota for each.
  2. Label each trace against a written rule, with a second labeler on a sample.
  3. Turn each promoted trace into a case record.
  4. Scrub private data, and check the case still behaves the same.
  5. Group near-duplicates, then split whole groups into sets.
  6. Version the set, and log every opening of the sealed part.
  7. Retire and refresh on a schedule.

Which traces should go into the set?

Draw from five sources and fix a quota for each before you look, because each source is blind in a different place. The five sources and the quotas are my arrangement. The limits of each method are from the evals FAQ by Hamel Husain and Shreya Shankar (modified September 2026), which tabulates sampling methods with one limitation apiece.

That FAQ says of random sampling that “A small batch can miss rare cases,” of using a classifier to flag failures that “It favors problems the classifier already knows how to find,” and of feedback that “It misses problems that users do not report.” Its advice for the mix is short: “Keep some random traces in every batch.”

Source How to draw it What it finds Where it is blind Use it for
Random Uniform over all runs in a time window Typical traffic, and failures no signal describes Rare cases, in a small batch rate, unknown failures
Worst by signal Sort by step count, cost, retries or a low judge score; take the top Failures that signal describes Failures that look normal on that signal known failures
User-flagged Negative feedback, complaints, human takeovers Failures a user noticed Failures nobody reported known failures
Stratum quota A fixed minimum per intent, tool or customer type Rare kinds of request Needs a stratum label on each trace first rare cases
Costly action Every run in the window that called a tool that moves money, deletes or sends outside Errors that are rare and expensive; runs where not acting was right Says nothing about ordinary requests costly cases

One rule follows from the table, and it is easy to break: only the random slice estimates a rate. The other four are chosen because they are likely to hold trouble, so the share of failures in them describes your sampling plan, not your product.

The quota for rare strata is arithmetic. If a kind of request is 2% of traffic, a random batch of 60 traces is expected to hold 1.2 of them, and holds none with probability 0.9860, about 30%. A random sample is the wrong tool for that stratum; a quota is the right one.

Passes belong in the set too. The lab guide to agent evals warns that “One-sided evals create one-sided optimization” (Anthropic, January 2026). The book’s version, in Chapter 16: “Whatever the rung, include negative cases.” A set made only of failures cannot report that a change broke something that used to work.

How do you label a trace so the label means something?

Label each trace pass, fail or cannot tell against a written rule, record which version of the rule you used, and have a second person label a sample without seeing your labels. The post on agent observability covers reading a trace and gives a review note to fill in; this section starts where that note ends.

The rule states what an acceptable end state is for each kind of request. Chapter 16 gives the test for whether a rule is sharp enough: “if two domain experts, given your criterion and an output, can disagree in good faith about whether it passed, sharpen the criterion, not the agent.”

The second labeler is how you find out. The FAQ’s instruction is to “Have annotators label the same examples independently before they discuss them.” When the two labels differ, ask which clause of the rule allowed both readings, rewrite that clause, raise the rule’s version number, and relabel the traces the clause touches. For cases that stay ambiguous, the book says to “route them to one designated human arbiter rather than a committee.”

A reviewed label can still be wrong, and at a measurable rate. A study of ten widely used benchmark test sets estimated “an average of at least 3.3% errors across the 10 datasets” (Northcutt, Athalye and Mueller, 2021).

Those were classification labels on public benchmarks, a different task from yours. The transferable point is that a test set labeled once and never audited carries errors. So when a case fails and the failure looks unfair, suspect the label first.

How do you turn one trace into a case?

A trace becomes a case when you can write down five things: the input, the world it ran against, the expected end state or property, the grader, and where it came from. Three conditions decide whether a trace is promoted at all, and they are set out in the observability post: the behavior is one you intend to guard, the expectation can be written as a check, and the input and tool results can be rebuilt with no live dependency.

Chapter 15 describes a failing trace as “two assets wearing one file format”: the reproduction, and the seed of an eval case. The case keeps the input and the recorded world. It drops the transcript as a reference, because a case that asserts the old run’s exact steps fails every valid alternative path. The post on agent trajectory evaluation covers what to check on a path.

Chapter 16 adds one requirement that a trace cannot supply: “every task must be solvable, and the way to know is to include a reference solution that proves it.” For a promoted pass, the recorded run is the reference solution. For a promoted failure, someone has to produce one.

The record below is the row in the dataset. It wraps the test itself: the frozen world and the checks are specified in the test-case block in how to test AI agents, and this record adds what a dataset needs on top, which is provenance, a group, a label history and a set.

GOLDEN CASE RECORD

Case id        : [stable id, never reused]
Status         : [active | quarantined | retired]   Reason if not active: [...]
Added in       : [set version]                      Retired in: [set version or -]

PROVENANCE
  Source trace : [trace id, date of the run]
  Drawn by     : [random | worst by <signal> | user-flagged | stratum quota | costly action]
  Stratum      : [intent, tool or customer type used for quotas and for the split]
  Group id     : [session, user, incident or input template shared with other cases]

THE TASK
  Input        : [the request, scrubbed]
  World        : [fixture name and revision; frozen clock; recorded tool results]
  Expectation  : [end state or property, written out; "must not" lines for negative cases]
  Grader       : [code check | calibrated judge | person]
  Reference    : [a known output or action sequence that passes the grader]

LABEL
  Verdict on the source run : [pass | fail]
  Rule version              : [...]
  Labeler                   : [name]     Second labeler: [name or "not sampled"]
  Disagreement              : [none | resolved by <arbiter>, rule clause changed: ...]

SCRUB
  Replaced     : [names, ids, addresses, free text: what was mapped to what kind of value]
  Kept         : [the structure the behavior depends on]
  Checked      : [the source verdict reproduces on the scrubbed case: y/n]

SET
  Set          : [development | regression | sealed | in the prompt (excluded from all three)]
  Moved        : [date, from -> to, why]

A case used as an example inside the prompt belongs to none of the three sets. The book’s Chapter 26 has the phrase for what happens otherwise: “a leaked exam answer.”

How do you scrub private data without breaking the case?

Replace private values with consistent stand-ins, keep the structure the behavior depends on, and then check that the scrubbed case still produces the verdict the original run did. Deleting fields is the mistake: a failure that depended on a customer holding two subscriptions disappears when the second one is blanked.

The reason to scrub at all is that promotion makes a new copy of user data in a new place. Chapter 15, in “Traces and Spans,” says of the store the copy comes from: “A trace store is typically access-controlled for an operations team, and it was not designed as a regulated data repository.” An eval set in a code repository is read by more people and kept for longer.

The steps, which are mine:

  1. Map identifiers consistently. The same customer id becomes the same fake id in the input and in every recorded tool result, so references still resolve.
  2. Keep shape. Preserve counts, currencies, the order of events and the gaps between dates. Shift the dates and freeze the clock.
  3. Rewrite free text. Names, addresses and anything a user typed about themselves get replaced, keeping length and language.
  4. Read the result. The FAQ’s warning applies: “Redaction tools can miss sensitive information, so check their output.”
  5. Re-run the check. The same FAQ: “Check that those edits preserve the behavior you need to evaluate.” A promoted failure should still fail on the version that produced it.

What you may store, and for how long, is decided by your own data rules and not by this post. If a trace cannot be scrubbed and kept, write a new case that has the same structure and mark its source as “rewritten from” the trace.

How do you split the set, and where does leakage come from?

Split by group into a development set and a sealed set, and let a regression set fill from the development set over time. This is the three-set discipline applied to an agent’s cases. The book states it for a classifier in Chapter 26, and the figure below draws that version: examples inside the guide, a development set and a sealed test set. The mapping to an agent’s eval set in the table is mine: in-prompt examples are kept out of every set, and a regression set takes the third place.

The chapter’s development discipline made spatial: one pool of labeled tickets divided into three disjoint homes.
Figure 26.4 The chapter’s development discipline made spatial: one pool of labeled tickets divided into three disjoint homes. A handful live inside the guide as demonstrations; the development set is the sparring bench you re-run and study freely; and the sealed test set (in accent) is the envelope opened rarely, the only number you report. A ticket that tries to live in two sets at once is a leaked exam answer—struck out here, because disjointness is the whole point. Reuse this diagram
Set Its job Who reads the failures What its number means
Development The cases you iterate on after every change Anyone, freely Progress on known problems
Regression Cases the system already passes, run by the gate Whoever the gate stops Nothing broke; expected near 100%
Sealed Cases nobody has tuned against, opened at milestones Nobody who edits the system The estimate you report

The regression set starts empty. Chapter 16 describes how it fills: “when a capability task stabilizes, promote it to the regression suite and add something harder behind it.” What “stabilizes” means across repeated runs is the subject of pass@k versus passk. The split for a classifier’s labels, which adds the in-prompt examples as a set of their own, is in the post on building an LLM classifier.

Leakage reaches the sealed part of a golden dataset for LLM evaluation by three routes.

Shared examples. A case that appears in the prompt and in a measurement set is answered from the prompt.

Near-duplicates across the split. Two cases from one session, one user, one incident or one templated input are nearly the same case. If one lands in development and the other in the sealed set, fixing the first fixes the second, and the sealed score rises without the system getting better at anything unseen.

This has happened to careful people: an audit of two heavily used image benchmarks found that “3.3% and 10% of the images from the test sets of these datasets have duplicates in the training set” (Barz and Denzler, 2019). The fix is to assign a group id before splitting and to deal whole groups.

Your own peeking. Each time you read a sealed failure and then edit the system, or pick among variants by the sealed score, information flows from the set into the system. The evals FAQ makes the point about tuning a judge prompt on a development set: “Each time you use dev results to change the prompt or choose a model, information from those examples influences the evaluator.” The same holds for any set you choose by.

A small calculation shows the size of the third effect. Suppose five variants each truly pass 80% of cases, and you score all five on a 24-case sealed set and keep the best. The expected score of the winner is 89.0%, nine points above the truth, with no variant better than another. With ten variants it is 91.6%.

This is my arithmetic, and it assumes the variants’ results are independent, which overstates the effect for variants that are small edits of one another.

What does the splitter do, and what does it not?

The three-set splitter deals rows into sets stratum by stratum, leaves out any text that appears twice, and writes a manifest with a line for counting openings. It was built for a classifier’s labels, so two adjustments make it fit here: paste one row per group (a representative input, a comma, the stratum), and remove any case used in the prompt before pasting. Every case in a group then follows its row.

Its duplicate check is exact: the same text after trimming, lower-casing and collapsing spaces. It will not catch a paraphrase. Grouping is your job, before the paste.

The embed opens empty, set to no in-prompt examples, an even split and seed 53. Press “Load the sample” to fill it with the tool’s own invented tickets, which are a classifier’s rows and not the worked example below. On those settings it reads 58 rows, leaves out one duplicate and one row that repeats a text under a different label, and places 56: 28 in development and 28 sealed.

With JavaScript on, the Three-set splitter runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Three-set splitter on its own page to share a result by link.

Worked example: 120 traces to 49 cases

This is a worked example with illustrative numbers. The system is a support agent for a subscription product, and the trace store holds two weeks of runs.

Sample. The reading budget is 120 traces: 40 random, 25 worst by step count, 20 user-flagged, 20 from a quota on rare intents and 15 that called a refund or deletion tool.

Label. Against rule version 1, the verdicts are 71 pass, 41 fail and 8 cannot tell. The random slice holds 5 failures in 40, or 12.5%. The other four slices hold 36 in 80, or 45%.

Pooling them gives 41 in 120, about 34%, a number that describes the sampling plan. Only the 12.5% says anything about traffic, and with 40 traces its interval is wide.

A second person labels 30 of the 120 independently. The two agree on 26, and Cohen’s kappa is 0.73. All four disagreements turn on one question, whether a refund the tool reported as queued counts as done. The rule gets a clause and becomes version 2. The clause adopts the first labeler’s reading, that queued is not done, so relabeling the affected traces leaves the counts at 71, 41 and 8.

Promote. Of the 41 failures, 6 have no expectation that can be written as a check and 4 cannot be rebuilt because a tool result was not recorded. That leaves 31 failure cases.

From the 71 passes, 24 are promoted, spread across the intents, 8 of them negative cases. Candidates: 55.

Five of the traces, followed through:

Trace Drawn by Verdict Becomes Group Set
A billing question answered in three steps Random Pass A pass case, stratum “billing question” Its own Development
A plan lookup repeated 31 times on an empty result Worst by signal Fail A failure case: stop on a repeated empty result Its own Development
A refund reported done while the ledger shows it rejected User-flagged Fail A failure case: claim completion only on a settled status One incident, six traces Sealed, with its group
A customer with two subscriptions asks to cancel one; both are cancelled Stratum quota Fail A failure case with a state check on both subscriptions One user, two traces Development, with its group
A request to delete an account with no identity check; the agent declines Costly action Pass A negative case: no call to the deletion tool Its own Sealed

Group. The 55 candidates fall into 43 groups: one of six cases (the refund incident), two of three, three of two and 37 singletons. A cap of two cases per group, keeping the two hardest, removes 6 and leaves 49 cases. One vendor’s guide gives the same advice for near-duplicates: “keep the hardest one or two variants and drop the rest” (Langfuse, read October 2026).

Split. Whole groups are dealt stratum by stratum: 22 groups (25 cases) to development and 21 groups (24 cases) to the sealed set. The regression set is empty on day one.

Had the 49 cases been split by row, each of the six two-case groups would straddle the split with probability 0.51, and at least one would straddle in about 99% of random splits (my calculation, by simulation).

Whether 24 sealed cases can answer the question you will ask of them is a sample-size problem, and for a small difference between two versions they cannot. The sample-size post gives the arithmetic.

How do you version the set and count openings?

Give the golden dataset a version that changes on any addition, removal, or change to a label or an expectation, and store that version with every score. A score is comparable only with scores on the same version. To compare across versions, re-run the old system on the new version. The FAQ accepts the cost plainly: “As your eval set changes, its scores may no longer be directly comparable with older scores.”

The sealed set needs one more record: an opening log. Each line holds the date, the milestone, the system version, the score, and who saw what. When you seal the set, write down a budget of openings. The budget is my addition; Chapter 26 says only “Open the envelope only at milestones.” Five is an illustrative starting value; the right number is whatever you are willing to defend.

There is a tension here with advice the book treats as the most important in its chapter: read the failures, because some of them are grader bugs. The way through is to separate the roles. Someone who does not edit the system audits the sealed failures for fairness. If the person tuning the system reads them, the set has been spent.

When do you retire cases, and when the whole sealed set?

Retire the sealed set when either of two things is true: the person tuning the system has read its failing cases, or its openings have reached the budget. Chapter 26’s instruction for that moment is to say that “it has become a second development set,” then “fold its labels into development and label a fresh envelope.” The two triggers are my reading of “consulted often enough.” The first counts readings after the split. Whoever labeled the traces has seen every source run once, sealed cases included. That look is unavoidable, and it is the reason to split before any fix is written and, where two people are available, to give the sealed part’s labels to the one who will not tune.

Expect the fresh set to score differently for reasons unrelated to peeking. When researchers rebuilt two image benchmarks by the original recipe, models lost 3% to 15% on one and 11% to 14% on the other, and the authors concluded that the drops “are not caused by adaptivity” (Recht and colleagues, 2019). A new sample is a different sample. Score the current system on both sets before you retire the old one, so the step is on record.

A single case is retired for one of four reasons, and its record says which:

  • The behavior no longer exists.
  • The expectation became wrong, for instance after a policy change.
  • The case cannot be passed, or its grader is wrong and cannot be fixed.
  • It duplicates a harder case in the same group.

A case that always passes is not retired for that reason. It moves to the regression set. Retired cases stay in the file with their status, so old scores remain readable.

How do you keep the set representative as traffic drifts?

Repeat the sampling plan on a schedule, compare the set’s mix of strata with the latest window of traffic, and add cases where the two have come apart. The FAQ’s summary of the problem is “Eval datasets naturally get stale as your product and users change.” Chapter 16 names the remedy as feeding the set from production: “a suite fed this way cannot drift far from reality, because reality is its diet.”

Three checks at each refresh, which are mine:

  1. New strata. Does the latest random slice contain a kind of request with no case in the set? Add a quota for it.
  2. Missing strata. Does the set hold a stratum that traffic no longer contains? Retire those cases under the first reason above.
  3. New failures. Label the new random slice with the current rule. Failures with no matching case go to development; a fresh share of new groups goes to the sealed set, unread.

New sealed cases must come from traces the person tuning the system has not studied. That is the reason to add to the sealed set from the random slice at the moment of sampling, before anyone reads the failures.

Where does this procedure break?

Building a golden dataset for LLM evaluation from traces breaks first where there is no traffic. With no traces, the first set is written from imagination and from the checks you already run before a release, and it is replaced as traces arrive. The lab guide’s starting size applies either way: “20-50 simple tasks drawn from real failures is a great start.”

It also breaks where the world cannot be rebuilt. If tool results were not recorded, a trace cannot become a case, which is an argument for the capture described in LLM tracing with OpenTelemetry.

A small sealed set gives a wide interval, and no discipline about openings narrows it. The second labeler measures agreement on a sample, not correctness.

Last, the set describes the traffic it was drawn from. Chapter 16 says it of every suite: “Suites drift from usage, agents occasionally learn a loophole in a grader, and the world moves.”

The one thing to keep

A golden dataset for LLM evaluation is a claim that these cases, with these expectations, stand for what the system must do. The claim is only as good as its paperwork: where each case came from, who judged it and by which rule, which set it sits in, and how many times the sealed part has been looked at.

The eval-set material is in Chapter 16, Evaluating Agents (in the full book), and the trace-to-dataset habit closes Chapter 15, Observability and Debugging (in the full book). The guide to evaluating and observing agents collects the related posts and tools, and you can see the formats.

Questions readers ask

What is a golden dataset in LLM evaluation?
A golden dataset is a versioned collection of evaluation cases whose expectations a person has reviewed. Each case holds an input, the fixed world the system runs against, the expected end state or property, the grader that checks it, and provenance: the source trace, the labeler, the rule version and the set the case belongs to.
Should a golden dataset contain only failures?
No. A set made only of failures cannot tell you when a change breaks something that used to work, and it rewards a system that refuses or over-acts. Keep failures, a slice of passes spread across the kinds of request you serve, and negative cases where the right behavior is to do nothing.
How do I split a golden dataset into dev and test sets?
Group the cases first, so that cases from the same session, user, incident or templated input share a group, then deal whole groups into a development set and a sealed set, stratum by stratum. With no training set to feed, something close to an even division is defensible. A regression set then fills over time with development cases the system passes reliably.
How does leakage happen in an LLM eval set?
Three ways. A case used as an example in the prompt also sits in a measurement set. Two near-identical cases land on opposite sides of the split. Or the person tuning the system reads the sealed set's failures, or picks among many variants by its score, and the set stops being unseen.
When should I retire a golden dataset?
Retire the sealed set when the person tuning the system has read its failing cases, or when the number of openings reaches the budget written down when it was sealed. Retire a single case when the behavior no longer exists, the expectation became wrong, the case cannot be passed, or it duplicates a harder case.

Sources

  1. Anthropic (2026). Demystifying evals for AI agents
  2. Hamel Husain and Shreya Shankar (2026). AI Evals: Everything You Need to Know (FAQ)
  3. Björn Barz and Joachim Denzler (2019). Do We Train on Test Data? Purging CIFAR of Near-Duplicates
  4. Curtis G. Northcutt, Anish Athalye and Jonas Mueller (2021). Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
  5. Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt and Vaishaal Shankar (2019). Do ImageNet Classifiers Generalize to ImageNet?
  6. Langfuse (undated; read October 2026). Golden dataset evaluation: build and maintain LLM test sets
  7. kbdiaz (Hacker News) (2026). Hacker News comment on hand-curating a golden dataset from traces
  8. senordevnyc (Hacker News) (2026). Hacker News comment on labeling traces and freezing passes