What is the three-set discipline for an LLM classifier?
The three-set discipline is a way of dividing your labeled examples so that the person who writes a classifier’s instructions cannot also, without noticing, grade them. The splitter above takes rows of text and label and deals them into the three sets that Chapter 26 of the book describes. It checks that no text sits in two of them, and it produces the files you need to keep the split honest.
The classifier in question has no training step. It is a model given a written guide: definitions of each class, boundary rules, and a few worked examples. The chapter’s instruction for the labels you have is one sentence: “Divide those labels into three disjoint sets with three jobs.”
- In-guide demonstrations. These are “the handful of examples living inside the classifier itself, chosen by the standards of the examples section”.
- The development set. This is “a labeled sample you re-run after every revision of the guide, whose failures you read freely and respond to”.
- The held-out test set. The chapter calls it the sealed envelope: “labeled once, locked away, consulted rarely, and the only number you ever report as the classifier’s accuracy”.
The glossary entry gives the short definition. The rest of this page explains each choice the splitter makes and says which are the book’s and which are the tool’s.
How do you fool yourself without the split?
You fool yourself by measuring the guide on the same examples you wrote it from. The chapter tells the story step by step. You draft the guide with twenty real tickets open in another window. You run the classifier on those twenty and get nineteen right. You add a clause for the twentieth and reach twenty out of twenty.
The chapter’s verdict on that score: “The number is real, and it describes almost nothing you care about, because the instrument was fitted to those tickets; you have measured the textbook on its own worked examples.”
The risk does not go away because there is no training. Of overfitting, the chapter says a classifier with no training “is exactly as immune as you are.” Each edit that chases one failure in the development set moves the guide toward those particular rows. The sign to watch for is “a widening gap: the development number climbs while the sealed number stands still, or slips.”
Why split evenly between development and test?
You split evenly because the examples that used to go to training are no longer needed there. In classical machine learning most labels feed the training set and evaluation gets what is left. Here the guide needs only a handful of examples, and the chapter draws the conclusion: “every label you can afford now buys measurement”.
So the arithmetic turns over, and “something close to an even division between development and test becomes defensible”. The chapter’s example is two hundred labeled tickets, “a hundred to iterate against and a hundred sealed”, and it marks those figures as illustrative.
The splitter starts at 50% and lets you move the share between 10% and 90%. Small classes and rounding can move the share you get away from the share you asked for, so the result states the share actually reached. When that falls outside 40% to 60%, a band that is the tool’s own, it also says how many rows the smaller set is left with.
How large each set needs to be is a separate question. The eval sample size calculator answers it: you give it the accuracy you expect and the margin you can tolerate, and it returns a count.
How does the splitter deal the rows?
The splitter deals the rows class by class, in an order fixed by a seed. All of this section is the tool’s own logic; the chapter says what the sets are for and leaves the mechanics open.
- Duplicates come out first. Two rows count as the same text when they match after trimming, lower-casing and collapsing runs of spaces. The first copy stays and later copies are left out.
- Each class is shuffled with a small seeded generator. The same rows and the same seed always give the same split, so a colleague can reproduce it from the manifest.
- Demonstrations are taken from the front of each class, as many as you asked for. A class never gives up its last two rows, so every class with two or more rows still has something to measure with.
- The rest is divided by the development share. A class with two or more rows left gets at least one in each set, and a class with a single row sends it to development. Rounding is carried from one class to the next so the totals stay close to the share you chose.
Splitting class by class is what statisticians call stratifying. It matters most when one class is rare: a plain shuffle of a hundred rows puts all four examples of a rare class on the same side about one time in eight.
The result table shows the counts for every class. A class with fewer than five rows in the test set is marked thin. Five is the tool’s cut-off. A per-class score from four rows moves by 25 points when one ticket changes, which is too coarse to act on.
Why are the demonstrations only placeholders?
The demonstrations are placeholders because the tool picks them at random, and the chapter wants each one chosen for a reason. The splitter cannot read your tickets for meaning. It reserves the slots so the counts add up, and it tells you to fill them yourself.
The chapter’s standard for a demonstration is that it plays a role: “Three roles cover nearly every example that earns its tokens.” The prototype is “the unmistakable center of a class”. The edge “sits deliberately close to a boundary and shows which side it falls on”. The contrastive pair is two inputs “near-identical in wording, on opposite sides of one boundary”. When the labels are levels on an ordinal scale, the pairs go between adjacent levels.
More examples do not help much. On in-context learning the chapter is blunt: “Examples are a stencil.” The model reads them and does not accumulate knowledge from them, “and the second fifty teach almost nothing the first three did not.”
When you swap a placeholder for a better example, move the row you removed into the development set. Take the replacement from the development set too. Never take it from the test set: choosing it would mean reading the sealed rows.
What counts as a leak between the sets?
A leak is any row the classifier has effectively seen before it is measured on it. The chapter’s sentence on this is the reason the tool exists: “An example that sits in the guide and also in a measurement set is a leaked exam answer, and the metric it inflates has quietly stopped measuring classification.” It adds: “Disjointness is the entire point.”
The splitter guards against the mechanical version of this, which is the same text pasted twice. It also reports a second problem that the same check turns up: one text carrying two different labels. Only the first copy is placed, and the result tells you to settle the label before measuring. A disagreement between labelers is information about your class definitions, so it is worth reading before you delete either row.
The check has a limit you should know. It compares text after light cleaning and nothing more. A ticket and a reworded copy of it, or two tickets from the same customer about the same incident, will pass as different rows. If your data has near-copies of that kind, group them before pasting so they land on the same side.
When should you open the sealed test set?
You should open the sealed test set at milestones and not between them. The chapter’s guards against overfitting the guide are procedural. The first: “Prefer revisions that state a general rule to revisions that enumerate a case.” The second is short: “Open the envelope only at milestones.”
The third guard applies when the test set has been looked at so often that its failures steer your edits. The chapter’s phrase for that state is that the test set “has become a second development set”, and the remedy is to “fold its labels into development and label a fresh envelope.”
The manifest supports this with one line to keep by hand: “Test set opened: 0 times”. Add a line with the date and the milestone each time. The file also records the seed, the settings, the counts and a short fingerprint of each set. The fingerprint is a simple checksum that changes if a row is added, removed or relabeled, or if its third column changes. It ignores changes of case and spacing in the text. It detects accidental changes and is not a security measure.
What goes in the changelog?
The changelog records each revision of the guide and what it did to the development score. The chapter asks for “every revision recorded with what changed, which observed failure motivated it, and the development score before and after.” It calls the result “this paradigm’s training log”.
The skeleton the tool produces has those columns, plus a second table for the times the envelope was opened. The habit that goes with it is to re-run the whole development set after every edit, “because no edit is local”. A clause tightened for one class can pull rows across a boundary two classes away.
What should you read besides accuracy?
You should read accuracy against what a classifier that learned nothing would score. On a skewed class mix, accuracy flatters. The chapter’s example is a mix where seven tickets in ten are bugs. There, “a classifier that answers bug unconditionally scores seventy percent while carrying no information”.
The splitter reports that floor for your own test set: the share of the largest class. The arithmetic is the tool’s. For one honest number the chapter points to Cohen’s kappa, which the judge agreement calculator computes from pasted pairs of labels. That calculator also lists which classes get confused with which. The chapter calls that table “a to-do list”, because each recurring swap is the next contrastive section to write.
What does the splitter leave to you?
The splitter leaves the labeling, the choice of demonstrations and the discipline to you. It cannot tell whether your labels are right, whether your sample looks like real traffic, or whether you opened the test file last night.
Its choices are labeled where they appear, and none of them is the chapter’s: the class-by-class split, the random placeholders, the duplicate rule, the fewer-than-five warning and the fold numbers. Where your situation calls for something else, such as grouping rows by customer, do that first and paste the result.
The chapter closes the section by noting how light the machinery is, “two labeled samples, a changelog, a habit of counting”, and what it is for: “keeping the person who writes the definitions from also, in effect, grading them.” The guide to agent evaluation places this beside eval sets and judges, and the full chapter, with the anatomy of the guide before it and scoring after it, is Chapter 26, Agents as Classifiers and Scorers (in the full book).
Questions readers ask
- Why split 50/50 and not 80/20?
- Because the 80 was for training, and there is no training here. A classifier built from a written guide needs only a handful of examples in the prompt, so nearly every label can go to measurement. Chapter 26 says the arithmetic inverts and a near-even division between development and test becomes defensible. Its example is two hundred labeled tickets, a hundred to iterate against and a hundred sealed, with the figures marked as illustrative.
- When should I retire the test set?
- When you have looked at it often enough that its failures are steering your edits. At that point it is working as a second development set, and the chapter says to name that and retire it: move its labels into the development set and label a fresh one. The manifest's count of how many times the set was opened is there so you can see this coming.
- Why must the prompt examples be excluded from both measurement sets?
- Because the classifier has already been shown the answer to them. An example that is in the guide and also in a measurement set gets classified correctly for a reason unrelated to the guide's quality, and the score goes up without the classifier getting better. The chapter calls it a leaked exam answer.
- Does my data leave the browser?
- No. The page makes no network request with your rows, and the split, the duplicate check and the five files are all produced by the script running in your tab. The rows are also kept out of the page's link: only the four settings are stored there, so sharing or reloading the link never carries your data.
- Can I use cross-validation instead of a sealed test set?
- Use it beside the sealed set, for a different question. Rotating which slice is held out tells you how much the score moves from slice to slice. It cannot give you a clean final number, because you are the one revising the guide and you remember the failures you studied in earlier folds. The tool numbers folds inside the development set and leaves the test set alone.
Sources
- Trevor Hastie, Robert Tibshirani, Jerome Friedman (2009). The Elements of Statistical Learning, 2nd edition, Chapter 7: Model Assessment and Selection
- Gavin C. Cawley, Nicola L. C. Talbot (2010). On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation (JMLR 11:2079–2107)
- Cynthia Dwork, Vitaly Feldman, Moritz Hardt, et al. (2015). The reusable holdout: Preserving validity in adaptive data analysis (Science 349(6248):636–638)
- Shachar Kaufman, Saharon Rosset, Claudia Perlich, Ori Stitelman (2012). Leakage in Data Mining: Formulation, Detection, and Avoidance (ACM TKDD 6(4))