To build an LLM classifier without a training set, write a label spec that defines every class at its edges, ask the model for a structured verdict with its reasons first, and test it on hand labels it has never seen. The prompt replaces the training set. The labels don’t disappear; they move from training to measurement.
A classifier assigns each input one label from a fixed list: a support ticket to billing or bug, an expense claim to in-policy or out. The classical way to build one is to label thousands of examples and train a model on them. A language model arrives with its learning already done, so you configure it with words instead, and the finished classifier is a document you can read, diff and version.
That convenience has a trap in it, and this post is mostly about the trap. It follows Chapter 26, Agents as Classifiers and Scorers (in the full book) of AI Agents, Engineered, whose one-line summary of the approach is that “the definitions are the model.” Every number in the worked example below is illustrative and was computed with this site’s judge agreement calculator, which you can run on your own labels.
How do you build an LLM classifier without a training set?
You build it in six steps: write the label spec, choose a handful of in-prompt examples, fix an output contract that puts reasons before the label, hand-label a development set and a sealed test set, iterate against the development set only, then open the test set once and report what it says. Each step has a failure it prevents.
| Step | What you produce | What it prevents |
|---|---|---|
| 1. Label spec | Per class: essence, includes, excludes, boundary rules; an exit class | The model settling borderline cases its own way, differently each run |
| 2. Examples | A few prototypes, edges and contrastive pairs, each with a reason | A catalog that buries the definitions and costs tokens on every call |
| 3. Output contract | Rubric answers, then a closed label, a review flag, a short rationale | Snap verdicts, unparseable prose, hard cases hidden as confident labels |
| 4. Two labeled sets | A development set and a sealed test set, disjoint from the examples | Grading the classifier on the cases it was written against |
| 5. Iteration | A changelog: each edit, the failure behind it, the scores before and after | Whack-a-mole edits that fix one class and break another |
| 6. One sealed number | Kappa, flag rate and confusions on the test set | Reporting the number you tuned instead of the number you’d get |
The order of the steps matters less than the separation between them. Steps 1 to 3 are the classifier; steps 4 to 6 are the evidence that it works, and they are what a reviewer will ask about.
What goes into a label spec for an LLM classifier?
A label spec gives every class four parts: a one-line essence, inclusions, exclusions, and boundary rules for the known collisions with its neighbours. It adds an exit class for inputs that fit nothing and a flag for inputs that fit two classes. The exclusions and boundary rules carry most of the value.
The book’s anatomy is exact: “a one-line essence; inclusions that pull the common cases in; exclusions that push the neighbors out; and boundary rules that arbitrate the known collisions.” Its reason for weighting the edges is shorter: “a class is defined at its edges; the centers were never in doubt.”
Take a ticket that reads “I was charged twice after the app crashed during checkout.” Billing or bug? Both readings are defensible, so something has to decide. In a trained model the decision comes from how annotators happened to label similar tickets years ago. In a prompted one it comes from a sentence you write, such as “classify by the remedy the customer is asking for: money back is billing, make-it-work is bug.” That rule is a business decision, and putting it where people can read and change it is much of what this route buys.
Give the menu an exit. Without an other class, the book warns, “a model forced to choose among categories that do not fit will manufacture a confident wrong label.” The exit is for inputs that belong to no class. Inputs that belong to two get the separate review flag described below, so that other never becomes the place hard cases go to hide.
One vendor’s guide to ticket routing makes the same point from the other side: the model’s ability to route tickets “is directly proportional to how well-defined your system’s categories are” (Anthropic documentation, read October 2026). Before you build an LLM classifier’s prompt, copy the template below into your repository and fill it in.
LABEL SPEC: [classifier name] v[n], [owner], [date]
Task: [one sentence: what is classified, and which decision the label feeds]
Input: [what one item contains; what the call must not contain, e.g. earlier verdicts]
CLASSES (closed list, plus an exit)
[class-a]: [one-line essence]
Includes: [the common cases]
Excludes: [neighbouring cases, and the class each one goes to]
Boundary vs [class-b]: [the rule that settles a known collision]
[class-b]: ...
other: [inputs that fit no class; not for inputs that fit two]
RUBRIC QUESTIONS (answered before the label, each from the item's own text)
1. [the deciding condition of class-a, asked as a question]
2. [the deciding condition of class-b, asked as a question]
3. [the tie-break, asked as a question]
Flag: needs_review = true when the answers point to different classes
or the input cannot be read.
OUTPUT CONTRACT (field order matters)
answers: one short sentence per rubric question
label: one of [class-a | class-b | ... | other] (closed enumeration)
needs_review: true | false
rationale: one or two lines citing the answers
IN-PROMPT EXAMPLES (each with the reason it is there)
[prototype | edge | pair] "[text]" -> [label], because [deciding feature]
MEASUREMENT
Development set: [n] items, labeled by [who], disjoint from the examples
Sealed test set: [n] items, labeled by [who] against this spec, opened on [dates]
Bar, written before opening: [kappa >= x; flagged <= y%; recall on the costly class >= z]
CHANGELOG
v[n] [date]: [change] | motivated by: [dev item ids] | dev before -> after: [accuracy, kappa, flagged]
The template holds for cases beyond tickets. For expense claims (in-policy, out-of-policy, missing-receipt, other), the boundary rule might read “a claim over the limit with a receipt is out-of-policy, not missing-receipt,” and the rubric questions become “Is there a receipt?” and “Is the amount under the category limit?” For a severity scale, each level is written as a class, which the section on scores comes back to.
How many examples belong in the prompt?
A handful, each chosen to teach one thing: a prototype or two per class, an edge case where a boundary is jagged, and a contrastive pair for each boundary the model actually confuses. The book’s image is a textbook rather than a database, and its verdict on the reflex to add more is blunt: “Examples are a stencil.”
Research on in-context learning supports keeping the catalog short. Min et al. (2022) found that “randomly replacing labels in the demonstrations barely hurts performance on a range of classification and multi-choce tasks” (the typo is theirs), across 12 models. What drove performance was the demonstrations showing “(1) the label space, (2) the distribution of the input text, and (3) the overall format of the sequence.”
Examples also carry hazards. Zhao et al. (2021) showed that “the choice of prompt format, training examples, and even the order of the training examples can cause accuracy to vary from near chance to near state-of-the-art,” and traced it to a bias toward answers that are “placed near the end of the prompt or are common in the pre-training data.” An example list with six billing tickets and one bug ticket is a thumb on the scale.
The contrastive pair is the role newcomers skip and the one that teaches most. Two tickets, worded almost the same, land on opposite sides of one boundary: “I clicked Upgrade and my card was charged twice” is billing; “I clicked Upgrade and the screen froze. No charge yet.” is bug. Everything is held constant except the feature that decides.
Which boundaries deserve a pair is a measurement question, not a guess. The development set tells you which two classes keep swapping, and that swap earns the next pair. The worked example below shows the loop.
Should the classifier reason before it picks a label?
Yes. Put two or three rubric questions in the output schema ahead of the label, answered from the input’s text, then the label, then a one-line rationale. The book gives the mechanism in one sentence: “text written before the label can inform the label, while text written after it can only defend it.”
The questions are the class definitions turned around. If billing hinges on whether money moved wrongly, ask that outright. For the ticket triage spec the three are: Does the ticket report money moving wrongly? Is the product behaving contrary to its design? What remedy is the customer asking for?
With those questions in place, a wrong label then points at a sentence of the guide: a question answered against the text is an evidence failure, correct answers with a wrong label are a definitions failure, and an answer no rule anticipated is a gap in the rubric.
Field order is easy to get wrong and expensive when you do. In a study of format restrictions, Tam et al. (2024) found that under one model’s JSON mode, “100% of” the responses “placed the”answer” key before the “reason” key, resulting in zero-shot direct answering instead of zero-shot chain-of-thought reasoning.” The reasoning field was there; it simply came too late to count.
The same paper is cited as evidence that format constraints hurt quality, and that claim is disputed. Its own results split by task: “stringent formats may hinder reasoning-intensive tasks but enhance accuracy in classification tasks requiring structured outputs.” A rebuttal from a company that builds structured generation (Kurt, 2024) argued that the gaps came from the prompts, not the constraint. For a classifier, both readings point the same way: make the label field a closed enumeration, constrained by the interface where it can be (the structured output for tool calls post covers what each grade guarantees), and put the reasoning fields first.
Add one more field, a review flag the model sets when its answers point at different classes. That turns the genuinely two-headed inputs into a queue for a person instead of confident wrong labels. The idea has a name in machine learning, selective classification or the reject option, which trades coverage for lower error (Geifman and El-Yaniv, 2017).
Don’t read the rationale as a confession. Turpin et al. (2023) found that chain-of-thought explanations “can systematically misrepresent the true reason for a model’s prediction.” The book’s own wording is to read the worksheet “as evidence, useful for debugging and audit, never as a faithful printout of what happened inside the weights.” The rationale tells you where to look; only the test set tells you whether the classifier works.
How should you split your labels into dev and test sets?
Keep three disjoint sets: the few examples inside the prompt, a development set you re-run after every edit and read freely, and a sealed test set you label once and open rarely. With no training set to feed, something close to an even split between development and test becomes defensible.
This is the three-set discipline, and disjointness is the whole point of it. “An example that sits in the guide and also in a measurement set is a leaked exam answer,” the book says. The tempting shortcut is to fix a failure by pasting the failed ticket into the prompt, re-run, and watch it pass, which measures the examples, not the classifier. One open-source quick-start for prompted classifiers demonstrates its correction loop in exactly this way: it adds the misclassified inputs as examples, then predicts those same inputs again.
The book’s illustrative arithmetic: “from two hundred labeled tickets, a hundred to iterate against and a hundred sealed.” Classical practice spends most labels on training; here every label you can afford buys measurement. Whether a hundred is enough for your question is a statistics problem, worked through in how many eval examples you need.
Label in an order that respects how criteria form. Shankar et al. (2024) named the effect criteria drift: “users need criteria to grade outputs, but grading outputs helps users define criteria.” So label the development set first and let the spec change while you do. Label the test set afterwards, against the spec as it then stands, before the classifier has seen a single test item.
Overfitting sounds like a disease of training, and a classifier with no training might seem immune. “It is exactly as immune as you are,” the book replies. Every clause bent around one odd development ticket fits the guide a little closer to those hundred tickets. The symptom is a development number that climbs while the sealed one stays put.
Worked example: four versions of a ticket classifier
The example is a ticket classifier with four classes (billing, bug, feature-request, other) and a review flag. It has 200 hand labels, 100 for development and 100 sealed, each with the same mix: 30 billing, 40 bug, 18 feature-request, 12 other. All numbers are illustrative.
| Version | Dev correct | Dev accuracy | Kappa | Flagged | Most confused pair |
|---|---|---|---|---|---|
| v1: definitions, bare label | 72 of 100 | 0.72 | 0.60 | 0 | billing and bug, 13 swaps |
| v2: + remedy rule, + one billing/bug pair | 82 of 100 | 0.82 | 0.74 | 0 | billing and bug, 5; bug and feature-request, 5 |
| v3: + rubric questions first, + review flag | 85 of 93 | 0.91 | 0.88 | 7 of 100 | bug and feature-request, 3 |
| v3 on the sealed test, opened once | 81 of 94 | 0.86 | 0.80 | 6 of 100 | billing and bug, 5 |
Accuracy and kappa are computed on the items the classifier labeled; flagged items went to a person. Read the table from the top. Version 1’s confusions put billing and bug at the head of the list with 13 swaps, so version 2 adds the remedy rule and one contrastive pair, and that pair drops to 5. Version 3 adds the rubric questions and the flag, and the flag takes 7 tickets off the board.
Then the envelope opens. Development accuracy was 0.914 and the sealed test gives 0.862, a gap of 5.2 points, and the billing and bug confusion is back at the top with 5 swaps. That gap is the overfitting the previous section described, at a size you should expect, and 0.86 is the number that gets reported.
The team wrote its bar down before opening the envelope: kappa of at least 0.75 and no more than 10% flagged. The sealed kappa of 0.80 and flag rate of 6% clear it.
The approximate 95% interval around kappa runs from 0.70 to 0.90, so a true value below the bar can’t be ruled out on 94 items. For routing tickets to queues, where a wrong label costs a detour, that is enough to ship. For a label that triggers a refund, it isn’t.
Below, the calculator opens on the sealed-test matrix. Change any cell to see how kappa, the lazy baseline and the ranked confusions move, or switch to pasted labels and enter your own.
With JavaScript on, the LLM Judge Agreement Calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the LLM Judge Agreement Calculator on its own page to share a result by link.
Which numbers should you report for an LLM classifier?
Report three: Cohen’s kappa on the sealed test with its interval, the share of items flagged for a person, and the confused pairs ranked by count. Accuracy can sit beside them, but alone it flatters any classifier on a skewed class mix.
Chapter 26 gives the problem in one sentence: “if seven tickets in ten are bugs, a classifier that answers bug unconditionally scores seventy percent while carrying no information.” In the worked example, answering bug every time would score 38 of 94, an accuracy of 0.40 with a kappa of zero. Cohen’s kappa corrects agreement for what chance alone would produce (Cohen, 1960); the arithmetic and its interval are worked step by step in is LLM as a judge reliable, which applies the same instrument to a model grading outputs.
The flag rate belongs next to accuracy because the two trade against each other. The book’s warning: “an instrument can be made to look arbitrarily accurate by teaching it to flag everything hard.” Of the 100 sealed tickets, 81 were labeled correctly, 13 wrongly and 6 went to a person. All three counts go in the report.
The confused pairs are the to-do list. Each recurring swap is the next boundary rule or contrastive pair in the guide, which is how the examples and the measurement form one loop. For a judge-shaped version of this classifier, the same reading applies, and the companion post on writing an LLM as a judge rubric covers the rubric side.
When do you stop iterating on the prompt?
Stop when the development set has stopped teaching you anything: the remaining confusions are items your own labelers dispute, or two revisions in a row fix one pair only by breaking another. Then open the envelope, compare it with the bar you wrote down, and decide.
The first condition needs a number you may not have yet. Have two people label the same fifty development items independently and compute kappa between them. If they agree at 0.6, no classifier can be shown to agree with either of them much above that, and the work left is on the spec, not the prompt. That ceiling is this post’s suggestion, using the same statistic, not a rule from the book.
The second condition is whack-a-mole, which the book names in its evaluation chapter. The book’s reason is that “no edit is local” in a guide read as one document: a clause tightened for billing can pull feature-requests across a boundary two classes away. Re-run the full development set after every edit, log the scores, and stop when the log shows you trading errors instead of removing them.
Two outcomes are possible when the envelope opens. If the sealed number clears the bar, ship and keep the changelog. If it misses, read the gap.
A large gap between development and test means the guide was fitted to the development set, and the fix is to replace case-by-case clauses with general rules. A development number that plateaued below the bar means the prompted route may be the wrong one, which is the next question.
Don’t edit the guide in response to individual test failures. Once the test set’s failures steer your edits it has become a second development set, and the book’s advice is to name that and retire it: fold its labels into development and label a fresh sealed set.
- No test item appears in the prompt’s examples or in the development set, including near-duplicates.
- The test labels were made against the current label spec, before the classifier ran on any test item.
- The bar (kappa, flag rate, recall on the class whose errors cost most) is written down.
- The changelog records every revision with the failure behind it and the development scores before and after.
- The development set has stopped teaching: remaining confusions are ones your labelers also dispute, or the last two revisions traded errors.
- Each item will be classified in a fresh call, with no earlier verdicts in the input.
- You have decided what a miss means: general rules for a large gap, a different route for a plateau.
The fresh-call item matters more in production than in testing. The book’s rule is to “classify every item on a fresh desk”: a run of earlier bug verdicts in the input makes the next bug a little easier to emit, so item four thousand is no longer judged by the same instrument as item four.
When should you fine-tune instead of prompting?
Fine-tune, or distill, when the categories have stopped moving, volume is high, and per-item cost or latency matters more than a guide people can read. Prompt when labels are scarce, the categories are young, volume is modest, or someone must be able to read why an item landed where it did.
The published comparisons lean toward fine-tuning once labeled data exists in quantity. Bucher and Martini (2024) found that “smaller, fine-tuned LLMs (still) consistently and significantly outperform larger, zero-shot prompted models in text classification,” with the fine-tuned model’s performance beginning “to saturate after around 200 labels” (in their ablation, against a zero-shot baseline that was not one of the generative models). Edwards and Camacho-Collados (2024), across 16 datasets, found zero- and few-shot results “considerably lower in comparison to smaller models fine-tuned on the entire training set.”
Two caveats keep those findings in proportion. Both compared zero-shot or one-shot prompts, not a guide with boundary rules and contrastive pairs, and both measured specific models at a specific date. Treat them as evidence that a trained model is a serious competitor once you have labels, not as a fixed accuracy gap.
| Condition | Prompted classifier | Fine-tuned or distilled model |
|---|---|---|
| Labels available | Dozens to a few hundred | Hundreds to thousands, or model-made labels with audits |
| Categories | New or still changing | Stable long enough that relabeling is rare |
| Volume and unit cost | Modest; per-call cost acceptable | High; per-item cost or latency is the constraint |
| Legibility | Someone must read why | A score on a test set is enough |
The two roads meet. A prompted classifier that works is a labeling machine for the training set you never had.
Wang et al. (2021) found it cost “50% to 96% less” to reach the same downstream performance with model-made labels than with human ones, and Gilardi et al. (2023) found a model’s zero-shot accuracy exceeded crowd workers’ on four of five annotation tasks. Both are dated measurements; the pattern is what transfers.
The book’s framing: “The guide remains the source; the small model is a compiled build.”
Before any of that, try the cheaper middle step: a fast screen that labels the obvious bulk and passes anything uncertain or flagged to the full classifier. The review flag you built is the escalation trigger. If the classifier’s cost ends up inside a per-task price, the cost-curve reasoning in AI agent pricing models applies to it too. The full decision table lives in the post on LLM classifier vs fine tuned classifier.
What changes when the classifier outputs a score?
Very little, if you treat the score as a set of ordered classes. Use a small ordinal scale, say one to four, and write each level the way you wrote each class, with boundary rules between adjacent levels. In the book’s words, “a score you can trust is a classification wearing numbers.”
A request for severity on a scale of “Zero to one, two decimals” produces numbers that look precise and mean nothing, because no sentence says what 0.7 is. The book’s test for whether two levels are real is whether you can write the sentence that separates them; if you can’t, merge them. Contrastive pairs go between neighbours, since on an ordered scale the confusable classes are always adjacent. Writing those level descriptors well is a rubric problem, and the rubric post linked above takes it up.
One measurement note: plain kappa counts a 2-versus-3 disagreement the same as a 1-versus-4. Weighted versions of kappa exist for ordered labels; this post’s worked example is nominal and doesn’t use them.
Where does this method break down?
It breaks in four places: very many classes, labels people can’t agree on, inputs that change under the classifier, and stakes the sealed number can’t cover.
Many classes. A guide can carry boundary rules for a handful of classes, not hundreds. One forum question asks how to classify into 700 categories by prompt (AI Stack Exchange, 2023). My suggestion, not the book’s, is a hierarchy: classify into a few parent groups, then within the parent, each level with a short guide and its own sealed test.
Disputed labels. If your labelers agree at a kappa of 0.6, the ceiling from the stopping section applies, and the problem is the spec.
Drift. New products and new complaints arrive. Sample production items into a fresh labeled set on a schedule and compare it with the sealed number; the observability side of that work belongs with the post on why an AI agent works in demo and fails in production.
Stakes. Prompts are sensitive to changes that preserve their meaning: Sclar et al. (2023) measured differences “of up to 76 accuracy points” from formatting alone in one open model, few-shot. Re-run the development set after any change to the prompt, the model or the interface, and gate actions by what a wrong label costs, not by how good the classifier looks.
The one thing to keep
When you build an LLM classifier, the prompt is the part you can write in an afternoon. The part a reviewer will trust is the sealed test, labeled against a written spec, disjoint from everything you tuned on, opened rarely, and reported with kappa and the flag rate beside it. Build that first and the prompt has something to answer to.
Chapter 26, “Agents as Classifiers and Scorers”, builds the full instrument, from definitions to packaging and the screen-then-detail cascade (in the full book). The agents in practice guide places this post among its neighbours, the eval sample size calculator sizes your test set, or you can see the formats.
Questions readers ask
- Can you build an LLM classifier with no labeled data at all?
- You can build one, but you cannot know whether it works. The prompt needs no training set; the measurement does need labels. A few hundred hand labels, split between a development set and a sealed test set, are enough to iterate honestly and to report one number you can defend.
- How many examples should go in the classifier's prompt?
- A handful, each there for a stated reason: a prototype or two per class, an edge case where a boundary is jagged, and a contrastive pair for each boundary the development set shows the model confusing. Research on in-context learning finds that demonstrations teach mostly the label space, the input distribution and the format, so dozens more add cost without adding much.
- Should the classifier explain its answer?
- Yes, but before the label, not after. Ask a few short rubric questions first, then the label, then a one-line rationale. Treat the written reasoning as a debugging aid; studies of chain-of-thought find explanations can misstate why a model answered as it did, so the sealed test, not the rationale, says whether the classifier works.
- Is accuracy the right metric for an LLM classifier?
- Not on its own. On a skewed class mix, accuracy rewards a classifier that always answers the majority class. Report Cohen's kappa, which corrects for chance, read the confused pairs, and show the share of items the classifier flagged for a person, because flagging every hard case also raises accuracy.
- When should I fine-tune instead of prompting?
- When the categories have been stable long enough that relabeling is rare, volume is high, and per-item cost or latency matters more than a guide people can read. A prompted classifier that already works can label the data for the fine-tuned one, with people auditing a sample, a route called distillation.
Sources
- Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, et al. (2022). Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
- Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein (2021). Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Melanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane Suhr (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design
- Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin (2024). Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
- Will Kurt (2024). Say What You Mean: A Response to 'Let Me Speak Freely' (a structured-generation company's rebuttal)
- Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Shreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
- Yonatan Geifman, Ran El-Yaniv (2017). Selective Classification for Deep Neural Networks
- Martin Juan José Bucher, Marco Martini (2024). Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
- Aleksandra Edwards, Jose Camacho-Collados (2024). Language Models for Text Classification: Is In-Context Learning Enough?
- Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu (2021). Want To Reduce Labeling Cost? GPT-3 Can Help
- Fabrizio Gilardi, Meysam Alizadeh, Maël Kubli (2023). ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- Anthropic documentation (2026). Ticket routing, a use-case guide in one vendor's documentation (read 7 October 2026)
- Lamini (2023). LLM Classifier README, an open-source prompted-classifier library (read 7 October 2026)
- Jacob Cohen (1960). A Coefficient of Agreement for Nominal Scales (Educational and Psychological Measurement 20(1):37–46)