Home / Blog / Agents at work / LLM Classifier vs Fine Tuned Classifier: A Deci…

Agents at work

LLM Classifier vs Fine Tuned Classifier: A Decision Table

LLM classifier vs fine tuned classifier: a decision table with three verdicts, the studies on both sides, and the test each must pass. Copy the recipe.

By Enrique Gutiérrez · Published · 23 min read

The choice of LLM classifier vs fine tuned classifier turns on four things: how stable the categories are, how many items arrive, how fast each answer must come, and who reads the reasons. Prompt while categories move and volume is modest, when someone must read the reasons, or when labels are few. Train a model to serve alone when categories are stable and the latency budget is tight, or when spare human labels and a pattern-driven task favor it. Put a cascade in front of the prompt when volume is high and an item can wait.

A classifier assigns each input one label from a fixed list, and three candidates can do that job. A prompted LLM reads a written guide on every call. In a fine-tuned model, training on your labels has changed the weights; it may be a small encoder or a fine-tuned LLM. The cascade puts a cheap screen in front of the full prompted classifier and escalates what the screen cannot settle.

Below: a table that returns one of those three verdicts, the evidence behind each row, a recipe for the cascade, and the test every candidate has to pass. I assume you can already build an LLM classifier from a written label spec. The worked numbers are illustrative.

What does the book say about an LLM classifier vs fine-tuned classifier?

One paragraph near the end of Chapter 26, Agents as Classifiers and Scorers (in the full book) holds the book’s whole answer. It names four conditions that favor the prompted classifier and three that favor the trained one, then shows how the two roads join. It prints no decision table and no thresholds.

Its first sentence carries the prompted side:

“The LLM classifier earns its keep when labels are scarce; when the categories are young and still moving, because a definitions edit costs minutes where a relabeling campaign costs a quarter; when volumes are modest; and when legibility carries weight, because an auditor can read a guide and nobody can read weights.”

Then comes the other side: “The trained model is the better bet in the opposite corner: categories stable for years, volumes vast, and per-item cost and latency floors that no per-call model price will meet.” The paragraph then joins them. Once the prompted classifier runs well, “its verdicts (human-audited at whatever rate the stakes demand) are a labeled dataset, the very thing you never had”, and a small model can be trained to imitate them.

That small model is, in the book’s phrase, “a compiled build” of the guide. The upkeep rule follows: “When the categories move, you edit the source and recompile.” The same chapter describes a cascade, and Chapter 8, Retrieval and Knowledge (in the full book) adds a rule for any training run: “exhaust the cheaper road first”.

Everything past that paragraph is this post’s own. The decision table, its rows, its thresholds and its tie-break order are my construction, and so is the rule that all candidates sit one sealed test. A threshold taken from a study names the study; a round number is marked illustrative.

Which of the three should you choose?

Read the table from the top and stop at the first row that is true of your task; that row’s last column is the verdict. The order is the tie-break, so a task that fits two rows gets the earlier row’s answer.

# If this is true of your task Prompted LLM Fine-tuned model Cascade (hybrid) Where the row comes from Verdict
1 You may not train on the data, or nobody on the team can run a training pipeline Needs a guide and a test set; no training data Ruled out until the constraint lifts Only with an untrained screen: rules or a smaller prompted model This post’s row; the book says the classical route needs “a training pipeline, someone who knows how to operate one” prompted
2 Every verdict needs a written reason that someone reads: an audit, an appeal, a score on an ordinal rubric Writes rubric answers and a rationale beside each label A small encoder returns a label and a score, with nothing to read A screen may decide only items that need no reason, and none qualify The book: “an auditor can read a guide and nobody can read weights” prompted
3 Categories are moving, and volume is modest Edit the guide, re-run the development set Relabel and retrain at each change A second model to keep in step, for a saving modest volume cannot repay The book: “a definitions edit costs minutes where a relabeling campaign costs a quarter” prompted
4 Categories are moving, and volume is high enough that the full prompt’s per-item cost is the standing complaint Correct, and paid for on every item The training set goes stale at each rename A cheap screen takes the obvious bulk and is rebuilt after each edit to the guide The book’s screen-then-detail cascade; Nie et al. (2024) hybrid
5 Categories are stable, volume is high, and every item has a latency budget that a large-model call cannot meet Latency and cost on every item One-off training cost, then fast and cheap per item The escalated share waits on two models, so escalation must leave the response path The book’s “opposite corner”; one 2025 study measured 0.02 s per item for a fine-tuned encoder and 0.6 to 1.4 s for most LLMs it tested fine-tuned
6 Categories are stable, volume is high, and the latency budget leaves room for a large-model call on the escalated share, or the work runs in batch Works; cost grows with every item Works alone if a training set exists; uncertain items still get a guess The small model answers the confident bulk; the rest goes to the full guide The book’s cascade joined to its distillation bridge; the combination is this post’s hybrid
7 Categories are stable, volume is modest, you hold several hundred to a thousand human-made training labels beyond the test set, and the task turns on surface patterns Trailed in most of the cited comparisons at this label count; tied or led in a few Ahead in most of those comparisons, most clearly near 1,000 labels No volume to justify two models Bucher and Martini (2024) and Wang, Qu and Ye (2024), on their tasks; Zhang, Huang et al. (2025) for pattern-driven tasks. An accuracy call that departs from the book’s “when volumes are modest”; run the break-even below before training fine-tuned
8 None of the above: few labels, modest volume, or a task that leans on world knowledge The default; build it first Nothing to train on yet Nothing to screen yet The book’s “when labels are scarce”; Zhang, Deng et al. (2023) for few-shot settings prompted

Two words in the table need fixing down, and both windows are illustrative. Categories are moving if one was renamed, split or merged in the last quarter, or you expect that soon. They are stable after a year without such a change; anything in between counts as moving.

“Fine-tuned” in the last column means trained weights answer every item alone, wherever the training labels came from. “Hybrid” means two stages in sequence. A distilled model serving by itself therefore counts as fine-tuned, and the same model placed in front of the full guide is the hybrid’s screen.

Why are the rows in that order?

The order runs from what you cannot change to what you can measure. Constraints come first, then the need for written reasons, then label stability, then volume, then latency, then labels in hand and task type, and last the default.

A contract that forbids training (row 1) is settled outside engineering. A required reason (row 2) outranks cost, because a model that cannot explain its verdict fails at any price. Moving categories (rows 3 and 4) outrank volume, since a training set that goes stale each quarter never pays back.

Volume and latency (rows 4 to 6) then decide between the cheaper serving shapes, and published accuracy evidence breaks the tie only at the end (rows 7 and 8). “High volume” has no universal number. The book’s illustration sets fifty tickets a day against fifty thousand, and at the low end its advice is to “run the full instrument on everything and think no more about it.”

What have the studies found, in both directions?

The published comparisons split by condition. With a few hundred to a thousand or more task labels to train on, fine-tuned models matched or beat zero- and few-shot prompts in several studies. With few labels, or on tasks that need world knowledge, prompted models led or tied in others. Each result below belongs to its datasets and its date.

Where did a fine-tuned model come out ahead?

Three studies read for this post found a fine-tuned model ahead of zero- or few-shot prompts on most of their tasks, after training on anything from a few hundred task labels to a full training set.

Bucher and Martini (2024) (arXiv 2406.08660v2) classified sentiment, approval, emotions and party positions in news, tweets and speeches. Three generative models prompted zero-shot, plus one zero-shot inference model, faced several fine-tuned encoders. Their finding: “fine-tuning with application-specific training data achieves superior performance in all cases”.

Their ablation on training-set size reads: “the sweet-spot tends to lie between 200 and 500 training observations”. That range belongs to their tasks, and they grant that “zero-shot results are decent for common tasks such as sentiment analysis”.

Edwards and Camacho-Collados (2024) (arXiv 2403.17661v2) covered 16 datasets, binary, multiclass and multilabel. They found that “fine-tuning smaller and more efficient language models can still outperform few-shot approaches of larger language models”. Their fine-tuned models used the entire training set of each dataset.

Wang, Qu and Ye (2024) (a preprint, arXiv 2411.05050v1) ran five political-science tasks with 2 to 22 classes. An encoder fine-tuned on 200, 500 and 1,000 samples faced zero- and few-shot prompting. Prompting gave “reasonable performance”, yet the prompted models “generally fall short” of the fine-tuned encoder or “at best, match” it, “particularly as the training set reaches a substantial size (e.g., 1,000 samples)”.

The five tasks did not agree at the low end. On the 22-class task, zero-shot prompting beat the encoder trained on 200 samples, and few-shot prompting was “equal to or slightly better than” it at 500 and 1,000. That 1,000 belongs to their five tasks.

Where did a prompted model lead or tie?

Four studies point the other way or complicate the picture: in few-shot settings, on tasks that need world knowledge, in one multilingual comparison, and where the fine-tuned winner was itself a large model.

Zhang, Deng, Liu, Pan and Bing (2023) (arXiv 2305.15005v1) studied sentiment analysis across 13 tasks and 26 datasets, comparing LLMs with small language models (SLMs). In their words, “LLMs significantly outperform SLMs in few-shot learning settings, suggesting their potential when annotation resources are limited.” The same abstract reports that LLMs “lag behind in more complex tasks requiring deeper understanding or structured sentiment information”.

Zhang, Huang, Liu, Gao and Hu (2025) (arXiv 2505.18215v1) chose six datasets they describe as high-difficulty. Overall, “BERT-like models often outperform LLMs”. By task type the result divides: “BERT-like models excel in pattern-driven tasks, while LLMs dominate those requiring deep semantics or world knowledge.”

Kuzman Pungeršek and colleagues (2025) (arXiv 2511.07989v2) compared the two on sentiment, topic and genre in several South Slavic languages. They found that “LLMs demonstrate strong zero-shot performance, often matching or surpassing fine-tuned BERT-like models”. They still call the fine-tuned models “a more practical choice for large-scale automatic text annotation”.

Their reasons were less predictable outputs, slower inference and higher computational cost, and the paper gives seconds. The fine-tuned model ran at “0.02 seconds per instance”, while “most LLMs have inference times between 0.6 and 1.4 seconds per instance”. Those figures come from their models and hardware, at their date.

Zhao, Chen, Zhang and Yang (2024) (arXiv 2412.08587v2) add a third candidate. On two datasets, news topics and intents, a fully fine-tuned decoder with 70 billion parameters outperformed a fine-tuned large encoder. “Fine-tuned” can mean a large model too, with a large model’s serving cost.

What did none of them test?

None of the prompted arms described above used a written guide with boundary rules and contrastive pairs. They are zero-shot or few-shot prompts. A guide-driven classifier may therefore do better than the prompts these papers measured, by an amount I have found no study of.

I read the LLM classifier vs fine-tuned classifier literature as a map of conditions. It suggests which candidate to build first, and which one wins on your tickets is a question for your own test set.

How many labels does each route need?

Both routes need the same sealed test set, labeled by people. A prompted classifier needs a development set beside it. A fine-tuned one needs a training set on top, which the studies above place between a few hundred labels and a full training set, depending on the task.

The post’s argument hinges here. The book states it for the prompted route: “The need for labeled truth remains, because measurement runs on it; what changes is that every label you can afford now buys measurement.” The prompt removed the training set and left the test set exactly where it was.

Some teams call that test set a golden dataset for LLM evaluation. The book’s terms are the held-out test set and the sealed envelope, one part of its three-set discipline.

So the labels for measurement are a cost both routes share, and they drop out of the comparison. What fine-tuning costs extra is a second labeled set for training, a pipeline, someone to operate it, and the upkeep of all three whenever the categories or the inputs shift.

On Hacker News, coder68 wrote in an August 2025 comment: “To some degree manual labeling has to be done anyway, just to validate that any approach works at all”.

What does that do to the cost comparison?

The shared test set turns the cost question into a break-even on volume. The prompted route pays on every item. The fine-tuned route pays once for training labels, training and a pipeline, then a small amount per item, and part of the one-off again at each retrain.

In arbitrary, illustrative units: a one-off cost of 3,000, a prompted cost of 0.002 per item and a small-model cost of 0.00002 per item give a break-even near 1.5 million items. At 5 million items a day that is reached in under a day. At 300 items a day it takes about 5,050 days, close to 14 years.

Wang, Qu and Ye describe the same structure: “Since fine-tuning is performed only once, the associated costs can be considered sunk”. Each prompted call typically costs more than an encoder’s inference, they add, “at least for now.” Whether a given saving repays the engineering time is a question of AI agent ROI, the return weighed against the risk of a wrong label. Prices date within months, so put your own measured costs into the division.

Can you train the small model on the LLM’s labels?

Yes. Training a small model to imitate a working prompted classifier is called distillation, and it removes most of the labeling cost of the fine-tuned route. The student inherits the teacher’s mistakes along with its skill, so the audit of the teacher’s labels is the price.

The book’s free glossary defines fine-tuning as “Continuing a model’s training on your own examples, so that the weights themselves change.” Distillation is that, with examples a model labeled.

Pangakis and Wolken (2024) (arXiv 2406.17633v1) measured the result on 14 classification tasks from recent social-science articles. Classifiers trained on an LLM’s labels landed a median 0.039 F1 below classifiers trained on human labels. Against the LLM itself, prompted few-shot, the median difference was 0.006.

The second finding matters more for design. Teacher and students alike, the paper reports, “perform significantly better than all other models on recall, but noticeably worse on precision”. The student copied the teacher’s tilt.

So audit the teacher’s labels before training, and grade the student against human labels afterwards. Agreement with its teacher only shows that the copy is faithful.

How does the cascade combine the two?

A cascade runs a cheap screen on every item and sends only the uncertain ones to the full prompted classifier. The screen can be rules, a smaller prompted model, or the distilled student. The guide stays the single source, and both stages are builds of it.

The book introduces the cascade with one sentence: “Volume changes the economics before it changes anything else.” Its rule for the screen is as short: “its one design requirement is that it knows what it may not decide”. The book’s cascade pairs a slim prompted screen with the full prompted pass; using the distilled student as the screen is this post’s combination.

The screen-then-detail cascade, with band widths standing in for volume.
Figure 26.5 The screen-then-detail cascade, with band widths standing in for volume. A cheap screen clears the unmistakable bulk on a wide path, and only the uncertain minority (in accent) is escalated to the full instrument. Of those, only a thin flagged trickle reaches a person. Most items never touch the expensive call—which is what makes the cascade pay, and the abstain flag is exactly its escalation trigger. Reuse this diagram

If the screen costs a tenth of the full call and 15% of items escalate, the cascade costs 0.10 + 0.15 = 0.25 of the full classifier per item (illustrative). If half the items escalate, the same sum gives 0.60, and you now maintain two models for a modest saving. Chapter 19, Cost, Latency, and Performance (in the full book) names that state: “a cascade where half the traffic escalates is a cheap tier that is failing its job.”

Published savings are upper bounds in their own settings. Nie and colleagues (2024) (arXiv 2402.04513v3) built a cascade that ends with an LLM and whose smaller models learn from its outputs over time. Across four benchmarks they report accuracy on par with the LLM “while cutting down inference costs by as much as 90%”.

Escalated items pay in latency. Chapter 10, Workflows and Composition Patterns (in the full book) puts it this way: “every request pays the cheap call and the escalated fraction pays twice, waiting on two models in sequence”. Hence row 5 of the table, which sends tight latency budgets to a model that serves alone.

HYBRID RECIPE: screen, escalate, audit, retrain
(pseudocode; the guide is the source, each model is a build of it)

SETUP, once
  guide        = the written label spec, under version control
  detail(item) = full prompted classifier: rubric answers, label, review flag
  screen(item) = cheap model: label plus confidence
                 (distilled small model, smaller prompted model, or rules)
  validation   = human-labeled items, disjoint from the sealed test set
  T            = lowest threshold at which the items the screen keeps
                 meet the bar on validation (the sealed test stays closed)
  note         : a confidence an LLM states in words tends to run high
                 (Xiong et al., 2024) and is no calibrated probability;
                 check it on validation before using it as a threshold

FOR EACH ITEM, in a fresh call
  s = screen(item)
  if s.confidence >= T and s.label is one the screen may decide:
      emit(s.label, source = "screen")
  else:
      d = detail(item)
      if d.needs_review: queue_for_person(item, d.worksheet)
      else:              emit(d.label, source = "detail")

ON A SCHEDULE (weekly is illustrative)
  audit   : sample emitted items from BOTH sources, confident ones included;
            a person labels them blind -> adjudicated labels
  watch   : escalation rate, screen error rate on audited items
            (escalation nearing half the traffic = the screen is failing)
  retrain : screen <- adjudicated labels (sealed-test items never enter)
  re-test : the whole cascade, end to end, after each retrain, on a
            sealed sample opened rarely; draw fresh production items
            once the old sample's results have steered a change

WHEN A CATEGORY CHANGES
  edit the guide -> relabel affected validation and test items
  -> rebuild the screen -> pick T again on validation -> re-test

Under a tight latency budget the same loop runs with one change. The small model answers every item alone, and an item below the threshold takes a safe default and joins the review queue. The audit, the retraining and the re-test stay as written.

Which confidence can you put a threshold on?

Neither route hands you one for free. A trained classifier outputs a probability that has measured well in-domain in one study and still needs checking on your data. A prompted LLM’s stated confidence tends to run high. Both become usable only after a check on held-out labels.

Calibration means stated probabilities match observed frequencies: of the items scored 0.9, about nine in ten come out right.

For fine-tuned encoders, Desai and Durrett (2020) (arXiv 2003.07892v3) found the models they tested “calibrated in-domain” when used out of the box. Out of domain they needed more help. Their tasks were natural language inference, paraphrase detection and commonsense reasoning, so the result covers inputs that resemble the training data and promises nothing about drifted ones.

For prompted models, Xiong and colleagues (2024) (arXiv 2306.13063v2) found that “LLMs, when verbalizing their confidence, tend to be overconfident”, on question-answering and reasoning tasks. No method they tried consistently beat the others.

No study read for this post measured calibration for a guide-driven prompted classifier on business categories. The book’s device avoids the probability altogether: a flag the model sets “when the rubric answers pull in different directions”. Judge a flag by its rate and by the error rate of the items it let through.

What test must every candidate pass?

Every candidate sits the same sealed test set, drawn from production and labeled by people against the current guide. The book states the principle for the prompted classifier, which is to be “trusted exactly as far as the sealed number says”.

Accuracy alone misleads on a skewed class mix, so the report is per class. Precision for a class is the share of items given that label that deserved it. Recall is the share of items deserving the label that received it. A distilled student’s inherited tilt shows up here.

Then comes the comparison. Suppose the prompted candidate gets 172 of 200 sealed items right and the fine-tuned one gets 181 of 200: 86.0% against 90.5%, a gap of 4.5 points. The 95% Wilson intervals run from 80.5% to 90.1% and from 85.6% to 93.8%, and they overlap. An unpaired two-proportion test gives a two-sided p of 0.162 (the calculator rounds it to 0.16).

The calculator below opens on exactly that comparison. It shows that a 4.5-point gap on 200 items each does not separate the two candidates. Detecting a gap of that size with 80% power takes about 803 items per candidate.

With JavaScript on, the Eval sample-size calculator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Eval sample-size calculator on its own page to share a result by link.

A rough 95% interval on the difference runs from about 1.8 points in the prompted candidate’s favor to 10.8 in the fine-tuned one’s (normal approximation, my arithmetic). Zero sits inside it, so the report reads “no difference shown yet”.

The calculator’s test is unpaired, and yours need not be. Both candidates labeled the same 200 items, so a paired test on the items where they disagree can settle the same gap with fewer items. The post on how many eval examples you need works through that design.

  • One sealed test set, sampled from production and labeled by people against the current guide, grades every candidate.
  • No sealed item appears in a prompt, a development set, a training set, or the pool a teacher model labels.
  • The bar is written before any candidate runs: per-class precision and recall floors, an escalation ceiling, and latency and cost ceilings.
  • Precision and recall are reported per class, with the item count behind each.
  • All candidates label the same items, and the difference is reported with its interval; an interval that includes zero is logged as “no difference shown yet”.
  • Slices are checked one by one: the rarest class, the longest inputs, each language or segment, the newest month.
  • Latency and per-item cost are measured on your own traffic, for the whole path, escalations included.
  • A distilled model is graded against human labels, and a cascade end to end, with its escalation rate and the error rate of items the screen decided alone.
  • Any confidence threshold was chosen on validation data and checked there for calibration.
  • A re-test date is set, with fresh production items and a named owner.

If the label feeds an agent that acts on it, the wider panel of AI agent evaluation metrics, task success and cost per successful task among them, grades what happens after the label.

Worked cases: two tasks through the table

Two tasks, taken row by row, show the table and the checklist at work. A third stops at row 2 and gets the prompted verdict: scoring 2,000 documents a month on a five-point ordinal scale, an LLM judge writing a reason for each score.

Case 1: 300 support tickets a day, eight categories that product keeps renaming

Row 3 gives the verdict: prompted. The team may train on its tickets and has an engineer who could, so row 1 is false. Nobody reads a reason per ticket, so row 2 is false. The categories changed last quarter and 300 a day is modest, which is row 3.

At this volume the illustrative one-off cost above takes years to repay. Ticket routing is also the queue-triage shape that recurs among AI agent use cases for startups, where a wrong label costs a detour.

One adjustment makes the checklist run. Eight classes on a 200-item sealed set (illustrative) leave the rarest class a handful of items, so its recall is nearly unmeasured. Sample extra items for rare classes, and relabel affected test items after each rename.

Two events flip the verdict. Volume may rise until per-item cost is the standing complaint, which is row 4 and a cascade. Or a year may pass without a rename while people correct several hundred audited verdicts into human-checked labels, which opens row 7. At 300 a day that would be a call on accuracy alone, since the one-off cost still takes years to repay.

Case 2: 5 million short messages a day, one stable policy label, a tight latency budget

Row 5 gives the verdict: fine-tuned. Assume the team may train on the messages, so row 1 is false. No reader needs a reason per message, and the policy has held for more than a year, so rows 2 to 4 are false. Volume is high and the latency budget excludes a large-model call on each message, so row 5 is true.

Row 5 leaves the source of training labels open. If a moderation history exists, use it. If none does, a prompted classifier labels a sample offline, people audit a share of those labels, and the small model is trained on the result. The sealed test set stays human-labeled either way.

Two checklist items bite here. A policy label is usually rare, so a random sample holds few positives; draw extra positives for the test set and weight back to the production rate when you report precision. A distilled model’s recall tilt then decides how many messages reach the review queue.

The verdict flips to hybrid if the latency budget loosens (row 6). If the policy starts to move while the budget stays tight, row 4 also says hybrid, and the escalation has to leave the response path: the small model still answers every message alone, and the full guide relabels a sample offline to rebuild it after each edit.

When should you revisit the choice?

Revisit the LLM classifier vs fine tuned classifier decision on the re-test date and on four events: a category change, a shift in volume, a failed re-test, and a new model on either side.

A rename or a split sends you back to rows 3 and 4. Volume that makes per-item cost the standing complaint moves a task from row 3 to row 4, or from row 8 to row 6. A year on stable categories opens row 6 at high volume, and row 7 once several hundred human-checked labels exist.

A re-test below the bar may point at the inputs, because production traffic drifts under every route. Any change of model, on either side, voids the last sealed number until the test is run again.

Where does this decision table fall short?

The table is untested as a table. It orders the book’s conditions and the cited findings into one sequence, and nobody has run a study on that sequence. It also has no row for class count, or for data that may not be sent to a hosted model (there a prompted verdict means a model you run yourself; the constraint removes the hosted service and leaves the prompted route), and the prompted arms in its sources were zero- and few-shot prompts.

Subjective labels strain every route. If two of your own labelers often disagree, no candidate can be shown to beat that ceiling. Distillation may be weakest there: Li, Zhu, Lu and Yin (2023), studying model-generated training data, found subjectivity “negatively associated with the performance of the model trained on synthetic data.” Fix the guide first; the judge agreement calculator measures the ceiling.

The one thing to keep

Whichever way you settle LLM classifier vs fine tuned classifier, the first step is shared: the sealed test set. Label it by hand, draw it from production, and run every candidate on it. After that the choice is a sequence. Prompt first, add a screen when volume asks for it, and train a small model once the categories hold still and the audited labels exist.

Chapter 26, “Agents as Classifiers and Scorers”, develops the full instrument and the cascade (in the full book). The agents in practice guide places this post among its neighbors, the eval sample size calculator sizes the comparison, or you can see the formats.

Questions readers ask

Is an LLM better than a fine-tuned BERT-style model for text classification?
It depends on the conditions, and the published studies say so. Fine-tuned encoders matched or beat zero- and few-shot prompts in several comparisons once they were trained on a few hundred to a thousand or more task labels. Prompted models led in few-shot settings and on tasks needing world knowledge. Run both candidates on one sealed test set from your own traffic before deciding.
How many labeled examples do I need to fine-tune a classifier?
No single number holds across tasks. One 2024 study of political and sentiment tasks placed the sweet spot between 200 and 500 training observations; another tested 200, 500 and 1,000 samples and found the encoder ahead on most of its tasks, most clearly at 1,000. Those figures belong to those tasks. Add the sealed test set on top, because training labels cannot grade the model.
Can I train a small classifier on labels from an LLM?
Yes, and the route has a name, distillation. In one 2024 study across 14 tasks, classifiers trained on model-made labels landed a median 0.039 F1 below classifiers trained on human labels, and they copied the teacher's higher recall and lower precision. Audit a sample of the teacher's labels first, and grade the student against human labels.
Is it cheaper to run a small model than to call an LLM for classification?
Per item, usually yes; in total, it depends on volume. A prompted classifier pays on every item. A fine-tuned one pays once for labels, training and a pipeline, then little per item, and pays again at each retrain. Divide the one-off cost by the per-item saving to get the break-even volume, using costs measured on your own traffic.
Do I still need labeled data if I use a prompted classifier?
Yes. The prompt replaces the training set, but measurement still runs on labels. A prompted classifier needs a development set to iterate against and a sealed test set to report from, both labeled by people. The same sealed set then grades any fine-tuned or distilled model you try later.

Sources

  1. Martin Juan José Bucher, Marco Martini (2024). Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification (arXiv 2406.08660, v2 read)
  2. Aleksandra Edwards, Jose Camacho-Collados (2024). Language Models for Text Classification: Is In-Context Learning Enough? (arXiv 2403.17661, v2 read)
  3. Wang, Qu and Ye (2024). Selecting Between BERT and GPT for Text Classification in Political Science Research (preprint, arXiv 2411.05050, v1 read)
  4. Zhang, Deng, Liu, Pan and Bing (2023). Sentiment Analysis in the Era of Large Language Models: A Reality Check (arXiv 2305.15005, v1 read)
  5. Zhang, Huang, Liu, Gao and Hu (2025). Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs? (arXiv 2505.18215, v1 read)
  6. Kuzman Pungeršek, Rupnik, Porupski, Dinić and Ljubešić (2025). State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting? (arXiv 2511.07989, v2 read)
  7. Zhao, Chen, Zhang and Yang (2024). Advancing Single and Multi-task Text Classification through Large Language Model Fine-tuning (arXiv 2412.08587, v2 read)
  8. Pangakis and Wolken (2024). Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels (arXiv 2406.17633, v1 read)
  9. Li, Zhu, Lu and Yin (2023). Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations (arXiv 2310.07849, v2 read)
  10. Nie, Ding, Hu, Jermaine and Chaudhuri (2024). Online Cascade Learning for Efficient Inference over Streams (arXiv 2402.04513, v3 read)
  11. Xiong et al. (2024). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs (arXiv 2306.13063, v2 read)
  12. Desai and Durrett (2020). Calibration of Pre-trained Transformers (arXiv 2003.07892, v3 read)
  13. coder68, Hacker News (2025). Hacker News comment in a thread on LLMs for text classification at volume (27–29 August 2025)