Home / Blog / Evaluating and observing agents / Is LLM as a Judge Reliable? A Calibration Metho…

Evaluating and observing agents

Is LLM as a Judge Reliable? A Calibration Method That Tells You

Is LLM as a judge reliable? Only as far as you have measured it. A calibration method with Cohen's kappa, a worked example and the four biases to test.

By Enrique Gutiérrez · Published · 15 min read

Is LLM as a judge reliable? An LLM judge is exactly as reliable as the agreement you have measured between it and your own human labels, on a sample drawn from your real traffic, and no further. Out of the box it is an uncalibrated instrument with documented biases; once calibrated, it can grade thousands of outputs about as well as a second human grader would.

That answer turns a yes-or-no question into a measurement you can run this week. This post gives the method: a small hand-labeled sample, an agreement statistic that cannot be gamed by a lazy judge (Cohen’s kappa), a worked example with real arithmetic, the four biases to test before you trust any verdict, and the rule for when you should not be using a judge at all. It follows the treatment in Chapter 16 of AI Agents, Engineered, whose verdict fits in one line: “A judge you have not calibrated is an opinion you have automated.”

Is LLM as a judge reliable enough to trust?

An LLM judge is reliable enough to trust when its chance-corrected agreement with your human labels reaches the level two careful humans reach with each other, measured on a few dozen real outputs or more. Until you have that number, the honest answer is “unknown,” however confident the judge’s rationales sound.

The reason the question gets asked so often is that the technique arrived with a strong headline. The study that legitimized it, Zheng and colleagues’ 2023 paper “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” reported that strong judges “can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans.” The book reads that result the right way: “a good judge approximates a human grader about as well as a second human does. It does no better, and the same study catalogued the systematic ways it does worse.”

So the useful version of “is LLM as a judge reliable” is narrower and answerable: reliable for which task, against whose labels, with what agreement, and checked how recently? The rest of this post answers each part.

What is LLM-as-a-judge?

LLM-as-a-judge is the practice of using a language model, armed with a written rubric, to score or compare outputs the way a human reviewer would. It exists because the outputs teams care most about (a research brief, a refactored module, a delicate customer reply) cannot be graded by a string match or a schema check, and humans cannot sit inside a test suite that runs on every commit.

A judge can work in two framings. Pointwise grading scores one output alone, as pass/fail or on a scale. Pairwise grading asks which of two candidates is better. Chapter 16 calls pairwise “often the steadier framing,” because absolute scores drift and bunch so badly that “a ten-point scale behaves, in practice, like a noisy three-point one.”

Neither framing is bias-free, which is why calibration applies to both. How to write the rubric itself, with ordinal levels that hold up, is the subject of the companion post on writing an LLM-as-a-judge rubric.

The same object shows up under other names. In the evaluator-optimizer pattern, a second model grades the first and sends it back to revise; the grader there is a judge, and everything below applies to it. The glossary entry for evaluator-optimizer gives the short definition.

What did the original study actually measure?

The original study measured agreement between a judge’s pairwise preferences and human experts’ preferences on open-ended questions, and the famous number depends on how ties are counted. With ties excluded, the strongest judge agreed with humans 85% of the time against 81% between humans; with ties counted, the figures were 66% and 63%.

That distinction rarely survives the trip into blog posts. The paper’s own sentence is precise: “The agreement under setup S2 (w/o tie) between [the strongest judge] and humans reaches 85%, which is even higher than the agreement among humans (81%).” In the setup that keeps ties (where random guessing would score 33%), pairwise judge-human agreement was 66% and human-human agreement was 63%. Both readings support the same conclusion, that a strong judge sits about where a second human sits. Neither supports the conclusion that a judge is right 85% of the time on your task.

Two further results from that paper matter for anyone deploying a judge. Position effects were large: with the default prompt, the strongest judge returned the same verdict after the two answers were swapped in only 65% of comparisons, and one weaker judge favored whichever answer came first in 75% of them. And judges struggled to grade math and reasoning answers, sometimes even when they could solve the problem on their own, because the candidate answers misled them. The authors’ fix for the second problem was to give the judge a reference answer.

Why is raw accuracy the wrong number for a judge?

Raw accuracy is the wrong number because it rewards a judge for matching the majority class without discriminating anything. When most outputs are good, a judge that always says “good” earns a high accuracy while carrying no information, and nothing in the accuracy figure warns you.

The book puts it in one sentence I would pin above every eval dashboard: “Raw accuracy is the number everyone quotes and the one that lies: if seventy percent of your outputs are good and the judge answers”good” every single time, it scores seventy percent while carrying zero information.” Chapter 26 meets the same trap in a classifier and calls it “Chapter 16’s lazy judge, returned in a new uniform.”

Cohen’s kappa is the fix. It is a chance-corrected agreement statistic: it measures how much better than lucky guessing the judge’s agreement with your labels actually is. The formula is short:

κ = (p_o − p_e) / (1 − p_e)

Here p_o is observed agreement (plain accuracy) and p_e is the agreement you would expect if the judge and the human labeled independently at their own base rates. A judge that always says “good” on a 70/30 sample has p_o = 0.70 and p_e = 0.70, so κ = 0. The laziness is priced at exactly zero, which is what you want.

The gap between the two numbers is not a curiosity. A 2026 study of 21 judges and roughly 541,000 judgments (Norman, Rivera and Hughes, 2026) found that the drop from exact-match agreement to Cohen’s kappa was “universal (33-41 pp on MT-Bench).” That is how much exact-match agreement flatters the same judge: thirty-odd percentage points of apparent reliability that vanish once chance agreement is priced in.

How do you calibrate an LLM judge against human labels?

You calibrate a judge the way you calibrate any instrument: hand-label a representative sample yourself, run the judge over the same sample, and compare. Chapter 16 asks for “a few dozen outputs at minimum, drawn from real traffic, including some truly bad ones,” then agreement reported as kappa, split into precision and recall where the verdict gates something.

In practice the procedure has six steps:

  1. Draw the sample from real traffic, not from examples you wrote while designing the rubric. Seed it deliberately with failures, or a 90%-good stream will leave you with three bad outputs to learn from. If you already collect traces, the post on LLM tracing with OpenTelemetry covers getting them into a form you can sample.
  2. Label it before you look at the judge’s verdicts. Use binary pass/fail where you can. Hamel Husain, who has set up evaluation for many teams, makes the same call: “Tracking a bunch of scores on a 1-5 scale is often a sign of a bad eval process.”
  3. Run the judge with the rubric you intend to ship, reasoning first and verdict last. The book’s reason: “a verdict stated first and justified after is a snap judgment wearing a lab coat.”
  4. Compute accuracy, kappa, and the confusion table, and read the off-diagonal cells as a to-do list.
  5. Split out precision and recall for the class you care about. A judge that misses bad outputs and a judge that fails good ones look identical in accuracy and cost you very different things.
  6. Check the level, not just the agreement. High agreement on ranking can coexist with a judge that is systematically harsher or kinder than you; compare the judge’s pass rate with yours.

Keep the calibration labels separate from anything the judge’s prompt contains. Chapter 26 calls an example that sits both in the guide and in a measurement set “a leaked exam answer,” and its three-set discipline (demonstrations, development set, sealed test set) applies to a judge as much as to a classifier.

The chapter’s development discipline made spatial: one pool of labeled tickets divided into three disjoint homes.
Figure 26.4 The chapter’s development discipline made spatial: one pool of labeled tickets divided into three disjoint homes. A handful live inside the guide as demonstrations; the development set is the sparring bench you re-run and study freely; and the sealed test set (in accent) is the envelope opened rarely, the only number you report. A ticket that tries to live in two sets at once is a leaked exam answer—struck out here, because disjointness is the whole point. Reuse this diagram

Worked example: computing Cohen’s kappa from a confusion table

Here is the whole calculation on an illustrative sample of 50 hand-labeled outputs, 35 of which you judged good and 15 bad. Three judges are compared: the lazy one, a first-draft judge, and the same judge after a rubric revision. Kappa separates them far more sharply than accuracy does.

Read each judge as four counts: human-good outputs the judge passed, human-good it failed, human-bad it passed, and human-bad it failed.

Judge Good→pass Good→fail Bad→pass Bad→fail Accuracy Expected p_e Kappa Bad outputs caught
Lazy (always pass) 35 0 15 0 0.70 0.70 0.00 0 of 15
Draft rubric 32 3 6 9 0.82 0.604 0.55 9 of 15
Revised rubric 33 2 3 12 0.90 0.588 0.76 12 of 15

Take the draft judge step by step. It agreed with you on 32 + 9 = 41 of 50 outputs, so p_o = 0.82. It said “pass” 38 times (0.76 of the sample) and you said “good” 35 times (0.70), so chance agreement on “good” is 0.76 × 0.70 = 0.532.

It said “fail” 12 times (0.24) and you said “bad” 15 times (0.30), so chance agreement on “bad” is 0.24 × 0.30 = 0.072. Together p_e = 0.604, and κ = (0.82 − 0.604) / (1 − 0.604) = 0.216 / 0.396 ≈ 0.55.

The same steps for the revised judge give p_o = 0.90, p_e = 0.72 × 0.70 + 0.28 × 0.30 = 0.588, and κ = 0.312 / 0.412 ≈ 0.76. On Landis and Koch’s 1977 bands, the draft is “moderate” (0.41–0.60) and the revision “substantial” (0.61–0.80). The book’s footnote treats 0.7 to 0.8, the neighborhood where human graders on judgment tasks tend to agree, as “a sensible bar for a judge meant to stand in for a human.”

Now the column accuracy hides. The draft judge let 6 of 15 bad outputs through. If its verdict gates a deploy, as a regression gate does, that 40% miss rate is the number that matters, and an 82% accuracy dressed it up as respectable.

How wide is the interval around kappa?

With 50 labels the 95% interval around kappa is roughly plus or minus 0.2 to 0.3, wide enough to separate a useless judge from a useful one but not a decent judge from an excellent one. Using the standard large-sample approximation, the revised judge’s kappa of 0.76 has an interval of about 0.56 to 0.96; the draft’s 0.55 spans about 0.28 to 0.81.

This is why the book warns that “Kappa computed on ten examples is a random number; you need a few dozen labels before it means anything.” It is also why the intervals of the draft and revised judges overlap here: on 50 labels you can say the revision is probably better, not that it is certainly better. For planning how many labels a given precision costs, the eval sample-size calculator does the binomial arithmetic. Treat the approximation itself as rough; at small samples, a bootstrap over your labeled items gives a more honest interval.

Which LLM-as-a-judge biases should you test for?

Test for four biases before trusting any judge: position bias, verbosity bias, self-preference and leniency. Chapter 16 says all four “will be waiting in your first judge,” and each has a cheap test you can run on your calibration sample and a mitigation that is mostly mechanical.

Bias What it looks like Evidence Test Mitigation
Position Preference depends on which answer comes first Zheng et al. (2023): strongest judge consistent under swap in 65% of pairs; one judge favored the first answer in 75% Score every pair in both orders; report the flip rate Accept only verdicts that survive the swap; score the rest a tie
Verbosity Padded answers score higher Zheng et al. (2023): a “repetitive list” attack fooled two judges on 91.3% of 23 answers, the strongest on 8.7% Correlate word count with score State in the rubric that length is not quality; length-match test pairs
Self-preference A model rates its own outputs, and its family’s, more warmly Panickssery et al. (2024): self-recognition correlates with self-preference Compare scores on your model’s outputs against another model’s of equal human-rated quality Use a judge from a different model family when stakes are real
Leniency Grades cluster at the top of the scale Practitioner observation reported in Chapter 16 Look at the score distribution Reserve the top grade for flawless; prefer pairwise or binary

Two cautions keep the table honest. Zheng and colleagues saw signs of self-preference but wrote that “our study cannot determine whether the models exhibit a self-enhancement bias”; the controlled evidence came later, from Panickssery, Bowman and Feng (2024). And bias sizes vary by judge and protocol: the 2026 study above found verbosity bias small across its cohort under a single pairwise rubric while finding severe position bias in two judges that otherwise scored above 0.95 on test-retest reliability. Consistency and fairness are separate properties, so measure both on your own setup.

Self-preference deserves one more sentence. It is the quiet flaw in any loop where a model grades its own homework, and it is a large part of why the evaluator in an evaluator-optimizer loop works better when it sees the output with fresh context, or comes from a different model.

When should you not use an LLM judge at all?

You should not use an LLM judge whenever code can decide the question: an exact value, a schema, a passing test suite, a numeric tolerance. The book’s rule is blunt: “the best judge is the one you did not need.” A code check is cheaper, deterministic, immune to the bias catalogue, and does not drift when someone upgrades the judge’s model.

Chapter 16 organizes graders as a ladder, climbed from cheap to expensive, stopping “as low as you can”:

Rung Grader Use it when Cost and stability
1 Programmatic check (an oracle) An objectively right answer exists: parses, tests pass, no leaked identifiers Fast, free, deterministic
2 Rubric decomposed into binary checks The criterion is fuzzy but splits into yes/no questions Steadier than any scale; tells you what broke
3 Model as judge Real open-endedness leaves no alternative Expensive, probabilistic, needs calibration and re-checks

Even at the top rung you can usually pull part of the judgment down a level. Give the judge a reference answer when one exists, since judges are much better at checking against a known-good answer than at spotting a subtle error on their own. Split the rubric into several small graders, each owning one dimension, so a failure tells you what broke rather than only that something did. And include negative cases in your eval set, tasks asserting what the agent should not do.

How often should you re-calibrate an LLM judge?

Re-calibrate whenever the judge’s model, its prompt or rubric, or your traffic changes materially, and spot-check on a schedule in between. The book’s reason is that “your data drifts and so, when its underlying model is updated, does the judge.” A calibration number describes one judge on one distribution at one time.

A practical routine has three triggers. First, any change to the judge itself (a new model version, a rewritten rubric, a new reference set) re-runs the full calibration sample before the new judge grades anything that gates a decision.

Second, a scheduled spot-check, where a person labels a fresh slice of recent outputs and you compare against the judge, catches drift in your traffic. Third, any surprising movement in judge-reported quality should send you to read the transcripts before you believe the number. If the judge’s pass rate jumps after a prompt change you thought was minor, the cheapest explanation is a judge problem, not a quality breakthrough.

When the calibration sample has been consulted so often that its failures are steering your rubric edits, Chapter 26’s advice applies: admit that it has become a second development set, fold it in, and label a fresh one.

Where does this method stop working?

This calibration method stops working where your human labels stop being trustworthy, where the task needs knowledge the judge lacks, and where high agreement measures the wrong thing. Kappa tells you the judge agrees with you. It cannot tell you that you were right.

Three limits are worth stating plainly. If your own labelers disagree with each other at a kappa of 0.5, no judge can be validated above that ceiling, and the work to do is on the rubric, not the judge; the comparison post on LLM-as-a-judge vs human evaluation covers how to split grading between people and models. If the outputs involve math, code or chained reasoning without a reference answer, expect the judge to miss subtle errors, as Zheng and colleagues observed.

And a judge can be consistent and agree with labels on your sample while missing the property you care about when it changes; a judge that never changes its verdict is perfectly stable and perfectly useless. Agreement is evidence of validity, not proof.

None of this argues against judges. It argues for treating one as the book does throughout: an instrument, “genuinely useful once calibrated, biased in documented directions out of the box, and never to be confused with ground truth.”

The one number to bring to the meeting

When someone asks whether your LLM judge is reliable, the answer is one sentence with a number in it, never a yes or a no.

A good version reads: “On 60 hand-labeled outputs from last month’s traffic, it agrees with us at a kappa of 0.74, catches 11 of 13 bad outputs, and survives position swaps on 92% of pairs.” (Those figures are illustrative.) If you cannot yet say that sentence, you have an opinion you have automated, and the method above is a week’s work at most.

Chapter 16, “LLM-as-a-Judge” (in the full book), builds the judge, its rubric habits and its calibration discipline; Chapter 26’s “Developing It Without Fooling Yourself”, also in the full book, carries the same discipline into classifiers. Chapter 16 in the online reader shows where the section sits (locked in the free edition); you can also browse the free evaluation pillar, or see the formats.

Questions readers ask

Is LLM as a judge reliable enough to replace human evaluation?
For scaling a judgment you have already defined and measured, often yes; for defining that judgment in the first place, no. A calibrated judge can grade thousands of outputs the way your labelers would, but the labels it is calibrated against still come from people, and periodic human spot-checks are what keep it honest.
How many human labels do I need to calibrate an LLM judge?
A few dozen at minimum, drawn from real traffic and including clearly bad outputs. Fifty labels are enough to tell a useless judge from a useful one, but the 95% interval around kappa is still about plus or minus 0.2 to 0.3, so you need more to distinguish a decent judge from a very good one.
What is a good Cohen's kappa for an LLM judge?
Around 0.7 to 0.8, roughly where careful human graders tend to agree with each other on judgment tasks. Landis and Koch's widely used bands call 0.61 to 0.80 substantial agreement, but the bands are conventions rather than laws, so set the bar by what the verdict gates.
How do I fix position bias in an LLM judge?
Run every pairwise comparison twice with the answers swapped, accept a verdict only when it survives the swap, and score inconsistent pairs as a tie. Zheng et al. (2023) proposed this conservative fix after finding that even their strongest judge gave the same verdict in both orders only 65% of the time.
Should the judge be a different model from the one being graded?
Where the stakes are real, yes. Panickssery et al. (2024) found that models can recognize their own outputs and that self-recognition correlates with rating those outputs more favorably, so a judge from a different family removes one source of inflated scores.

Sources

  1. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
  2. J. Richard Landis, Gary G. Koch (1977). The Measurement of Observer Agreement for Categorical Data (Biometrics 33(1):159–174)
  3. Arjun Panickssery, Samuel R. Bowman, Shi Feng (2024). LLM Evaluators Recognize and Favor Their Own Generations
  4. Justin D. Norman, Michael U. Rivera, D. Alex Hughes (2026). Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
  5. Hamel Husain (2024). Using LLM-as-a-Judge For Evaluation: A Complete Guide