What does an LLM judge agreement calculator measure?
An LLM judge agreement calculator measures how well a model grader agrees with your own labels once luck is taken out. You hand-label a sample, run the judge on the same items, and compare. The headline number is Cohen’s kappa, a chance-corrected agreement statistic, reported with an interval, a per-class breakdown and the confusions that tell you what to fix.
The method is Chapter 16’s, and the chapter’s verdict on skipping it fits in one line: “A judge you have not calibrated is an opinion you have automated.” The book’s recipe is short: “Hand-label a representative sample (a few dozen outputs at minimum, drawn from real traffic, including some truly bad ones), then run the judge over the same sample and measure agreement.” This calculator does the measuring. You can paste the pairs as two columns (human label, judge label), or type the counts into a confusion matrix if you already have one. Two classes or six, the arithmetic is the same.
An LLM-as-a-judge is a language model armed with a written rubric, scoring or comparing outputs the way a human reviewer would. The calculator treats it the way the book does: as an instrument, “genuinely useful once calibrated, biased in documented directions out of the box, and never to be confused with ground truth.”
Why does raw accuracy lie about a judge?
Raw accuracy lies because it rewards a judge for matching your base rate without telling good outputs from bad ones. When most outputs are good, a judge that always says “good” earns a high accuracy and carries no information. Kappa prices that laziness at exactly zero, which is why the book asks you to report it alongside accuracy.
Chapter 16 puts it more sharply than I can: “Raw accuracy is the number everyone quotes and the one that lies: if seventy percent of your outputs are good and the judge answers “good” every single time, it scores seventy percent while carrying zero information.” The calculator shows this floor on every result. It finds your majority class, reports the accuracy a judge would get by always answering it, and tells you how many points your judge clears it by.
Chapter 26 meets the same trap in a ticket classifier, where a model that answers bug to everything scores seventy percent on a sample that is seven-tenths bugs. The book calls it “Chapter 16’s lazy judge, returned in a new uniform.” The cure is the same in both places: count agreement beyond chance, then look past the single number.
How is Cohen’s kappa calculated?
Cohen’s kappa compares observed agreement with the agreement you would expect if the judge and the human labeled independently at their own rates. Observed agreement p_o is the share of items on the diagonal of the confusion matrix. Expected agreement p_e multiplies each class’s human share by its judge share and sums. Then κ = (p_o − p_e) / (1 − p_e).
The worked example comes from the companion article Is LLM as a judge reliable?, and the presets of this LLM judge agreement calculator reproduce it exactly. Fifty outputs, 35 of which you judged good and 15 bad, graded by three judges (the numbers are illustrative):
| Judge | Good→pass | Good→fail | Bad→pass | Bad→fail | Accuracy | p_e | Kappa | Approx. 95% interval |
|---|---|---|---|---|---|---|---|---|
| Lazy (always pass) | 35 | 0 | 15 | 0 | 0.70 | 0.700 | 0.00 | −0.42 to 0.42 |
| Draft rubric | 32 | 3 | 6 | 9 | 0.82 | 0.604 | 0.55 | 0.28 to 0.81 |
| Revised rubric | 33 | 2 | 3 | 12 | 0.90 | 0.588 | 0.76 | 0.56 to 0.96 |
Take the draft judge. It agreed with you on 32 + 9 = 41 of 50 items, so p_o = 0.82.
It said “pass” 38 times (0.76) and you said “good” 35 times (0.70); it said “fail” 12 times (0.24) and you said “bad” 15 times (0.30). Expected agreement is 0.76 × 0.70 + 0.24 × 0.30 = 0.604, and κ = 0.216 / 0.396 ≈ 0.55. The revised judge gives κ = 0.312 / 0.412 ≈ 0.76.
The interval is the tool’s addition. It uses the standard large-sample approximation, SE = √(p_o(1 − p_o) / (N(1 − p_e)²)), and reports κ ± 1.96·SE, the formula McHugh (2012) gives in her review of the statistic. Treat it as rough at small samples; a bootstrap over your labeled items gives a more honest interval when N is in the dozens.
How many labels do you need before kappa means anything?
You need a few dozen labels before kappa means anything, and the calculator raises a warning below 30, a cut-off it chooses because the book gives no exact number. With fewer, the interval is so wide that a useless judge and a good one can produce the same number. The book is blunt: “Kappa computed on ten examples is a random number; you need a few dozen labels before it means anything.”
The table above shows why. On 50 labels, the draft judge’s interval runs from 0.28 to 0.81 and the revised judge’s from 0.56 to 0.96. They overlap.
You can say the revision is probably better; you cannot say it is certainly better. That is the honest state of a 50-label calibration, and the interval is on screen so nobody rounds it away in a meeting.
Two things make the labels count. Draw them from real traffic, not from examples you wrote while designing the rubric, and seed the sample with failures, or a stream that is ninety percent good will leave you three bad outputs to learn from. For planning how many items a given precision costs, the eval sample-size calculator does the binomial arithmetic, and the article How many eval examples do you need? explains the trade-off.
What kappa should an LLM judge reach?
An LLM judge meant to stand in for a human grader should reach a kappa of about 0.7 to 0.8. That is roughly where careful human graders agree with each other on judgment tasks, so a judge in that neighborhood is about as good a substitute for you as a second person would be. The calculator draws that bar on the scale.
The book’s footnote ties the number to the standard reference: humans on such tasks “tend to agree at a kappa around 0.7 to 0.8—the top of what Landis and Koch’s widely used interpretation bands call “substantial” agreement.” Landis and Koch’s 1977 paper proposed labels for the whole range, and the calculator shows them as reference:
| Kappa | Landis–Koch label |
|---|---|
| below 0 | poor |
| 0.00–0.20 | slight |
| 0.21–0.40 | fair |
| 0.41–0.60 | moderate |
| 0.61–0.80 | substantial |
| 0.81–1.00 | almost perfect |
Bands like these are conventions, not laws, and other fields draw them tighter. McHugh, writing for clinical research, argues that “any kappa below 0.60 indicates inadequate agreement among the raters.” The right bar depends on what the verdict gates. A judge that ranks drafts for a human editor can live lower than one that blocks a deploy as a regression gate.
Zheng and colleagues’ study, the one that legitimized the technique, deserves the same care. Their 2023 paper reported that strong judges “can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans.” The book’s reading is that “a good judge approximates a human grader about as well as a second human does,” and no better.
Is the judge harsher or kinder than you?
The level check answers whether the judge is systematically harsher or kinder than you, which agreement alone cannot show. For a pass/fail judge it compares the judge’s pass rate with yours; for numeric scores it compares the averages. A judge can agree with you often and still pass six points more of everything.
Chapter 16 names the trap: “high agreement on ranking can coexist with a judge that is systematically harsher or kinder than you; check the level, not just the correlation.” In the worked example the draft judge passes 76% of outputs against your 70%, a six-point lean toward kindness, and the calculator says so in words. It calls a gap of two points or less (0.1 on a numeric scale) level; that tolerance is the tool’s own choice. When the pass label is not obvious, choose it from the menu.
The per-class table splits the same question further. “For a pass/fail judge, split out precision and recall as well: a judge that misses bad outputs and a judge that fails good ones are different problems hiding behind the same accuracy.” The draft judge’s recall on “bad” is 0.60: it caught 9 of 15 bad outputs. If its verdict gates anything, that miss rate is the number to bring to the meeting, and its 82% accuracy dressed it up as respectable.
How do you read the confusion matrix as a to-do list?
You read the confusion matrix by its off-diagonal cells, folded into pairs of classes that swap with each other, biggest first. Each recurring pair is a boundary the rubric does not yet draw clearly. Chapter 26 turns that observation into a working habit: write the next section of the rubric about the pair that confuses most.
The book’s sentence: “read the confusions in pairs (which classes swap with which), because that table is a to-do list: each recurring swap is the next contrastive section of the guide.” The three-class ticket preset shows it working. Of 120 tickets, bug and billing swap 11 times, more than every other confusion combined, so the next revision of the guide should contain a passage that tells bug from billing with one example of each.
Chapter 26 adds two guards, and the calculator implements both. “Watch the abstain rate beside accuracy, too; an instrument can be made to look arbitrarily accurate by teaching it to flag everything hard, and the flag volume is where that trick shows up.” Type the abstentions into the matrix, or write “abstain” in the judge column, and the rate appears beside the agreement. The tool highlights a rate above 15%; that warning line is its own, not the book’s.
The second guard is the three-set discipline: in-guide demonstrations, a development set and a sealed test set, kept disjoint. “An example that sits in the guide and also in a measurement set is a leaked exam answer.” The optional leak check compares your calibration items against the examples inside the judge’s prompt and lists any exact matches.
How do you test a judge for position and verbosity bias?
You test position bias by running each pairwise comparison in both orders and counting how many verdicts survive the swap, and verbosity bias by correlating each output’s word count with its score. Both tests run on data you already have from calibration, and the calculator’s optional sections do the counting.
For position, the book’s mitigation is mechanical: “run every pairwise comparison both ways, accept only verdicts that survive the swap, score the rest a tie.” Paste the two verdicts per comparison, naming the winning answer rather than the slot, and the tool reports the survival rate, the share that end as ties, and how often the first slot won. For scale, the book notes that in the foundational study even the strongest judge’s verdict survived the swap “in only about two-thirds of comparisons.”
For verbosity, the instruction is to “check your results for a correlation between word count and score; if you find one, your judge is grading by the pound.” The tool computes Spearman’s rank correlation, which tolerates skewed length distributions. A positive correlation is not proof of bias, because longer answers are sometimes more complete. Add your own scores as a third column: if the judge’s correlation clearly exceeds yours, length is buying grades. What counts as “clearly” is the tool’s rule of thumb, not the book’s: it flags a gap of 0.2 or more between the two correlations (or a correlation of 0.3 or more when you give no scores of your own), and it labels strength with the usual bands at 0.1, 0.3 and 0.5.
Where does this calculator stop being useful?
This LLM judge agreement calculator stops being useful where your own labels stop being trustworthy. Kappa tells you the judge agrees with you; it cannot tell you that you were right. If two of your labelers agree with each other at 0.5, no judge can be validated above that ceiling, and the work to do is on the rubric.
Three further limits are worth saying plainly. Kappa depends on the class mix: the same judge can score differently on a balanced sample and a skewed one, so compare kappas only across samples drawn the same way. The leak check catches exact text matches, not paraphrases. And a calibration describes one judge on one distribution at one time; the book’s reason to re-check on a schedule is that “your data drifts and so, when its underlying model is updated, does the judge.”
The best judge, Chapter 16 adds, is the one you did not need: if code can decide the question, a code check is cheaper and immune to every bias above. When a judge is the right tool, run it through the LLM judge agreement calculator and bring one sentence with numbers in it to the decision. The free evaluation pillar places this tool among the other measurements an eval set needs, the pass@k calculator covers repeated runs, and the full treatment is in Chapter 16 and Chapter 26, where the judge reappears as a classifier shipped as the product itself.
Questions readers ask
- My judge agrees with me 85% of the time. Why is kappa so much lower?
- Because part of that 85% would happen by chance. If most of your outputs are good and the judge says good most of the time, the two of you agree often without the judge discriminating anything. Kappa subtracts that expected agreement, so a judge that only matches your base rate scores zero however high its accuracy.
- How many labels do I need before kappa means anything?
- A few dozen at minimum, and the calculator warns below 30, a cut-off of its own. Even at 50 labels the approximate 95% interval spans roughly plus or minus 0.2 to 0.3, enough to tell a useless judge from a useful one but not a good judge from an excellent one.
- What kappa should an LLM judge reach?
- About 0.7 to 0.8, the neighborhood where careful human graders tend to agree with each other on judgment tasks. Landis and Koch's bands call 0.61 to 0.80 substantial agreement, but the bands are conventions; set the bar by what the verdict gates.
- What is position bias and how do I test for it?
- Position bias is a judge favoring an answer because of where it sits in the prompt. Run every pairwise comparison twice with the answers swapped, keep only verdicts that survive the swap, and score the rest a tie. The calculator reports the survival rate and how often the first slot won.
- Why does the tool complain when calibration items overlap my prompt examples?
- Because an item the judge has already seen as a worked example will be graded correctly for the wrong reason. Chapter 26 calls it a leaked exam answer: it inflates agreement without measuring judgment, so keep calibration items and in-prompt examples disjoint.
Sources
- Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- J. Richard Landis, Gary G. Koch (1977). The Measurement of Observer Agreement for Categorical Data (Biometrics 33(1):159–174)
- Jacob Cohen (1960). A Coefficient of Agreement for Nominal Scales (Educational and Psychological Measurement 20(1):37–46)
- Mary L. McHugh (2012). Interrater reliability: the kappa statistic (Biochemia Medica 22(3):276–282)