LLM as a judge vs human evaluation is a question of allocation: code grades whatever has an objectively right answer, a model judge grades open-ended output at scale, and people set the criterion, label the calibration sample, arbitrate hard cases and read failures. A small human-labeled set proves the judge before it grades alone.
This post gives the pieces: a two-question sorting rule, an allocation table, a two-stage calibration, a worked example and a template to copy. It follows Chapter 16 of AI Agents, Engineered and marks every place where the arrangement is mine.
How reliable a judge is, which biases to test and how to compute the agreement statistic belong to the post that asks is LLM as a judge reliable. Here the space goes to who grades what.
LLM as a judge vs human evaluation: what can each grader see?
Each of the three graders sees a different part of an output: code sees structure, values and state; a model judge sees meaning measured against a written rubric; a person sees whether the criterion itself is right. Most write-ups of LLM as a judge vs human evaluation leave out the cheapest grader of the three.
The three-way split has a named source. An engineering guide from Anthropic, “Demystifying evals for AI agents” (published January 2026), states that “Agent evaluations typically combine three types of graders: code-based, model-based, and human.” Its recommendation: “We recommend choosing deterministic graders where possible, LLM graders where necessary or for additional flexibility, and using human graders judiciously for additional validation.” It is a vendor’s engineering guide, cited here for the taxonomy only.
Chapter 16 of the book arranges the same ground differently. It describes a ladder of three rungs: programmatic checks, a rubric decomposed into binary checks, and a model as grader. The human stands beside the ladder, defining the criterion, calibrating the judge and reading the failures. Where the two framings differ, this post follows the book.
| Grader | What it can see | What it cannot see | What it costs you |
|---|---|---|---|
| Code check | Format, exact values, schemas, test results, the state of a database or calendar, the tool calls in a trace | Meaning, tone, whether an argument holds | Machine time and the effort of writing the check |
| Model judge | Meaning, measured against a rubric and a reference, on thousands of outputs | Its own biases, anything outside its prompt, whether the rubric is the right one | Model calls on every run, plus the human hours that calibrate it |
| Person | Whether the criterion fits the product, domain context, the failure nobody predicted | Thousands of outputs; the same output the same way on a tired afternoon | Expert hours |
The cost column is qualitative: the book gives no price per graded example, and I found no primary source that does.
What does code see that the others miss?
Code sees the world the agent changed, and it reports the same answer every time. The book’s ordering rule is short: “On how to grade, climb the ladder from cheap to expensive and stop as low as you can.” Of the bottom rung it says: “Fast, free, deterministic; wherever an objectively right answer exists, nothing beats them.”
Testers call a check like this an oracle, a source of truth outside the thing being judged. Chapter 16 never uses that word and speaks of programmatic checks and ground truth; the book’s glossary supplies it.
What does a model judge see, and where is it blind?
A model judge reads for meaning, which is the one thing code cannot do, and it does so at a volume no reviewer can match. LLM-as-a-judge exists because, in the book’s words, “the gold-standard grader, a careful human, reads at human speed and bills at human rates.” It adds: “You cannot put a person inside a test suite that runs on every commit.”
Its blind spots are documented: a judge favors answers by position, rewards length, rates its own style warmly and clusters at the top of a scale. The sibling post covers the test and the mitigation for each. For allocation, every documented bias is a reason to move a criterion down to code or to a reference check wherever possible.
A person, the third grader, sees whether the question being graded is the right one. People have limits too, and the book names one: an undefined criterion produces a metric that is “measuring the mood of whoever graded that day.”
Which grader gets which criterion?
A criterion goes to code when two careful people would agree on the verdict and a program can compute it, to a model judge when they would agree and no program can, and back to a person for sharpening when they would disagree. Two questions, asked in that order, sort nearly every row of a grading plan.
- Can two people who know the domain, given the criterion and an output, agree on pass or fail? If they cannot, the criterion goes back to its owner to be sharpened or split. Nothing grades it yet.
- Can code decide it? If an exact value, a schema, a test, a count or a state query settles the verdict, a code check grades it.
- Otherwise a model judge grades it, with a reference answer whenever one exists, and only after a human-labeled sample has proven it.
The rule is my arrangement of the book’s prose, and the book does not state it as a three-way rule. Its pieces are all in Chapter 16. Question 1 is the chapter’s working standard for a criterion: “if two domain experts, given your criterion and an output, can disagree in good faith about whether it passed, sharpen the criterion, not the agent.” Question 2 is the bottom rung of the ladder. Step 3 is the chapter’s closing instruction on judges: “Reach for a judge only where real open-endedness leaves no alternative, and even there, anchor it with a reference when you can.”
The table applies the rule to the grading jobs the chapter discusses. The filter narrows it by grader, and every row stays on the page.
| Grading job | Why this grader | What proves the grader | Grader |
|---|---|---|---|
| An objective property of the output: it parses, a value is exact, a schema holds, no internal identifier leaked | An objectively right answer exists, and code reports it identically on every run | A few known-pass and known-fail outputs, sorted correctly by the check | code |
| Whether the world changed: the calendar entry, the database row, the generated code runs | The book: grade the state; a transcript can claim success while nothing changed | One known-good and one known-bad end state, sorted correctly | code |
| Safety and efficiency of the path: a destructive tool was never called, the step count stayed in bounds | The two things the chapter says an outcome cannot certify | One trace with the forbidden call and one without, sorted correctly | code |
| A fuzzy criterion that splits into yes-or-no questions | Binary verdicts “force sharper thinking and steady a metric faster than any scale”; each sub-question is sorted again | Whatever proves the grader each sub-question lands on | code, judge |
| Open-ended output no code can grade: a research brief, a refactored module, a reply to an angry customer | Only a model reads for meaning at the volume a test suite needs | Two-stage agreement on a human-labeled sample (below) | judge |
| A sampled slice of live traffic | The chapter assigns this to “a calibrated judge” | The same calibration, re-measured on a schedule against fresh human labels | judge |
| Defining the acceptance criterion | Someone must decide what “would I accept this?” means before anything is measured | Two people label the same outputs with it and agree | person |
| Ruling on ambiguous cases | The chapter wants “one designated human arbiter rather than a committee” | Rulings written back into the criterion, so the next labeler reaches the same verdict | person |
| Labeling the calibration sample | The judge is calibrated “against ground truth, which here means you” | A second person labels the same sample; their agreement is measured first | person |
| Reading failing runs | “An aggregate score cannot tell you whether a low number means a bad agent or a bad grader” | Sampled failures pass the chapter’s test: they “should seem fair” | person |
| Scheduled spot-checks of the judge and of production traces | Data drifts, and the judge drifts when its underlying model is updated | The judge’s agreement with the new labels, compared with the last measurement | person |
The fourth row carries two graders because the split sends its sub-questions to either side. The person rows describe jobs around grading: in this plan a person grades every output only inside the calibration sample.
Grading a finished run offline is a different act from approving an action while the agent is still running. That runtime control has its own design, and approval gates sized by consequence cover it.
What do the agreement studies show, and what do they leave open?
The published measurements show that the ceiling for a judge depends on the task and on the statistic, and that agreement between people moves just as much. Four are worth knowing precisely.
Zheng and colleagues (2023), “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” which I read in full. The setup was 80 multi-turn questions, answers from 6 models, and “58 expert-level human labelers,” described as mostly graduate students, giving around 3K votes. The statistic is raw pairwise agreement, with no correction for chance. On the first turn, judge and human agreed 66% of the time against 63% between two humans when ties were counted, where random agreement is 33%. With ties excluded the figures were 85% and 81%, where random agreement is 50%.
Its limits: the task was preference between two chat answers, the statistic is uncorrected, and the authors report limited capability “in grading math and reasoning questions” among their judges.
Thakur and colleagues, “Judging the Judges” (arXiv 2406.12624, first posted in 2024; I read version 6 in full). They chose “a clean scenario in which human alignment is very high”: short factual answers judged against references. Three human judges labeled 1200 sampled answers; measured against their own majority vote, they averaged a Scott’s pi of 96.2 ± 1.07 on a scale of 100, a chance-corrected statistic, and percent agreement of 98.52% ± 0.42%. The abstract on the paper’s arXiv page says of the best judge models that “they are still quite far behind inter-human agreement and their assigned scores may still differ with up to 5 points from human-assigned scores.”
The authors name the limit themselves, under the heading “Simplicity of the task”: judges are usually deployed on harder tasks, where people agree less.
Bavaresco and colleagues, “LLMs instead of Human Judges?” (arXiv 2406.18403). I read the abstract and the limitations section only. Across 20 datasets with human annotations and 11 models, they report that models “display substantial variability depending on the property being evaluated, the expertise level of the human judges, and whether the language is human or model-generated.”
Mukherjee and colleagues (2026), “The Geometry of LLM-as-Judge” (arXiv 2606.03043), abstract only. Across 42 judges on two Indic benchmarks, they report that “On subjective rubrics, judges agree with one another as much as humans do yet reach only 58-66% of human agreement” and, the line that matters for allocation: “On the one rubric with a verifiable answer, most of that gap closes.”
What do the four studies have in common?
Read together, the four show that the ceiling is task-specific. Two people agreed 63% to 81% of the time on chat preferences; on short factual answers, three people agreed with their own majority vote 98.52% of the time. Judges matched or trailed them depending on the task and the statistic, and checkability narrowed the gap.
None of them measured your criterion on your outputs. A published figure tells you a judge can work on a task of that kind; whether it works on yours is a number you produce yourself.
One Hacker News commenter (tempfile, July 2025) put the circularity objection plainly: “using a system to test itself gives you no reliable information. You need a control.” The calibration procedure below is that control.
Are human labels a gold standard?
Human labels are the best reference available and a noisy one, so the first measurement in any calibration is how well two people agree with each other. The book calls a careful human the gold-standard grader and, in the same chapter, describes how fast that standard degrades when the criterion is loose.
Two sources describe how human labels go wrong, each with a narrow scope. Hosking, Blunsom and Bartolo (“Human Feedback is not Gold Standard,” arXiv 2309.16349, abstract only) report that “the assertiveness of an output skews the perceived rate of factuality errors” in human annotation. Shankar and colleagues (2024, “Who Validates the Validators?”) name a phenomenon they call criteria drift, in which “users need criteria to grade outputs, but grading outputs helps users define criteria” as they go. Their evidence is a qualitative study with nine industry practitioners, so treat it as an observation with a name.
A Hacker News commenter (brookst, July 2025) wrote: “I have yet to meet the human who is 100% accurate at simple tasks done thousands of times.” The book offers a footnoted yardstick, that human graders on judgment tasks tend to agree at a kappa around 0.7 to 0.8, and calls it rough. It measured no human agreement itself.
The practical consequence is a ceiling. If two people applying your criterion agree only moderately, a judge cannot be shown to do better, because there is nothing steadier to compare it with. Measuring agreement between people first tells you whether to work on the judge or on the criterion.
How do you prove a judge with a small human-labeled set?
A judge is proven in two stages: two people label the same sample to establish the human ceiling, and then the judge is compared with the settled labels from that sample. The book’s own instruction covers the second stage, and the first is this post’s addition.
The book’s version reads: “Hand-label a representative sample (a few dozen outputs at minimum, drawn from real traffic, including some truly bad ones), then run the judge over the same sample and measure agreement.” The seven steps below add a second labeler in front of that instruction.
- Write the criterion as a pass-or-fail question with one example of a pass and one of a fail. Writing an LLM-as-a-judge rubric with levels that hold up is a subject of its own; these steps assume a rubric exists.
- Draw the sample from real traffic and seed it with truly bad outputs.
- Have two people label every item independently. Neither sees the other’s labels, and neither sees a judge’s verdict.
- Measure their agreement, raw and chance-corrected. This number is the human ceiling for the criterion as written.
- Read every disagreement. Where the two split because the criterion was vague, sharpen it. Then one arbiter rules on each disputed item, and the ruled labels become the calibration set.
- Run the judge over the same sample and measure its agreement with the calibration set. Split its errors by direction: bad outputs it passed, and good outputs it failed.
- Decide, record and schedule. Accept the judge for this criterion when it reaches the human ceiling and its error direction suits what the verdict gates. Record both agreement figures and the next re-check date.
Step 5 resolves a tension inside the book’s advice. The chapter prefers one arbiter, because “a single consistent taste is worth more to a metric than a negotiated average.” Two labelers look like the committee it warns against. They serve a different purpose here: the pair tests the criterion once, and the arbiter still owns every final label. That reconciliation is mine.
Teams often call the result of step 5 a golden dataset for LLM evaluation. Chapter 16 never uses the word and speaks of a hand-labeled sample and ground truth. Building that labeled set from real production traces is a larger job than the calibration sample needs, and this post stops at the sample.
How many labels, and who labels them?
The book gives one count, “a few dozen outputs at minimum,” and no number of labelers.
One practitioner gives more specific figures. In a guide to building and validating a judge (published 2024, revised 2026), Hamel Husain writes: “I start with around 30 examples and keep going until I do not see any new failure modes.” For validating an automated judge he asks for more: “Aim for about 100 examples per failure mode,” and “Below 60 examples, the confidence intervals are often too wide to support a useful conclusion.” These are one person’s heuristics with no study behind the counts, and they differ from the book’s wording, so I print both and merge neither.
On who labels, the same guide points to a principal domain expert, and the book says “you.” Both mean someone who would be believed if they overruled the judge. The worked example below uses 60 items because the arithmetic is easy to follow; that count is illustrative.
For what a given sample size can distinguish, the eval sample size calculator does the arithmetic.
Worked example: sixty outputs, two people, one judge
In this illustrative example, two people agree at a kappa of 0.64, and the judge reaches 0.72 against the labels their arbiter settled. Every number in this section is invented for the example and computed by hand; none is a measurement of a real system.
The setting is a support agent that handles refund requests. The criterion under test: the reply answers the question the customer asked and states the next step. Sixty replies are drawn from last month’s traffic, including replies that support staff had already flagged as poor.
Stage 1: what is the human ceiling?
Two people who know the refund policy, A and B, each label all sixty without conferring.
| B says pass | B says fail | Total | |
|---|---|---|---|
| A says pass | 38 | 4 | 42 |
| A says fail | 5 | 13 | 18 |
| Total | 43 | 17 | 60 |
They agree on 38 + 13 = 51 of 60, so observed agreement is 0.85. Agreement expected by chance is (42 × 43 + 18 × 17) / 60², which is 2112 / 3600 = 0.5867. Cohen’s kappa is (0.85 − 0.5867) / (1 − 0.5867) = 0.64. The approximate 95% interval runs from 0.42 to 0.86.
The calculator below opens on this table. It was built for a person and a judge, so its labels say “human” for the rows and “judge” for the columns. In stage 1 the rows are person A and the columns are person B, and its sentence about a judge standing in for a human describes person B.
With JavaScript on, the LLM Judge Agreement Calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the LLM Judge Agreement Calculator on its own page to share a result by link.
The tool reports κ = 0.64 on 60 items, in the band Landis and Koch call substantial, with an interval of 0.42 to 0.86, observed agreement 0.85 and chance agreement 0.59 (0.5867 rounded). It notes that always answering “pass” would score 0.70 accuracy, 15 points under the observed 0.85. For the pass class it shows precision 0.88 and recall 0.90; for the fail class, precision 0.76 and recall 0.72. Its level check reads 72% against 70%, about level.
It also says 0.64 is below the bar the book suggests, about 0.7 to 0.8. Read for two people, that is the finding: the criterion still leaves room for good-faith disagreement. The tool’s to-do list shows where to look: nine disputed items, four that A passed and B failed and five the other way.
Stage 2: does the judge reach that ceiling?
The arbiter reads the nine disputed replies and rules four of them a pass and five a fail. The calibration set therefore holds 38 + 4 = 42 passes and 13 + 5 = 18 fails. Most of the nine turned on whether “we’ll be in touch” states a next step. The arbiter rules that it fails, and the ruling is written into the criterion.
The judge then grades the same sixty replies.
| Judge says pass | Judge says fail | Total | |
|---|---|---|---|
| Calibration set: pass | 39 | 3 | 42 |
| Calibration set: fail | 4 | 14 | 18 |
| Total | 43 | 17 | 60 |
Observed agreement is 53 / 60 = 0.883. Chance agreement is again 2112 / 3600 = 0.5867, because the totals happen to match. Kappa is (0.8833 − 0.5867) / (1 − 0.5867) = 0.72, with an approximate 95% interval of 0.52 to 0.91. Type 39, 3, 4, 14 into the calculator to see the same figures, along with a recall of 0.78 on the fail class.
What does the comparison tell you to do?
Three readings follow from the two tables, and each one maps to an action.
First, the judge has reached the human ceiling as far as this sample can show. Its 0.72 against the settled labels is at least as high as the 0.64 the two people reached with each other. The intervals overlap almost entirely, so the data support “as good as a second person” and nothing stronger. The comparison also flatters the judge slightly, since settled labels are cleaner than raw ones.
Second, the next gain is in the criterion. Tuning the judge’s prompt further would chase a target the people themselves hit only at 0.64. Re-labeling a fresh sample with the sharpened criterion is the step that can raise both numbers.
Third, look at the direction of the judge’s errors. It passed 4 of the 18 bad replies and failed 3 of the 42 good ones. The book’s guidance is that “which one you can live with depends on what the verdict gates.” A weekly quality trend can absorb four missed failures in eighteen. A check that blocks a release on reply quality would need a stricter judge or a human read of the passes.
What does the grading plan look like on paper?
A grading plan is one row per criterion, with the grader, the reference, what the verdict gates, the proof that the grader works, the arbiter and the trigger for a re-check. Copy the template below and replace its example rows with your own.
GRADING PLAN: <system name> owner: <name> date: <date>
criterion: <one pass-or-fail question>
split from: <the compound criterion it came from, or "none">
two people can agree: <yes | no: sharpen before grading>
code can decide: <yes | no>
grader: <code | judge | person>
reference available: <the expected value, state or answer, or "none">
what the verdict gates: <dashboard | merge | release | nothing yet>
proof of the grader:
code -> known-pass and known-fail cases the check sorts correctly
judge -> human-human agreement: <raw %, kappa, N, date>
judge-human agreement: <raw %, kappa, N, date>
bad outputs passed: <n of n> good outputs failed: <n of n>
person -> second labeler's agreement on the same items: <raw %, kappa, N>
arbiter: <one named person>
re-check trigger: <schedule, judge model change, criterion change>
--- two example rows (illustrative) ---
criterion: reply names the refund policy that applies to the case
grader: code
reference available: expected policy identifier, stored with each eval task
what the verdict gates: merge
proof of the grader: 6 known-pass and 6 known-fail replies sorted correctly
arbiter: support policy owner
re-check trigger: policy catalog changes
criterion: reply answers the question asked and states the next step
grader: judge
reference available: none
what the verdict gates: dashboard
proof of the grader: human-human 85%, kappa 0.64, N 60
judge-human 88%, kappa 0.72, N 60
bad passed 4 of 18; good failed 3 of 42
arbiter: support policy owner
re-check trigger: monthly, and on any change to the judge's model
The proof field has three forms, one per grader. The person form is for the low-volume case described under the limits below, where people read every output and no judge is used. A plan row with a grader and no proof is unfinished.
Tracing platforms that include annotation queues (Langfuse and LangSmith are two the book lists among its labeled examples of that category) and open-source evaluation libraries (pydantic-evals and promptfoo are two examples as of October 2026) package parts of the loop. They are examples of two categories; I verified no feature claims for any of them and recommend none.
How does one criterion move through the plan?
Take a criterion as it might first be written: the summary must be faithful and concise. It comes out of the plan as two rows.
Question 1 stops it. Two careful people will split on “faithful and concise,” because it holds two properties and defines neither. The criterion goes back to its owner, who splits it. “Concise” becomes a word limit the team chooses. “Faithful” becomes “every claim in the summary is supported by the source text.”
The word limit passes both questions, so code grades it, and a handful of summaries on each side of the limit prove the check. The faithfulness row passes question 1 and fails question 2, so a judge grades it with the source text as its reference. Its proof is the two-stage agreement on a labeled sample. Each row ends with one grader and one proof.
What stays with a person after the judge passes?
Four jobs stay with a person once a judge has passed calibration: owning the criterion, arbitrating the ambiguous cases, reading failures and production traces, and re-checking the judge. Calibration moves the per-output grading to the judge and leaves each of these where it was.
Owning the criterion. The standard the judge scales was set by a person, and only a person can revise it when the product changes. The book’s glossary calls a judge “an instrument that must itself be calibrated against human judgment before it is trusted,” which keeps a person upstream of every verdict.
Arbitrating. New ambiguous cases keep arriving. They go to the same arbiter, and each ruling goes back into the criterion.
Reading failures and traces. A score tells you that something failed and says nothing about why. The chapter adds that “someone should still read production traces on a schedule” even while a calibrated judge samples live traffic.
Re-checking the judge. The chapter tells you to “keep spot-checking on a schedule,” because data drifts and so does a judge whose underlying model is updated. The re-check trigger in the template is where that schedule lives.
A fifth job is my own addition: the human read before a consequential release. A Hacker News commenter (sdenton4, July 2025) drew the line this way: “it seems fine to use LLMs to judge your work in progress, but we should be requiring human evaluation for ‘final’ results.” The book’s version is a budget rule: “buy as much certainty as the cost of being wrong justifies, and no more.” For a decision that is costly to reverse, a person reading a sample of the runs is cheap certainty.
The book is direct about the price of all this: “Curating tasks and calibrating graders takes senior-human hours, the expensive kind.” The allocation exists to spend those hours where only a person can do the work.
Where does this allocation break down?
The allocation breaks down in three places: when no reference holds still, when the sample is too small to separate the judge from the ceiling, and when judge and people share a blind spot.
No stable ground truth. Calibration needs labels that stay put. For open-ended research with no reference answer, people may be unable to say what a pass is until long after the run. This procedure cannot prove a judge for such work.
Small samples. At sixty items the intervals in the worked example overlap almost entirely. That is enough to tell a useless judge from a usable one and too little to rank two decent judges.
Shared blind spots. Agreement between a judge and people can come from a bias they share. Bavaresco and colleagues write in their limitations that “there may be domains where human annotators and LLM evaluators appear aligned simply because they are affected by similar biases.” High agreement is evidence that the judge applies your criterion; whether the criterion tracks what users need is a separate question that production signals answer.
One alternative deserves a fair statement. A team with low volume and high stakes per output can skip the judge and have people read everything, with code checks underneath. The judge earns its row only when the volume exceeds what people will reliably read.
One plan to bring to the design review
The useful answer to “LLM as a judge vs human evaluation” is a table with a name in every row. Each criterion has one grader, one proof that the grader works, one arbiter and one date for the next check.
If you do one thing this week, run stage 1 before reopening the LLM as a judge vs human evaluation debate. Have two people label the same few dozen real outputs against your most contested criterion and measure how often they agree. That single number tells you whether your next hour belongs to the judge or to the criterion.
The method comes from Chapter 16, “Evaluating Agents”, which is in the full book; the online reader shows where the section sits and keeps it locked. Chapters 1 and 2 and the glossary are free to read online. The evaluation pillar collects the related posts and tools, including the judge agreement calculator on its own page. If the plan earned its place in your review, see the formats the book comes in.
Questions readers ask
- Is using an LLM to grade an LLM circular?
- It is circular until a control exists. The control is a sample that people labeled before they saw the judge's verdicts, with the judge's agreement measured against those labels. A judge that was never compared with human labels gives no independent information about the system it grades.
- Should final results still get human evaluation when a judge is in place?
- For a decision that is costly to reverse, yes: a person reads a sample of the runs behind the number before signing. The book's rule is to buy as much certainty as the cost of being wrong justifies, so a dashboard trend can rest on the judge while a consequential release gets a human read as well.
- What if the human labelers disagree with each other?
- Treat the disagreement as a finding about the criterion. Chapter 16 says that when two domain experts can disagree in good faith about a pass, you sharpen the criterion; then one designated arbiter rules on the remaining ambiguous cases. A judge cannot be validated above the agreement the people reach.
- How do you validate an LLM judge against human review?
- Have two people label the same sample of real outputs without seeing each other's labels or the judge's. Measure their agreement, settle the disagreements with one arbiter, then run the judge on the same sample and measure its agreement with the settled labels, split by the direction of its errors.
- Can a committee of LLM judges replace the human labels?
- One recent study points to no. Mukherjee and colleagues (2026, abstract only) report that on subjective rubrics judges agreed with one another as much as humans do yet reached only 58-66% of human agreement. Judges agreeing with each other measures consistency among models; a human label is still what they are checked against.
Sources
- Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Thakur, Choudhary, Ramayapally, Vaidyanathan, Hupkes (2024). Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
- Bavaresco et al. (2024). LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Mukherjee, Hamna, Bali, Sitaram (2026). The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
- Shankar, Zamfirescu-Pereira, Hartmann, Parameswaran, Arawjo (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
- Hosking, Blunsom, Bartolo (2023). Human Feedback is not Gold Standard
- Hamel Husain (2024). Using LLM-as-a-Judge For Evaluation: A Complete Guide
- Anthropic (2026). Demystifying evals for AI agents
- tempfile (2025). Hacker News comment on needing a control (handle: tempfile)
- brookst (2025). Hacker News comment on human evaluators at scale (handle: brookst)
- sdenton4 (2025). Hacker News comment on requiring human evaluation for final results (handle: sdenton4)