Home / Blog / Evaluating and observing agents / Writing an LLM as a Judge Rubric: Ordinal Scale…

Evaluating and observing agents

Writing an LLM as a Judge Rubric: Ordinal Scales That Hold Up

Write an LLM as a judge rubric with four anchored levels, a worksheet before the verdict and two bias tests. Copy the template and run the checks.

By Enrique Gutiérrez · Published · 23 min read

An LLM as a judge rubric is the written marking guide a model grader follows: one criterion, a short ordinal scale whose levels are each defined in a sentence, a worksheet the judge fills in before its verdict, and an output your code can parse. A level holds up when someone could be shown to have applied it wrongly.

One caveat belongs this early. I found no study that tests an anchored short scale against a bare 1–10 score head to head, and one study measured 1–10 tracking human raters better than 1–5. The case for a short anchored scale is therefore an argument from AI Agents, Engineered, plus indirect evidence that I report below with its limits.

This tutorial writes one LLM-as-a-judge rubric end to end for one criterion, then runs two bias tests on it.

What is an LLM as a judge rubric?

An LLM as a judge rubric is the document that tells a model grader what to look for and what each grade means. It plays the part an examiner’s marking guide plays in a school: the grader arrives educated and knows nothing about your local standard until you write it down.

The analogy is Chapter 26’s, introduced there for classifiers: “Building the LLM classifier is hiring an experienced examiner and writing the marking guide.” The same chapter extends it to the judge, which it describes as the same instrument pointed at your own agent’s outputs. An LLM-as-a-judge setup is, in that reading, a classifier whose classes happen to be grades.

Chapter 16 gives the habit in one line: “Write the rubric as if for a new hire: name the criteria, define what each grade means, include a worked example of a pass and a fail”. Most judge prompts I see skip the middle clause. They say “rate this answer 1–10 for quality” and define nothing.

Chapter 26 explains why that fails with a severity score: “nothing anywhere defines what 0.7 is, so no verdict of 0.7 can be wrong, and an instrument that cannot be wrong cannot be trusted”. A rubric exists to make the judge’s verdicts checkable.

Which scale types can a judge use?

A judge can use four families of scale: pass/fail, a short anchored ordinal scale, a long numeric scale, or a pairwise choice. An ordinal scale is a score made of ordered levels, like grades, where the gaps between levels carry no fixed size. Each family answers a different question and breaks in its own way.

Scale Good for What goes wrong What the cited evidence says
Pass/fail Release gates; criteria with a clear line Hides degree; a mixed item has to be forced one way Recommended over 1–5 by Husain and Shankar (2025), as opinion with no data offered
Binary checks, counted Fuzzy criteria that split into yes-or-no questions The count treats every check as equally important Chapter 16’s middle rung; no head-to-head study found
Anchored ordinal, 3–4 levels Degree of success, triage, trends Adjacent levels blur when a boundary sentence is missing No study found that isolates anchors on a short scale
1–5 with no descriptors Quick first look One value dominates; 3 versus 4 is undefined Liu et al. (2023): “one digit usually dominates the distribution of the scores”
1–10 or 1–100 Looks precise on a dashboard Most of the range goes unused Stureborg et al. (2024): on 1–100, the range 1–60 “is almost entirely ignored”
Pairwise (A or B) Head-to-head choices between two versions Slot order and surface features swing the choice Tripathi et al. (2025), from the abstract: preferences flipped in “about 35%” of cases with distractor features, against “only 9%” for absolute scores

The rows are mine; the right-hand column carries only what each source says. Three rows have no experiment behind them.

Chapter 16 adds a practitioner’s observation about the fifth row, that “a ten-point scale behaves, in practice, like a noisy three-point one”. The book gives no citation for it, so read it as experience.

What do the studies of scale length find, and what do they leave open?

The studies of scale length find that numeric range changes how well a judge tracks people, and they disagree on the direction. Four papers bear on the choice. None of them compares defined levels with undefined ones on a short scale, which is the comparison this tutorial depends on.

Li and colleagues (2026, arXiv preprint) ran six LLM judges on 0–5, 0–10 and 0–100 scales over 5,497 items from six benchmarks, with 12 human annotators scoring 150 pooled items. Their finding: “a 0-5 scale yields the strongest absolute agreement between LLM judges and human consensus”, “while 0-10 is consistently the weakest choice”. Fractional scores were allowed on every scale and no level descriptors were compared, so the study isolates range. The authors add that the best scale “can be benchmark-dependent”.

Stureborg, Alikaniotis and Suhara (2024, arXiv preprint) point the other way. On one summarization dataset, their table of average Kendall’s tau against human judgment reads 0.339 for a 1–5 star scale, 0.428 for a 1–10 score and 0.383 for a 1–100 score. In their words, “1-10 score performs best on average”. That is one dataset and one task, and it cuts against any blanket rule that fewer levels is better.

Liu and colleagues (2023, arXiv preprint) built a form-filling judge with criteria on a 1–5 scale and reported that “one digit usually dominates the distribution of the scores, such as 3 for a 1 - 5 scale”. Bunching is a property of short bare scales too.

Murugadoss and colleagues (2024, arXiv preprint) tested prompts with growing detail, up to a full rubric with per-score instructions. From the paper’s introduction: more granular rubrics improved Pearson correlation with human judgments “by only as much as 4%”, though “some individual metrics may benefit”. The criteria were generic text-quality ones.

What does the literature not show?

The literature I could find does not show that anchored ordinal scales beat 1–10 scores. No paper above holds the items and the judge fixed and varies only whether the levels carry written definitions on a short scale.

So the claim in this post’s title is narrower than it may sound. “Hold up” here means only that a level can be audited against a sentence. A published correlation gain is a separate claim, and I have none to cite.

I can offer one inference about the Murugadoss result, labeled as mine. A capable judge already knows what coherence means, so describing it adds little. A team’s local criterion, such as what counts as resolving a ticket under your refund policy, is something no model learned in training.

Why does the book prefer binary in one chapter and a small scale in another?

The book prefers binary checks for grading an agent’s outputs in Chapter 16 and recommends a small ordinal scale for scorers in Chapter 26, and it never spells out how the two fit. Chapter 16 says “binary verdicts force sharper thinking and steady a metric faster than any scale”. Chapter 26 says to use a small ordinal scale, “one to four, say”.

My own reading, which the book does not state: the two form a ladder. Start with yes-or-no questions. Keep an ordered scale on top of them only when the reader of the number needs degree, and when you can write the sentence between each pair of neighbors.

A practitioner on Hacker News makes the case for the scale side. In a 2025 comment, the user nimitkalra wrote: “Better to use a discrete scale of 1..3 or 1..5 and specify exactly what makes a sample a 2/5 vs a 4/5”. Hamel Husain and Shreya Shankar take the other side in their 2025 FAQ: “Numeric labels are advanced and usually not necessary.” Both are opinions from experience, offered without data.

The rubric below sits on both rungs at once. Its four levels are computed from four yes-or-no answers, so it collapses to pass/fail by moving one threshold. That design is this post’s, chosen so the tension costs you nothing.

How do you write the rubric, step by step?

You write an LLM as a judge rubric in five steps: one criterion, the levels, the boundary sentences, the worksheet, and the output format. The worked example grades a support agent’s reply on whether it resolves the problem the customer stated. Every number and policy line in it is illustrative.

Step 1: one criterion, in one sentence

Name a single criterion and say what is out of scope. “One criterion per scale” is this post’s rule; the book’s nearest statement is Chapter 16’s habit of “several small graders, each owning one dimension (accuracy, format, safety)”, which “tell you what broke”.

The best outside support comes from Stureborg and colleagues. When one judge output scored several attributes, the judge’s scores for two of them correlated at r = 0.979, while human raters’ scores for the same two correlated at r = 0.315. The judge had mostly produced one opinion and written it down twice.

Separate calls cost money. One GitHub user, chaholl, put the bind this way in 2026: “run five evals (five judge calls for one judgement), or run one and keep four dimensions as inert detail”. I take the first option for any score that gates a decision, and accept the correlation knowingly elsewhere.

Step 2: write the levels as mini-classes

Write each level as a small class with a name, what it includes and what it excludes. Chapter 26 gives the reason through its examiner: “descriptors make graders agree, and bare numbers make them guess”.

How many levels is a judgment call the book leaves open. I use four here because the criterion has four states I can tell apart in prose. Chapter 26 notes one side effect: “An even number of levels removes the comfortable midpoint”. Its verdict on odd versus even is “Neither is wrong.”

Chapter 16’s remedy for lenient judges applies to the top level: “reserve the top grade for flawless”. Level 4 below demands everything.

Step 3: write the boundary sentences

Write one sentence for each pair of neighboring levels that says which side an item falls on. Four levels need three. Chapter 26’s test for whether a level is real: “if you cannot write the sentence that separates level n from level n+1, you do not have two levels; merge them”.

Add a hard rule where one fact must override the rest. The book’s severity example carries one: “Boundary: money handled wrongly is never below 3.” Mine caps any reply that contradicts policy at level 1.

CRITERION
Resolution: does the reply resolve the problem the customer stated?

INPUTS THE JUDGE SEES
customer message, reference policy, support reply

OUT OF SCOPE (neither rewarded nor punished)
length, tone, apology, greeting, formatting. Length is not quality.

LEVELS
4  Resolved. The reply addresses the stated problem and gives the
   remedy the policy provides. The problem is fixed, or the customer
   can fix it from this reply alone. No further message is needed.
3  Resolved with a gap. The right remedy for the stated problem,
   but it is promised and not yet done, a step the customer must
   take is missing, or the customer must write back first.
2  Addressed, unresolved. The reply engages the stated problem and
   gives no specific remedy for it: reassurance, a promise to look
   into it, a request for details the customer already gave.
1  Missed or wrong. The reply answers only a different problem, answers
   nothing, or says something the reference policy contradicts.

BOUNDARIES
4|3  If the remedy is promised but not yet done, or the customer
     must guess a step or send another message: 3.
3|2  If the policy's remedy is neither carried out nor offered as a
     specific action: 2.
2|1  If the reply never engages the problem the customer stated: 1.
Cap  A reply that contradicts the reference policy is level 1,
     whatever else it does well.

Step 4: turn the definitions into a worksheet

Turn each boundary into a question the judge answers from the text before it names a level. Chapter 26 says where such questions come from: “they are your class definitions turned interrogative”.

Q1  What problem did the customer state? One sentence, from the
    customer message.
Q2  Does the reply address that same problem? yes / no, with a quote.
Q3  Does the reply contain the remedy the policy gives for that
    problem, carried out or offered as a specific action? yes / no,
    with a quote.
Q4  Does anything in the reply contradict the policy? yes / no,
    with both passages quoted.
Q5  Is the problem fixed, or fixable by the customer, from this
    reply alone, with no further message to support? yes / no;
    if no, name what is missing.

LEVEL RULE, applied in this order
Q2 = no   -> 1
Q4 = yes  -> 1
Q3 = no   -> 2
Q5 = no   -> 3
otherwise -> 4

The level rule is the three boundaries and the cap, restated as a lookup. A judge that answers the five questions has almost no room left to improvise the grade.

Step 5: fix the output format

Fix a structured output whose fields appear in the order you want them written: answers, then level, then a short rationale, then a flag. Chapter 16 asks for “a label your code can parse”. The flag follows Chapter 26, which adds one for cases where the rubric answers pull in different directions.

stated_problem:              text
addresses_stated_problem:    yes | no | cannot_tell, plus quote
remedy_present:              yes | no | cannot_tell, plus quote
contradicts_policy:          yes | no | cannot_tell, plus quotes
complete_without_followup:   yes | no | cannot_tell, plus what is missing
level:                       1 | 2 | 3 | 4
rationale:                   at most two sentences citing the answers
needs_human:                 true | false

Set needs_human when any answer is cannot_tell, when the customer stated more than one problem, or when the policy does not cover the problem. One addition of my own: have your code recompute the level from the four answers and flag any item where the judge’s stated level differs.

Why does the worksheet come before the verdict?

The worksheet comes first because of how generation works, and because it leaves a record you can debug. Chapter 26 states the order in one sentence: “The examiner fills in the worksheet before writing the grade.” Text written before the level can shape it; text written after can only justify it.

The measured evidence for the ordering is thinner than the folklore. Chiang and Lee (Findings of EMNLP 2023) compared a score alone, a score followed by an explanation, and an analysis followed by a score, using one judge model on two datasets. They found that a rating with no explanation “is suboptimal” and that asking for one “consistently improves the correlation” with human ratings.

On the order itself they were candid: “We do not see rate-explain to be significantly better (or worse) than analyze-rate”. Stureborg and colleagues went further in their own setup and settled on “non-CoT prompting at a temperature of 0” as their final recipe. So one study found the order made no significant difference, and a second preferred no reasoning step.

I still put the worksheet first, and I want to be exact about why. The supported claim is that asking for reasons helps. The ordering is a design choice I make for the audit trail: a wrong level now points at a specific answer, and a specific answer points at a sentence of the rubric.

The book attaches its own warning. Chapter 26 says the worksheet “supplements, and cannot replace” measurement, because a tidy rationale can sit on a wrong label.

Does each sample answer land on exactly one level?

Each of three sample replies lands on exactly one level when the worksheet is applied by hand. I wrote a clear top, a clear bottom and a boundary case, then walked each through the five questions. A rubric that cannot place its own author’s examples is not ready for a judge.

The customer wrote: “I was charged twice for my March subscription. Please refund one of the charges.” The illustrative policy says a duplicate charge is refunded in full to the original payment method, the agent issues the refund, and it arrives in 5 to 7 business days.

Reply Q2 addresses Q3 remedy Q4 contradicts Q5 complete Level
A: “I see two identical charges on March 3. I’ve refunded the second in full to the card you paid with. It should arrive in 5 to 7 business days, and you don’t need to do anything else.” yes yes no yes 4
B: “Thanks for reaching out! To update your payment method, open Settings, choose Billing, and select Change card.” no (not reached) (not reached) (not reached) 1
C: “I’m sorry about the double charge. Duplicate charges qualify for a full refund to your original payment method. Let me know if you’d like me to go ahead.” yes yes no no 3

Reply C is the boundary case. A generous reader calls it a 4 because every statement in it is correct. A strict reader calls it a 2 because nothing has happened yet.

The anchors settle it. The refund is offered as a specific action, so the 3|2 boundary puts it at 3 or above. The customer must send another message to confirm a request already made, so the 4|3 boundary puts it at 3.

What does the template look like?

The template is the worked rubric with the specifics removed, so you can fill it for your own criterion. Keep the section order, since the output fields follow it.

JUDGE RUBRIC: <criterion name>

CRITERION (one sentence)
<the single question this judge answers>

INPUTS THE JUDGE SEES
<task input> / <reference, if one exists> / <output to grade>

OUT OF SCOPE (neither rewarded nor punished)
<length, tone, formatting, ...>. Length is not quality.

LEVELS (3 or 4; each a name, what it includes, what it excludes)
4  <name>. <descriptor>
3  <name>. <descriptor>
2  <name>. <descriptor>
1  <name>. <descriptor>

BOUNDARIES (one sentence per neighboring pair)
4|3  <the fact that decides it>
3|2  <the fact that decides it>
2|1  <the fact that decides it>
Cap  <one fact that overrides the rest, if any>

WORKSHEET (answer from the text, with a quote, before the level)
Q1  <restate the task or request>
Q2  <yes/no question taken from the 2|1 boundary>
Q3  <yes/no question taken from the 3|2 boundary>
Q4  <yes/no question taken from the cap>
Q5  <yes/no question taken from the 4|3 boundary>

LEVEL RULE (in order)
<answer pattern> -> <level>, one line per level

OUTPUT (fields in this order)
worksheet answers -> level -> rationale (max two sentences)
-> needs_human (true when an answer is cannot_tell or the scale
   does not fit the input)

CONTRAST PAIR (one per boundary, the nearest two examples)
<example at level n> / <example at level n+1> / <why they differ>

The last block matters more than its size suggests. Chapter 26 says of examples placed between neighbors that “the pair that separates a 2 from a 3 is worth more than any amount of scale-polishing”. For the worked rubric, reply C and a reply that only says “we’re looking into it” make the 3|2 pair.

How do you test the rubric for position and length bias?

You test the rubric with three small runs: a repeat baseline, an order swap, and a length-padding test. The book gives the swap rule and the correlation check and stops there. The procedure, the item counts and every threshold below are this post’s own and illustrative; no source I read states a pass mark.

  • Repeat baseline. Score the same 30 items twice with an identical prompt and settings. Record how many items change level. If more than 3 of 30 change, fix the rubric or the settings before reading any test below.
  • Order swap, pairwise use. Take 30 pairs of replies to the same customer message and judge each pair in both orders. Hold the rubric, the judge, the settings and both replies constant; change only which reply sits first.
  • Read the swap. Count the pairs whose winner is the same in both orders, and score the rest as ties. The rubric fails when fewer than 27 of 30 pairs survive. Report the first slot’s share of decisive verdicts beside it to show which way the flips lean; it has no pass mark of its own, because survival already bounds it.
  • Order swap, pointwise use. Score the same 30 items with the levels listed 4 to 1, then 1 to 4, and with any examples reordered. Hold the wording constant. The rubric fails when reordering moves more than 3 items beyond the repeat baseline.
  • Build padded twins. Take 30 replies that people labeled at levels 1 to 3. For each, append a restatement of its own sentences that roughly doubles the length and adds no fact, step or promise.
  • Check the twins by hand. A person confirms that all five worksheet answers are the same for each twin as for its original.
  • Read the padding. Hold the customer message, policy, rubric, judge and settings constant. An item fails when its twin scores a higher level; the rubric fails when more than 3 of 30 twins move up.
  • Check length on real traffic. On your labeled sample, compute the rank correlation of word count with the judge’s level, and the same correlation with the human level. A judge figure well above the human one is the warning.

The repeat baseline comes first for a reason given by the position-bias literature below: a flip means something only when it exceeds run-to-run noise. The pointwise swap is my extension by analogy, since a single reply has no slot to swap. None of the studies cited here tested level order.

Take care with the last item on this criterion. A complete reply is usually longer than an evasive one, so word count and level should correlate somewhat even for a fair judge. The human column is what tells real quality from padding.

For the arithmetic, the judge agreement calculator has a swap check that reports the survival rate and the first-slot share, and a verbosity check that reports rank correlation with an optional human column. Its kappa is unweighted. On a four-level scale it counts a 3-versus-4 miss the same as a 1-versus-4 miss, so it does not measure ordinal distance.

What does the evidence say about LLM as a judge position bias?

The evidence says position bias is large in some judges, varies in direction, and is strongest when the two answers are close in quality. Chapter 16’s rule is the standard response: “run every pairwise comparison both ways, accept only verdicts that survive the swap, score the rest a tie”.

Wang and colleagues (2023, arXiv preprint) showed that changing only the order let one system beat another on 66 of 80 test queries. Their conflict rate, the share of items whose verdict changes on a swap, was 46.3% and 5.0% for the stronger of two judges on two candidate pairs, and 82.5% and 52.5% for the weaker. The rate was “negatively correlated with the score gap”, which is why the test above uses pairs of similar quality. The study used 80 questions and two judges from 2023.

Zheng and colleagues (NeurIPS 2023 Datasets and Benchmarks) found their strongest judge consistent under a swap in 65.0% of cases with the default prompt, rising to 77.5% with few-shot examples. Their rule is the one the book adopts. The companion post that asks whether an LLM judge is reliable covers that study and its tie caveat in full.

Shi and colleagues (AACL-IJCNLP 2025) studied 15 judges on 22 tasks, over 150,000 evaluation instances by their count. From the abstract and introduction: position bias in capable judges “is not a result of random variations”, and there is “a high volatility in the direction of preference, even within the same LLM judge when applied to different tasks”. The authors call the work exploratory. I take two things from it: measure the first-slot share as well as survival, and measure repeat noise first.

What does the evidence say about length bias?

The evidence says some judges reward added words that carry no added content, and that the size of the effect depends heavily on the judge. Chapter 16 gives the check: “check your results for a correlation between word count and score; if you find one, your judge is grading by the pound”.

The padding recipe in the checklist comes from Zheng and colleagues. Their “repetitive list” attack rephrased an answer’s list items and put the rephrased copy ahead of the originals. On 23 answers, two judges were fooled 91.3% of the time and the strongest 8.7%.

Saito and colleagues (2023, arXiv preprint) report, from the abstract and introduction, that judges “exhibit a preference for longer answers in creative writing tasks”, with a gap between judge and human preferences. They hedge it to their own problem setting. That gap is the reason the checklist compares the judge’s length correlation with the human one.

At benchmark scale, Dubois and colleagues (COLM 2024) report, from the abstract, that controlling for length raised an automatic evaluator’s Spearman correlation with a human-preference leaderboard “from 0.94 to 0.98”. That result concerns a leaderboard correction and says nothing direct about rubric wording.

When should you merge levels or go binary?

Merge two levels when their boundary sentence cannot be written or cannot be applied, and report pass/fail when nobody acts differently on neighboring levels. The first rule is Chapter 26’s merge rule, quoted in step 3. The triggers below are my practical reading of it.

Three signals say merge. Two people labeling the same items split between the same two neighbors again and again. One level is almost empty on real traffic. Or the worksheet question behind a boundary keeps coming back cannot_tell.

Go binary when the verdict gates a release. A gate needs one line, and the worked rubric gives you a choice of two: pass at level 4, or pass at level 3 and above. Which line to draw depends on what a miss costs, so the scale informs the decision and a person makes it.

Keep the ordinal scale when degree is the thing the reader wants. A support lead triaging replies cares whether a reply missed the problem or merely left the refund waiting on a confirmation. A trend line of mean level across weeks carries that, with the definitions underneath.

For finer resolution, the glossary gives the book’s advice: “prefer a small ordinal scale with each level defined like a mini-class over a continuous score the model cannot actually resolve, and recover resolution by averaging discrete judgments”. Chapter 26 states the principle behind it: “Fine resolution earned by repeating a well-defined coarse judgment is honest; fine resolution claimed by a single draw on an undefined fine scale is decoration.”

How do you calibrate the rubric against human labels?

You calibrate by having people label a sample with the same rubric, then comparing the judge’s levels with theirs. The bias tests above only show that order and length are not doing the grading. They say nothing about whether the judge agrees with you.

Three neighboring posts own the parts. The method for measuring agreement, with Cohen’s kappa and a worked table, is in the reliability post linked above. The post on LLM as a judge vs human evaluation covers the step before that: two people labeling first, to find the ceiling a judge can reach.

The tutorial on how to build an LLM classifier covers keeping prompt examples apart from the labeled sets you measure on. And the guide to AI agent evaluation metrics places a judge score among the other numbers a team reports.

One point is specific to ordinal rubrics. Plain kappa treats every disagreement alike, so read the confusion table by neighbors: which adjacent levels swap, and in which direction. Chapter 26 gives that advice for classifiers, and each recurring swap names the boundary sentence to rewrite.

Where does this rubric method stop working?

This way of writing an LLM as a judge rubric stops working where the criterion has no reference to check against, where the scale does not fit the input, and where the evidence base is too thin to lean on. Three limits deserve a plain statement.

First, the worked rubric depends on a policy text. Without a reference, question Q3 asks the judge to know the right remedy unaided, and Chapter 16 warns that judges are much weaker at spotting an error alone than at checking against a known-good answer.

Second, some inputs do not sit on the scale: a customer with two problems, a reply to an abusive message, a question the policy never anticipated. The flag routes those to a person. Watch the flag rate, because a judge can look accurate by flagging everything hard.

Third, most of the studies above used judges from 2023 and 2024, on summaries and open-ended questions. None graded support replies against a policy. They tell you which failures to test for in your own eval set, and your judge’s numbers remain yours to measure.

The one thing to keep

In an LLM as a judge rubric, a level is worth keeping when a sentence defines it and a test could catch it being misapplied. Write the sentence between each pair of neighbors, make the judge answer the questions that sentence implies, and run the swap and the padding before you trust the scores. Where no sentence can be written, merge.

The marking-guide frame, the worksheet and the ordinal scale come from Chapter 26, “Agents as Classifiers and Scorers”. The rubric habits, the four biases and the swap rule come from Chapter 16, “Evaluating Agents”. Both chapters are in the full book; the glossary is free to read online. The agent evaluation guide collects the related posts and tools, and when the rubric has earned its place in your suite you can see the formats.

Questions readers ask

What scale should an LLM judge use?
The smallest scale whose levels you can define in writing. Use pass/fail when the criterion splits into yes-or-no checks, and three or four anchored levels when the degree of success is what the reader of the number needs. The evidence on numeric range is mixed, so the argument for a short anchored scale rests on definability: a level with a written boundary can be checked, and a bare number cannot.
Is a 1–5 scale better than a 1–10 scale for an LLM judge?
The studies disagree. Li et al. (2026) found a 0–5 range agreed best with human consensus and 0–10 was consistently the weakest, aggregated across six benchmarks. Stureborg et al. (2024) found a 1–10 score correlated better with human judgment than a 1–5 score on one summarization dataset. Neither study tested levels with written descriptors.
Should the judge explain before it gives a score?
Ask for the explanation, and put it first for the audit trail. Chiang and Lee (2023) found that a score with no explanation is suboptimal, and found no significant difference between explaining before the rating and explaining after it. Treat the written reasoning as evidence for debugging and never as a trace of the model's computation.
How do I test an LLM judge for position bias?
Judge every pair twice with the two answers swapped, count the share of pairs whose verdict survives, and score the rest as ties. Also report how often the first slot wins among decisive verdicts, because Shi et al. (2025) report that the direction of the preference is volatile, even within one judge across tasks. Measure run-to-run noise first so a flip is read against it.
Can one judge prompt score several criteria at once?
It can, and the scores contaminate each other. Stureborg et al. (2024) found a judge's scores for two attributes in one output correlated at r = 0.979, against r = 0.315 for human raters on the same attributes. One criterion per call keeps the scores separable; batching is a cost decision to take knowingly.

Sources

  1. Weiyue Li, Minda Zhao, Weixuan Dong, et al. (2026). Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale (arXiv:2601.03444)
  2. Rickard Stureborg, Dimitris Alikaniotis, Yoshi Suhara (2024). Large Language Models are Inconsistent and Biased Evaluators (arXiv:2405.01724)
  3. Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, Chenguang Zhu (2023). G-Eval (arXiv:2303.16634)
  4. Bhuvanashree Murugadoss, Christian Poelitz, Ian Drosos, et al. (2024). Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions (arXiv:2408.08781)
  5. Cheng-Han Chiang, Hung-yi Lee (2023). A Closer Look into Automatic Evaluation Using Large Language Models (Findings of EMNLP 2023)
  6. Peiyi Wang, Lei Li, Liang Chen, et al. (2023). Large Language Models are not Fair Evaluators (arXiv:2305.17926)
  7. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023 Datasets and Benchmarks)
  8. Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, Soroush Vosoughi (2025). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge (AACL-IJCNLP 2025)
  9. Tuhina Tripathi, Manya Wadhwa, Greg Durrett, Scott Niekum (2025). Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation (COLM 2025)
  10. Keita Saito, Akifumi Wachi, Koki Wataoka, Youhei Akimoto (2023). Verbosity Bias in Preference Labeling by Large Language Models (arXiv:2310.10076)
  11. Yann Dubois, Balázs Galambosi, Percy Liang, Tatsunori B. Hashimoto (2024). Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators (COLM 2024)
  12. Hamel Husain, Shreya Shankar (2025). Why do you recommend binary (pass/fail) evaluations instead of 1-5 ratings (Likert scales)?
  13. nimitkalra (Hacker News handle) (2025). Hacker News comment on discrete scales for judges
  14. chaholl (GitHub handle) (2026). GitHub issue on scoring five dimensions in one judge call (AltairaLabs/PromptKit #1883)