Home / Blog / Patterns and multi-agent systems / The Evaluator Optimizer Pattern: A Second Model…

Patterns and multi-agent systems

The Evaluator Optimizer Pattern: A Second Model Grades the First

The evaluator optimizer pattern loops generate, check and revise. See which evaluator adds information, when to stop, and test it against plain retries.

By Enrique Gutiérrez · Published · 23 min read

The evaluator optimizer pattern is a loop. One model call generates a candidate, a separate evaluator grades it against written criteria, and the failures go back for revision until the check passes or a cap fires. The evaluator is often described as a second model. The book prefers a check that runs, with a model as the fallback.

The title’s “a second model grades the first” is the common description, and it matches the 2024 essay that named the pattern. Chapter 10, Workflows and Composition Patterns (in the full book) seats a verifier first: “the test suite, the compiler, the schema validator, the query that executes or fails to”. A model critic takes the seat only when no such check exists.

After this post you can say what your evaluator knows that your generator does not, write the loop with a stop rule that depends on nobody’s opinion, and compute whether the rounds beat simply trying again. The pseudocode, the evaluator-types table, the five-question test, the contract and all the arithmetic are mine. The loop, the ladder of evaluators, the convergence conditions and the failure shapes are the book’s.

What is the evaluator optimizer pattern, in the book’s words?

The book’s glossary, free to read online, defines the evaluator–optimizer as a workflow in which one call generates a candidate, an evaluator “(ideally a real verifier—a test suite, a schema check)” grades it, and the failures feed back for revision, “looping until the check passes or a cap fires”.

Its name comes from a 2024 engineering essay by Erik Schluntz and Barry Zhang. Its definition is one sentence: “one LLM call generates a response while another provides evaluation and feedback in a loop.” The essay presents no measurements, no cap and no cost, and both of its examples seat a model as the evaluator. Generator–critic and propose-and-check are other names for the same shape.

The book’s design rule is separation: “The generator proposes; something that is not the generator disposes”. Its warning about the alternative is blunt: “Collapse them into one prompt that both writes and approves, and you have built a rubber stamp with extra steps.” The post on agentic workflow patterns surveys the three shapes that sit beside this one.

Where does a verifier come first?

A verifier comes first wherever the work can be exercised by something outside the model. Chapter 4 defines a verifier as an external check that runs the work, and it states the rule the rest of the book points back to: “A real verifier beats the model grading itself.”

Chapter 10 applies that rule to the evaluator optimizer pattern in six words: “Wherever a verifier exists, seat it”. It adds an instruction about the feedback: “pass the failure through whole and unedited, because the stack trace is better feedback than any summary of it”.

Some work has no verifier: tone, faithfulness to a source, the register of a translation. There the book hands the seat to a model critic: “a separate call, a separate prompt, ideally a different model, judging against an explicit rubric”. It prices that choice in one clause: “Its verdict is an estimate where the verifier’s is a fact”.

How do you write the loop with a stop rule?

Write the loop with three exits (pass, no progress, cap) and return the best candidate seen whenever it leaves without a pass. The book describes the exits in prose and gives no code. The pseudocode is my rendering, in no particular language.

function refine(task, criteria, CAP, K):
    best = none          # best candidate so far, by score
    stale = 0            # rounds with no higher score
    seen = empty set     # fingerprints of candidates
    feedback = none      # the last draft and its failures

    for round in 1..CAP:
        candidate = generate(task, feedback)
        verdict = evaluate(task, candidate, criteria)
            # never sees the generator's conversation
        log(round, fingerprint(candidate), verdict)

        if verdict.pass:                      # every criterion PASS
            return candidate, exit = "pass", rounds = round

        if best is none or verdict.score > best.score:
            best = (candidate, verdict.score)
            stale = 0
        else:
            stale = stale + 1

        if fingerprint(candidate) in seen or stale >= K:
            return best.candidate, exit = "no_progress", rounds = round

        add fingerprint(candidate) to seen
        feedback = (candidate, verdict.failures)   # failures whole and unedited

    return best.candidate, exit = "cap", rounds = CAP

The score is the count of criteria that passed, whatever the evaluator is: tests green, schema fields valid, rubric lines marked PASS. Only the first exit means accepted. The other two hand back the best attempt with a flag, and the caller decides whether a person sees it or the task fails.

Why three exits and keep-best?

Three exits are needed because a pass may never come, and a loop whose only exit is a verdict has no guaranteed end. The book starts with two: “The loop ends at one of two doors: the evaluator passes the candidate, or a cap says enough.”

Its caps, in its own words: “A hard iteration ceiling, small (three to five rounds is the common figure, offered as illustration), plus a no-progress exit that stops early when the verdict or the candidate plateaus”. The “three to five” is the book’s illustration, labeled as one. The book also keeps the best candidate by the evaluator’s score, so a late round that polishes substance away cannot become the output.

The pseudocode makes two choices the book leaves open: a plateau is K rounds with no higher score, and a cycle is a repeated fingerprint.

One exit is missing on purpose: looping until the reviewer finds nothing. A Hacker News commenter put the problem well in 2026: “if you ask the agent to find things to fix, it’ll find things to fix.” A stop condition that waits for a model to run out of complaints may never fire, or fires on a critic that approves everything.

What must the evaluator know that the generator does not?

The evaluator must hold at least one piece of information the generator lacked when it wrote the candidate. Otherwise the second call repeats the first one’s beliefs. A Hacker News commenter asked exactly this in 2026: “what could the reviewer agent add of value to the agent writing the code?”

Chapter 4 gives the mechanism for the empty case. A model rereading its answer uses the weights that wrote it: “No new information enters; the second pass is another draw from the same lottery over the same desk”.

Which evaluator types add which information?

Each evaluator type adds a different kind of information and fails in its own way, and the table ranks them from most information added to least. The ladder behind it is the book’s: verifier, then independent critic, then self-review, across Chapters 4, 10 and 16. The rows, columns and wording are mine.

Read from the top and stop at the first row you can staff. If you have both a test and a rubric, the test is your evaluator and the rubric covers only what the test cannot examine.

Evaluator type Information it adds How it fails Row you reach when you
Executed check: tests, compiler, schema validator, a query that runs A measurement of the candidate, taken outside the model Passes whatever it does not examine; a thin suite accepts wrong work; tests the model wrote itself can be wrong in both directions have a test
Reference comparison for this input: a match in code, or a judge handed the known-good answer The answer itself Exists only where the input has a reference; a model doing the comparison brings judge error back in have a reference
Tool-grounded critic: a model that may run code or look facts up before it rules The tool results it chose to fetch The critic picks what to check; a claim it never checked can pass as verified have only a rubric
Rubric judge, different model, fresh context Written criteria, a clean desk, different weights Leniency and verbosity bias; its errors overlap with the generator’s have only a rubric
Rubric judge, same model, fresh context Written criteria and a clean desk Self-preference; it holds the generator’s beliefs about the facts have only a rubric
Self-review: the same model in the same conversation None beyond a reworded question Signs off its own work; on reasoning tasks it can revise right answers to wrong have nothing

Two notes keep the table honest. A reference set of labeled examples is a measuring instrument for the loop, used offline; the reference row applies only when the input being processed has its own known-good answer. And a tool-grounded critic applies only when the criteria include claims a tool can check.

Self-review run several times is a field recipe the book reports as the Rule of Five, and the book rules on using it alone: “Five reviews and then the test suite is a discipline. Five reviews instead of the test suite is a séance.”

What does the evidence say, sorted by feedback source?

Most of the studies line up once they are sorted by where the feedback came from. Refinement helps when the verdict carries information from outside the model, and it has lowered accuracy on reasoning benchmarks when the model judged itself. All are benchmark studies, mostly on models from 2023 and 2024. I found no production measurement of rounds-to-convergence for the evaluator optimizer pattern.

Where I read only a paper’s abstract, the entry says so. Chapter 4, in the full book, makes the same sort in prose.

What happens when the model judges itself?

When the same model supplies the feedback, the record splits by task: gains on quality and style tasks, losses on reasoning tasks. The favorable result is Madaan and colleagues (2023), the Self-Refine paper, in which one model acts as generator, feedback provider and refiner.

Its abstract reports “improving by ~20% absolute on average in task performance” across seven tasks. The full text caps the loop at four iterations and says “the marginal improvement naturally decreases with more iterations”. On math reasoning the paper calls its gains modest: one model’s feedback said everything looked good in 94% of instances. Gains there are “much bigger (5%+) if an external source can identify if the current math answer is incorrect”.

Huang and colleagues (2023, published at ICLR 2024) tested self-correction with no outside signal on reasoning benchmarks. They count one call for the first answer, three after one round and five after two. For one model, accuracy went 82.0 → 79.5 → 80.0 on CommonSenseQA and 49.0 → 49.0 → 43.0 on HotpotQA. For an older model the CommonSenseQA figure went 75.8 → 38.1 → 41.8.

The paper’s summary is “after self-correction, the accuracies of all models drop across all benchmarks.” It attributes earlier reported gains to stopping rules that used ground-truth labels, gains that “vanish when oracle labels are not available.” Its limits are reasoning benchmarks and 2023 models. It also asks that self-correction be compared with baselines using the same number of responses, which is the retry baseline below.

What happens with a sound verifier, with and without feedback?

With a sound external verifier, accuracy rose sharply in one study, and much of that rise appeared even when the loop fed back no critique and simply sampled again. The study is Stechly, Valmeekam and Kambhampati (2024, an arXiv preprint), on four test sets in three formal domains where a checker in code decides correctness.

Test set (100 instances each) Standard prompt Model verifies itself Sound verifier, pass/fail only Sound verifier, first error fed back No feedback, 15 samples No feedback, 25 samples
Game of 24 5% 3% 36% 38% 28% 42%
Graph coloring 16% 2% 38% 37% 40% 44%
Blocksworld 40% 55% 60% 87% 68% 72%
Mystery Blocksworld 4% 0% 10% 8% 9% 14%

The numbers are from the paper’s Table 1. The model as its own verifier lowered accuracy on three of the four, and the authors explain why: “even when the LLM generates a valid solution, the verifier LLM rejects it often enough that overall performance suffers”. Blocksworld is the exception, where self-verification beat the standard prompt.

With the sound verifier, accuracy rose on all four. The two right-hand columns use the checker only to pick a passing sample, and they land close to the feedback columns on three. The authors conclude that “merely re-prompting with a sound verifier maintains most of the benefits of more involved setups.” Blocksworld again differs: feeding back the first error reached 87% against 72% for 25 samples.

One model was tested, each cell is 100 instances, and the arms use different numbers of calls. A sound checker also has to exist, which is true of puzzles and rarely of prose.

One abstract-only read points the same way: Tyen and colleagues (Findings of ACL 2024) report that models “generally struggle” to find reasoning mistakes and improve when handed the error’s location.

What do tests plus an explanation buy, and what does a human add?

When executed tests say that code failed and a model explains why, repair works less often than when a person writes the explanation. That is Olausson and colleagues (2023, published at ICLR 2024), who studied self-repair on Python programming tasks with unit tests.

Their Table 1 compares repair success with the model’s own explanation and with a human’s. Overall it is 33.30% against 52.60%. By difficulty it is 42.64% against 62.21% on introductory problems, 19.33% against 45.67% on interview problems and 3.67% against 14.67% on competition problems.

The paper counts cost in tokens sampled and concludes that “when the cost of carrying out repair is taken into account, performance gains are often modest, vary a lot between subsets of the data, and are sometimes not present at all.” Its stated limits include self-contained Python tasks with executable tests.

Shinn and colleagues (2023), the Reflexion paper, used unit tests the model wrote for itself. The abstract reports 91% pass@1 on the HumanEval coding benchmark against the 80% it cites for the best earlier result, a dated snapshot. The full text names the weakness: with a flaky suite “it is possible that all tests pass on an incorrect solution and lead to a false positive label”.

What conclusion does that evidence support?

The evidence supports one narrow conclusion: the loop’s value tracks the quality of the information in the verdict. Kamoi and colleagues (TACL, 2024) surveyed the field by feedback source and put it in a clause: “the bottleneck is in the feedback generation”.

Their first finding is that “no prior work demonstrates successful self-correction with feedback from prompted LLMs, except for studies in tasks that are exceptionally suited for self-correction”. Their second is that “self-correction works well in tasks that can use reliable external feedback”. The survey covers work to mid-2024.

The studies also disagree: Madaan reports a large average gain from self-feedback, and Huang traces the gain on one of Madaan’s seven tasks to a weak first prompt. None of them measured a loop with a calibrated rubric judge on open-ended work, which is the case many teams are in. The book’s oracle entry states the rule that covers it: “A loop is exactly as trustworthy as its oracle”.

Should the critic be the same model or a different one?

A fresh context for the critic is supported by measurement, and a different model is supported more weakly: it reduces one bias and leaves the critic’s errors correlated with the generator’s. I call that leftover overlap a shared blind spot. The term is mine; the book speaks of a loop steering “toward the critic’s blind spots”.

The question gets argued by anecdote: in one 2026 Hacker News thread, advice to use a different reviewer model drew the reply “There’s no evidence of this”. Three papers, each read in abstract only, are the measured part.

Khullar, Hopkins, Wang and Roger (2026, a preprint) report that model monitors “fail to report high-risk or low-correctness actions more often when evaluation follows a previous assistant turn in which the action was generated, compared to when the same action is evaluated in a new context presented in a user turn”. The abstract covers four coding and tool-use datasets. It compares contexts and makes no comparison between models.

Panickssery, Bowman and Feng (2024, a preprint) report “a linear correlation between self-recognition capability and the strength of self-preference bias”. Kim, Garg, Peng and Garg (ICML 2025) report that on one leaderboard dataset “models agree 60% of the time when both models err”, and that “larger and more accurate models have highly correlated errors, even with distinct architectures and providers”.

What is still unmeasured?

I found no controlled comparison, inside a refinement loop, of a same-model critic in a fresh context against a different-model critic, scored on the final accepted output. Chapter 10 advises “ideally a different model” and cites no study for it; Chapter 16 grounds the advice in self-preference, one of the four judge biases it lists.

So the order I would act on is short. Never evaluate in the generator’s conversation. Use a different model where the stakes justify a second bill, and treat neither choice as independence. Calibrate whichever critic you pick against human labels; the post on whether LLM-as-a-judge is reliable walks that method through.

Does the loop beat simply trying again?

The loop beats plain retries only if feedback makes the next attempt more likely to pass than a fresh draw would be, and that is a measurable quantity. The baseline keeps the same check, discards the draft and its failure message, and samples again up to the same cap. Chapter 10 has a phrase for it: “a retry behind a gate is simply a fresh draw from the same lottery”.

The arithmetic uses illustrative inputs. Suppose a candidate passes a sound check 40% of the time and attempts are independent. Three attempts then produce at least one pass with probability 1 − 0.6³ = 78.4%. That is pass@3, the figure three rounds of the evaluator optimizer pattern must beat.

With JavaScript on, the pass@k and passk calculator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the pass@k and passk calculator on its own page to share a result by link.

The calculator opens on those inputs and shows 78.4% for pass@3. It also shows 6.4% for pass3, the chance that all three attempts pass, and a 72-point envelope between the two. Only pass@3 matters here, because a check picks the winner.

Now give the loop a repair rate, r: the chance that a failed candidate passes after one round of feedback. A three-round loop then accepts 1 − 0.6 × (1 − r)². At r = 0.4 the loop equals the retries at 78.4%; at r = 0.5 it reaches 85.0%, and at r = 0.3 it falls to 70.6%. A repair rate below the fresh-draw rate is possible, because the revision starts from a flawed draft.

The book’s Figure 4.3 shows repeated sampling with the answers voting. The retry baseline takes the same repeated draws and lets a check choose. Both lean on the draws being independent, and real attempts at one task are correlated: an input that fails once tends to fail again. Expect measured retries to score below the formula, and run both arms on your own eval set.

Self-consistency.
Figure 4.3 Self-consistency. The same hard question is sampled several times, and because the paths vary, so can their answers. Right routes tend to converge on the one correct answer while wrong routes scatter, each wrong in its own way, so a majority vote recovers the answer (in accent) that most paths agree on—evidence gathered without anyone knowing the right answer in advance. Reuse this diagram

What does an accepted output cost?

An accepted output costs the expected calls per task divided by the share of tasks the loop accepts. For the evaluator optimizer pattern the book gives the count: “every round is at least one generate and one evaluate, so an N-round loop costs roughly N times the calls and N times the latency of a single shot”.

With the illustrative inputs above (40% first-round pass, repair rate 0.5, cap of three), the loop stops after one round 40% of the time, after two 30% and after three 30%. Expected rounds are 1.90, and 85.0% of tasks end accepted. A check in code makes each round one model call, so the cost is 1.90 ÷ 0.85 = 2.24 calls per accepted output.

A model critic doubles the calls per round: 3.80 ÷ 0.85 = 4.47 calls per accepted output. Plain retries at 40% cost 1.96 expected attempts for 78.4% accepted, which is 2.50 generate calls per accepted output, or 5.00 with a model critic behind each attempt. Calls are a floor on cost, since each feedback round also lengthens the prompt.

What do the evaluator’s own errors do to the accepted set?

A model critic’s two error rates decide how much of the accepted set is good, and neither appears on the bill. This arithmetic is mine and illustrative; the inputs are printed so you can recompute it.

Take a generator whose first draft is good 60% of the time, and a critic that passes 90% of good drafts and 30% of bad ones. Round one accepts 0.60 × 0.90 + 0.40 × 0.30 = 66% of drafts, and 0.54 ÷ 0.66 = 81.8% of those are good. The cost is 2 ÷ 0.66 = 3.03 calls per accepted output. A critic that passes everything leaves the accepted set at 60% good for twice the calls.

Now look at what gets rejected. With the same critic, 34% of drafts are sent back and 17.6% of those were good (0.06 ÷ 0.34). Raise the generator to 95% good and the picture flips: 13% are sent back, and 73.1% of them were good (0.095 ÷ 0.13). Each of those enters a revision round that can make it worse, which is the rejection of valid work Stechly’s team described.

How do the book’s failure shapes show up in a trace?

The book names four failure shapes for the evaluator optimizer pattern, and each leaves a distinct mark in the per-round log the pseudocode writes. The names and descriptions are Chapter 10’s. The trace signatures and first responses are mine.

Failure shape (the book’s) What it is, per Chapter 10 What the per-round log shows First response
Rubber stamp Every candidate passes on round one, because the evaluator is too kind or too vague Exit “pass” at round 1 on nearly every task; known-bad candidates also pass Feed the evaluator labeled bad examples and measure its false-accept rate
Oscillation Round two’s fix reintroduces round one’s flaw A fingerprint repeats; the failing criterion alternates between two names Exit on no progress; give the generator both failures at once
Over-edit Late rounds polish substance away The score rises, then dips Keep the best candidate; lower the cap
Wrong hill The loop climbs toward a weak critic’s blind spot The loop’s score rises while a held-out check stays flat Calibrate the critic; add a check the loop never sees

The book lists the signals in one sentence: “a verdict that stops improving, a candidate that stops changing between rounds or cycles back to an earlier version, a score that dips after rising”.

Its planning figure is that “the first round of feedback typically buys most of the improvement, the second buys some, and beyond that the curve flattens fast”.

When does the book say the loop converges?

The book says the loop converges when three conditions hold at once: “the feedback is specific enough to act on”, “the flaw is of the kind revision actually repairs”, and “the target holds still between rounds”. These are the book’s conditions, stated with no measured curve behind them.

The second condition is where the evidence bites: revision repairs rendering, coverage and style, plus any fault something else has located. Chapter 4’s division of labor covers the rest: “let the world find; let the model fix”.

What goes in an evaluator contract?

An evaluator contract records the criteria, what the evaluator is given, the verdict format, the calibration record and the stop rule, so another engineer can audit the loop without reading its code. The template is mine. Its habits (binary criteria, a pass and a fail example, evidence before the verdict, a reference where one exists) come from Chapter 16’s treatment of LLM-as-a-judge.

The third verdict, UNKNOWN, is a practitioner’s idea from a 2026 Hacker News comment: “Each acceptance criteria gets a PASS/FAIL/UNKNOWN verdict, attached with evidence.” Its effect is unmeasured; I keep it because it stops an unverifiable criterion from being folded into PASS.

EVALUATOR CONTRACT

Task:
Evaluator type (one row of the table):
What it knows that the generator does not:

Criteria (each binary, each with examples)
  C1:                  pass example:          fail example:
  C2:                  pass example:          fail example:
What the criteria do not examine:

Evaluator receives: task input, candidate, criteria, reference (if any)
Evaluator never receives: the generator's conversation

Verdict
  per criterion: evidence, then PASS / FAIL / UNKNOWN
  pass: every criterion PASS (UNKNOWN counts as FAIL)
  score: number of criteria marked PASS
  failures: returned to the generator whole, with its last draft

Calibration (labeled examples, n =        date:        )
  false-accept rate (bad passed):
  false-reject rate (good failed):

Stop rule
  round cap (CAP):
  no-progress window (K):
  cost or time ceiling:
  exit without a pass: best by score, flagged, sent to:

Logged per round: fingerprint, verdicts, score, calls, exit

Baseline (same check, same cap)
  plain retries, accepted share:
  loop, accepted share:          calls per accepted output:

The line I would refuse to leave blank is the one about what the criteria leave unexamined. The book’s fine print on every success of this pattern is that “the loop optimizes exactly what the evaluator measures, and nothing else”.

Should you use the pattern for this task?

Use the evaluator optimizer pattern when all five questions below get a yes, and treat each no as a specific instruction. The first and third come from Chapter 10, the rest from the studies above. The test is mine.

  • Criteria. Can you write the criteria as binary checks, each with a pass and a fail example? If no, stop: the book says “there is no evaluator to build”.
  • Separation. Does something other than the generator’s own conversation produce the verdict (any row of the table except the last)? If no, do not gate on the loop.
  • Repairable flaw. Is the typical flaw one that revision repairs once it is located: a failing case, a missing field, a named criterion? If no, use plain retries behind the check.
  • Measured evaluator. Do you know the evaluator’s false-accept and false-reject rates on labeled examples, and are most of its rejections true? If no, calibrate first.
  • Baseline. On your own eval set, does the loop accept more, or cost fewer calls per accepted output, than plain retries at the same cap? If no, keep the retries.

Three cases show the table, the test and the contract agreeing. Code generation with a test suite takes the top row, stops on green tests, cap or no progress, and keeps the loop once it beats retries. A marketing-copy rewrite judged for tone takes the different-model rubric judge with keep-best, and waits at question four until the judge is calibrated. The third case gets the full walk.

Worked example: extraction with a schema and a reference set

An extraction task with a schema and a labeled reference set takes the executed-check row, and the reference set becomes the loop’s calibration record. The numbers are illustrative. The task is to pull twelve fields from each incoming invoice into a typed record, and the team holds 200 invoices with hand-checked answers.

The schema validator is an executed check, so it sits in the evaluator’s seat. A new invoice arrives with no known-good answer, so the reference set goes into the contract under calibration.

Then the stop rule. Pass means every field validates; the cap is three, K is one, and the validator’s error list goes back whole with the draft. On the 200 labeled invoices, 140 validate on round one (70%). The loop repairs 48 of the 60 failures within the cap, so 188 are accepted (94%), and 12 exit flagged for a person.

Then the test. Criteria, separation and repairable flaw all get a yes, since a validator error names the field. On the baseline, plain retries on the same 60 failures recover 39 (65%) against the loop’s 48 (80%). The independence formula predicted 1 − 0.3² = 91%, so this invented retry arm sits well below it, where correlated failures would put it.

Question four is where the reference set pays for itself. Of the 188 accepted records, 168 match the hand-checked answers and 20 are schema-valid with a wrong value: 14 from round one and 6 from repaired records. The accepted set is 89.4% correct, and the validator’s false accepts are 10.6% of it.

So the answer is yes, with a boundary. The loop raised schema-valid output from 70% to 94% and left every wrong value where it was. The next move is to turn the 20 misses into new executed checks, such as line items summing to the total.

Where does this advice stop applying?

This advice stops applying where its evidence does: on open-ended work judged by a rubric, on newer models, and in systems larger than one generate-and-check pair.

First, the studies are benchmark studies on models from 2023 and 2024, several in formal domains with a perfect checker. Second, every figure I computed assumes independence or fixed rates, and real loops have neither. Third, I found no measured optimum for the cap, so every cap, the book’s included, is a planning figure.

Fourth, the evaluator optimizer pattern is a workflow shape. A reviewer agent inside a team of agents adds handoffs and coordination cost on top of everything here, which the post on multi-agent system overhead covers. The AI agent design patterns cheat sheet puts the neighboring shapes on one page, each with its cost and its typical failure.

For the wider set of shapes, see the agent patterns guide; the judge agreement calculator helps with the calibration lines of the contract.

The check to keep

The evaluator optimizer pattern is worth its calls when the verdict carries information the generator lacked. The cheapest way to find out is to run it against plain retries behind the same check.

Chapter 16 supplies the rule for the evaluator’s seat: “the best judge is the one you did not need”.

Chapters 1 and 2 and the glossary are free to read online. The loop, its convergence conditions and its failure shapes are in Chapter 10, “Workflows and Composition Patterns”. The evidence and the judge are in Chapters 4 and 16. All three are in the full book; see the formats.

Questions readers ask

What is the evaluator optimizer pattern?
It is a workflow in which one model call produces a candidate, a separate evaluator grades it against written criteria, and the failures go back to the generator for revision. The loop ends when the check passes or a cap fires. The 2024 essay that named it describes two model calls; the book prefers a verifier, such as a test suite or a schema check, in the evaluator's seat wherever one exists.
Should the evaluator be a different model from the generator?
The measured part is narrower than the advice. One 2026 preprint found model monitors more lenient when they judged an action in the conversation that produced it than in a fresh context. A 2025 paper found that different models' errors overlap. So a fresh context is supported, a different model reduces self-preference, and neither makes the critic independent. A check that runs beats both.
How many iterations should an evaluator-optimizer loop run?
No measured optimum for production loops was found for this post. One 2023 paper capped its loop at four iterations and reported shrinking gains with each one. The book offers three to five rounds as an illustration and tells you to plan for two or three productive ones. Set a small cap, stop early when the score stops improving, and keep the best candidate.
Is the evaluator optimizer pattern the same as reflection?
They share a shape and differ in who judges. Reflection usually means the model critiques its own output. This pattern requires the evaluator to be separate from the generator, and it works best when the evaluator runs something. The book places self-review on the bottom rung of its ladder and calls it the weakest verifier you can buy.
When should I skip the pattern?
Skip it when you cannot write the criteria down, when the only available evaluator is the generator rereading its own conversation, when the typical flaw is one the model would have to discover unaided, or when the same number of plain retries behind the same check scores as well on your own eval set.

Sources

  1. Erik Schluntz, Barry Zhang (2024). Building effective agents
  2. Aman Madaan and colleagues (2023). Self-Refine: Iterative Refinement with Self-Feedback
  3. Jie Huang and colleagues (2023). Large Language Models Cannot Self-Correct Reasoning Yet (ICLR 2024)
  4. Kaya Stechly, Karthik Valmeekam, Subbarao Kambhampati (2024). On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks (preprint)
  5. Theo X. Olausson and colleagues (2023). Is Self-Repair a Silver Bullet for Code Generation? (ICLR 2024)
  6. Noah Shinn and colleagues (2023). Reflexion: Language Agents with Verbal Reinforcement Learning
  7. Ryo Kamoi and colleagues (2024). When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs (TACL 12: 1417–1440)
  8. Gladys Tyen and colleagues (2023). LLMs cannot find reasoning errors, but can correct them given the error location (Findings of ACL 2024; abstract read)
  9. Arjun Panickssery, Samuel R. Bowman, Shi Feng (2024). LLM Evaluators Recognize and Favor Their Own Generations (preprint; abstract read)
  10. Elliot Kim, Avi Garg, Kenny Peng, Nikhil Garg (2025). Correlated Errors in Large Language Models (ICML 2025; abstract read)
  11. Dipika Khullar, Jack Hopkins, Rowan Wang, Fabien Roger (2026). Self-Attribution Bias: When AI Monitors Go Easy on Themselves (preprint; abstract read)
  12. sonnig (Hacker News) (2026). Hacker News comment asking what a same-model reviewer adds
  13. ModernMech (Hacker News) (2026). Hacker News comment asking whether a review loop ever terminates
  14. sdevonoes (Hacker News) (2026). Hacker News comment disputing the different-model advice
  15. LeoStehlik (Hacker News) (2026). Hacker News comment describing PASS/FAIL/UNKNOWN verdicts with evidence