Home / Tools / Agent verifiability scorecard

Free tool · runs in your browser · from Chapter 1

Agent verifiability scorecard

Score an AI agent on twelve questions, one per Part of the book: how would you know it worked? Get a total out of 24, a radar by Part, and the chapters to open.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The twelve questions and their quotes are the book's. The 2/1/0 scoring, the three bands, the radar and the order of the gaps are the tool's own.

What is an agent verifiability scorecard?

An agent verifiability scorecard is a twelve-question check of one thing: for each part of your system, can you say how you would know it worked? It is the book’s first design question asked twelve times. Chapter 1 puts it plainly: “The first question to ask of any agent design is the one this book will ask over and over: what signal tells you it worked?” Each question here is that question asked of one Part of the book, in the words of the chapter that answers it.

The reason to ask it so often is that the model cannot answer it for you. “The model’s own confidence carries no information about whether it is right,” Chapter 1 says, “so trust must come from outside—a passing test, a validating schema, a source you can open, a human who approves.” A system that is impressive in a demo and has no outside signal is still a demo. The scorecard measures how much of yours has one.

The book’s compass.
Figure 1.1 The book’s compass. Four bearings decide the design forks in almost every later chapter, and the needle rests on the first of them—verifiability—because an agent is only as trustworthy as the signal you can use to check it. The other three, the scarcity of context, compounding error, and the simplest thing that works, are read alongside it, not instead of it. Reuse this diagram

How is the score calculated?

Each question is answered yes, partly or no, scored 2, 1 and 0, for a total out of 24. The tool shows the total as you go, and gives a band once all twelve are answered: 0 to 8 is a demo, 9 to 16 a pilot, 17 to 24 a product. The scoring and the bands are the tool’s own; the book does not grade systems. The radar spreads the same answers across the seven Parts of the book, so a team that scores well on evaluation but has no stopping condition can see the dent.

“Partly” exists because most real answers are partial. A team may trace every run but record only the template, not the rendered prompt; may have an eval set but run each task once; may have gates, but on every action rather than on the irreversible ones. Score “partly” when some of the clauses are true, and “yes” only when all of them are.

The band text leans on a line the book quotes in Chapter 25: “Demo is works.any(), product is works.all().” The chapter’s gloss is the whole argument in one sentence: “A demo needs one run to succeed; a product needs every run to.” The questions are a list of the places where a single good run tells you nothing about the next one.

What do the twelve questions cover?

The twelve questions follow the book’s Parts in order, one or two per Part, from the signal that defines success to the rollout that protects users from a bad change.

# Part The question, in short Chapter
1 I Can you name the outside signal that a run worked? 1
2 I Is exact work routed to something exact? 2
3 II A checkable exit, hard budgets, and a loud failure? 3
4 II Tool errors the model can act on, and a real verifier? 5, 4
5 III Do you know what was on the desk, and does all of it earn its keep? 7
6 IV Gates by consequence tier, in code, sending decisions? 12
7 IV For unattended runs: goal function, iteration cap, kill switch? 13
8 IV The lowest rung that works, and evidence for each climb? 14, 1
9 V A full trace of every run, and someone reading the worst? 15
10 V Evals from real failures, repeated, with a calibrated judge and a gate? 16
11 V The trifecta audit, the three dials, idempotency and receipts? 17, 18
12 VI, VII Cost and latency on a fixed bank, an automatic rollback, and the cost of checking? 16, 20, 23

Two of the questions are about the model’s limits rather than your plumbing. Question 2 comes from Chapter 2’s list of what wobbles: “Multi-digit arithmetic, precise counting, long chains of exact deduction.” The remedy it prescribes is “architectural and cheap: route exact work to something exact—a calculator, a database, executed code—and let the model orchestrate rather than compute.” Question 4 carries Chapter 4’s short rule: “A real verifier beats the model grading itself.”

The rest are about the harness, which is where most of the verification lives. Question 3 asks for an exit a program can check (“The exits you can trust most are the ones a program can check”) and budgets that fail loudly: “Save the full transcript, surface whatever partial work exists, and say plainly that the run was stopped rather than finished.” Question 9 asks for a trace of every run and the habit that makes it useful: “read traces, on a schedule, with human eyes.”

Why does each question point to a chapter?

Each question points to a chapter because the scorecard is an index, not a course. A “no” is not a verdict on the team; it is the address of the chapter that explains what to build. The result lists every “no” first and every “partly” after it, in the book’s order, each with the chapter to open. Chapters 1 and 2 are free to read on this site; the others open on their summaries.

The order matters because the questions depend on each other. Without question 1, the signal, there is nothing for question 10’s eval set to check and nothing for question 12’s canary to roll back on. Without question 9’s traces, you cannot grow question 10’s eval set from real failures, which is where Chapter 16 says the tasks come from: “Where do the tasks come from? From watching the agent fail.” When several answers are “no”, start from the top.

How do you score an unattended or high-stakes agent?

You score an unattended or high-stakes agent the same way, but with less room for “partly”, because nobody is watching to catch what the signals miss. Question 7 applies only to systems that run without a person in the loop, and its three clauses come from Chapter 13. The goal must be checkable: “A goal you can loop toward is a goal whose satisfaction a machine can check.” The cap is not optional: “a hard iteration cap is the floor beneath every goal function: it catches the case where ‘done’ never comes true.” And the kill switch must exist before it is needed: “one obvious, fast, tested way to stop every running loop and agent at once.”

For high-stakes actions, question 6 asks where the gates live. Chapter 12’s rule is that “the gate must live in your code, never in the agent’s judgment,” set by what the action can do (read-only, reversible, externally visible, irreversible) rather than by how sure the model sounds. The person at the gate should get something they can decide on: “Send a decision, not a transcript.” The consequence tier classifier helps sort your actions into those tiers.

What does the scorecard not tell you?

The scorecard does not tell you whether the agent is good; it tells you whether you would know. A system can score 24 and still be wrong often, as long as the signals catch it. A system can score 6 and work well for months, and no one would be able to say why or for how long. It also takes your answers on trust. A “yes” to question 10 is only as good as the eval set behind it, and a “yes” to question 11 only as good as the audit you ran.

Use it before a launch review, and again whenever a model, a prompt or a tool changes, because the score falls quietly when one does. For the individual checks, the lethal trifecta audit, the eval sample size calculator and the agent bug bestiary go deeper. Should this be an agent? asks question 8 at the start, before anything is built. The compass and its four bearings are set out in Chapter 1, What Is an Agent?, which is free to read.

Questions readers ask

Why is “what signal tells you it worked?” the first question?
Because a model's confidence says nothing about whether it is right. Chapter 1 makes verifiability the first bearing of the book's compass: an agent is only as trustworthy as the signal you can use to verify it, so trust has to come from outside the model, such as a passing test, a validating schema, a source you can open or a human who approves. Every other question on the scorecard is a version of that one for a different part of the system.
Can a system score “product” with no human in the loop?
Yes, if the signals are mechanical. Nothing on the scorecard requires a person to approve every run. Question 6 asks that the gates be set by consequence tier, so the irreversible actions wait for a human while the rest run free, and question 7 asks that anything unattended have a machine-checkable stopping condition, a hard cap and a tested kill switch. A product band means you can say how you know each run worked, not that someone watched it.
Which question should a team fix first?
Usually the first “no” in the list, and most often that is question 1. If you cannot name the signal that tells you a run worked, the later questions have nothing to measure: you cannot grow an eval set, gate a merge or roll back a canary without it. After that, the tool lists every “no” before every “partly”, in the book's order.
What do the bands mean?
They are the tool's own cut points on the total out of 24: 0 to 8 is a demo, 9 to 16 a pilot, 17 to 24 a product. The book's version of the distinction is a line it quotes in Chapter 25, “Demo is works.any(), product is works.all()”: a demo needs one run to succeed, a product needs every run to. The band is a rough reading of how many of your runs you could check.
Is this a security audit?
No. Question 11 asks whether you have run the lethal-trifecta audit and set the blast-radius dials, but the scorecard only records the answer. The lethal trifecta audit tool runs the audit itself, and the consequence tier classifier helps place the gates question 6 asks about.

Sources

  1. Anthropic (2024). Building Effective Agents
  2. Anthropic (2026). Demystifying evals for AI agents
  3. Hamel Husain (2024). Your AI Product Needs Evals