Home / Tools / ROI versus risk map

Free tool · runs in your browser · from Chapter 23

ROI versus risk map

Work out the return on an AI agent use case after the cost of checking its output, place it on a value-versus-risk map, and see when to refuse.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The two axes, subtracting the cost of checking before reading the chart, the four quadrant labels and the caption, the analyst's two conditions, the quadrant to refuse, the baseline before deploying, and the difference between deeds and words are the book's (Chapters 14 and 23). The tool's own: every formula, both break-evens, the two reference lines that divide the map, the logarithmic axes, the order in which verdicts override each other, and every default.

How do you estimate the return on an AI agent?

You estimate the return on an AI agent by taking the value of the work, subtracting what it costs to check that work, and then subtracting what its errors are expected to cost. The tool above does that arithmetic in your own unit and places the use case on the map Chapter 23 of the book draws for the budget meeting.

Chapter 23 calls the question “the only question that survives contact with a spreadsheet: is this worth it?” It answers with two axes and one correction. The axes come first: “Plot any candidate use case by how much the work is worth (value per instance times volume) and by what an error costs when it lands.” The correction is the part most business cases skip: “before reading the chart, add the third number: what it costs to check the work”.

Where a candidate use case earns its keep.
Figure 23.5 Where a candidate use case earns its keep. Plot the value of the work against the cost of an error, and subtract the cost of checking the work from its value before you read the chart. Internal, low-stakes lookups (in accent) are where to start; high-value work whose errors are costly earns an agent only behind gates and review; low-value work is a script, or nothing at all. Reuse this diagram

The figure has four labels: start here, gate heavily, low priority and not worth it. The tool redraws it with your use case as a point and keeps those labels as they are.

What does the tool compute?

The tool computes four lines per period and one net figure. The formulas are the tool’s own. The chapter names the three quantities and does not write an equation.

  • Value of the work = value of one instance × instances per period.
  • Checking = cost of checking one instance × instances per period.
  • Expected errors = instances per period × share where an error lands × cost of one landed error.
  • Net value = value − checking − expected errors.

The defaults describe an internal lookup counted in hours per month. Each answer saves a quarter of an hour, there are 2,000 a month, checking costs 0.02 hours per answer, 3% of answers carry an error that lands, and each such error costs an hour to put right. The value is 500 hours, checking takes 40, errors take 60, and the net value is 400 hours a month. The numbers are illustrative and are not measurements of any system.

The error rate needs one clarification. It is the share of instances where an error lands after checking, which is the figure that hurts. More checking would normally lower it, and the tool does not model that link. You enter both numbers as you measured or estimated them. The eval sample size calculator says how many cases you need before a small error rate means anything.

Why does checking come off the value first?

Checking comes off the value first because an agent that must be fully re-checked has saved nobody any work. Chapter 23 states the limit case in one sentence: “An agent whose output costs as much to verify as to produce by hand has a return of approximately nothing, however impressive the demo.”

The tool applies that sentence literally. When the cost of checking one instance is equal to or greater than the value of that instance, the headline says the return is nothing, and it says so in every quadrant of the map. The one exception is a use case with no volume or no value at all, where the tool asks for a number first.

Chapter 14 explains why this cost goes missing. It prices every rung of the ladder, from plain code up to an agent, with the same four lines: “A rung’s true quote has four lines: what it costs to build, what it costs to run, what it costs when it is wrong, and what it costs to know whether it worked.” Then it adds: “The fourth line is the one engineers leave off”. For plain code the fourth line is a unit test written once. For an agent, the same chapter calls it “the most expensive knowledge on the ladder.” The agent verifiability scorecard looks at how cheaply a given task can be checked at all.

Where do the map’s dividing lines come from?

The dividing lines come from two numbers you enter, because the book gives none. The figure is a sketch with a dashed line down the middle and another across it. It does not say how much value is “high” or how much error cost is “costly”, and a page that holds no prices cannot say it for you.

So the tool asks two questions of its own:

  1. What is the smallest value per period, after checking, that would justify a project? That is the horizontal line.
  2. What is the largest single-error cost your team absorbs as a correction and not as an incident? That is the vertical line.

The second question borrows a distinction Chapter 23 draws when it explains why internal work comes first: “the errors are recoverable (a wrong answer to a colleague is a correction, a wrong answer to a customer is an incident)”.

Each axis is drawn on a logarithmic scale centred on your line, so a point ten times above the line and a point ten times below it sit the same distance away. The solid point is your use case after checking. The hollow point above it shows where the same use case would sit if checking were free, which is how far the uncorrected spreadsheet overstates it.

What do the four quadrants mean?

The four quadrants are the figure’s reading of high or low value against cheap or costly errors. The caption gives three of them in a sentence each.

  • Start here. High value, cheap errors. The caption says “Internal, low-stakes lookups (in accent) are where to start”. The chapter gives the reason for starting there: the integration, the audit trail and the permission scoping all get proven where a bad week is cheap.
  • Gate heavily. High value, costly errors. The caption: “high-value work whose errors are costly earns an agent only behind gates and review”. The consequence tier classifier sorts individual actions into what may run unattended and what waits for a signature.
  • Low priority and not worth it. The caption treats the two bottom quadrants together: “low-value work is a script, or nothing at all.” The should-this-be-an-agent tool covers the choice between a script, a single model call, a workflow and an agent.

The quadrant and the net value are separate results, and they can disagree. A use case can sit in “start here” and still lose value, if errors are frequent enough. When that happens the tool reports both and lets the arithmetic decide the headline.

When does the tool say to refuse?

The tool says to refuse when the output is words, no person will review it before something irreversible depends on it, and one error costs more than your error line. This verdict overrides the map. Only the checking rule above outranks it.

It comes from the chapter’s account of the research agent, which it calls the analyst. The analyst pays off under two conditions together: “the question is valuable enough to justify the machinery, and a person will review the output before anything irreversible depends on it.” The machinery is costly. The chapter cites a measurement that puts a multi-agent research design at “roughly fifteen times the tokens of a chat”, and marks the corner to avoid: “The quadrant to refuse is the mirror image: questions where a wrong fact is dangerous and no one will check.”

Reading “dangerous” as “above the error line you entered” is the tool’s interpretation. Answer that a person always reviews the output, and the refusal lifts. The tool then treats the use case as its quadrant and its arithmetic describe it.

Why do words and deeds get different warnings?

Words and deeds get different warnings because the cost of an error can be capped for one and not for the other. This is the part of Chapter 23 that runs against intuition. An agent that issues refunds looks dangerous, and one that writes briefs looks safe.

The chapter reverses that. Actions submit to engineering: “permissions bound what can be touched, consequence tiers bound what runs unattended, approval gates bound the worst transaction”. A document has no such limit, because “a wrong fact costs nothing when it is written and an unknown amount when it is believed”. The chapter compresses the contrast into two sentences: “Deeds fail loudly at a size you chose. Words fail silently at a size the future chooses.”

The tool turns this into two different checks.

  • For words, it reminds you that the error cost you typed is a guess, and it asks whether a person reviews the output. The chapter’s position on the analyst is that its output “keeps a human in the loop at any level of measured skill”.
  • For deeds, it asks whether the worst single action is capped and whether the current process has a recorded baseline. If the action is not capped, the error cost you entered describes a typical error and not the worst one.

The practical conclusion is the chapter’s: “autonomy should therefore track error tolerance rather than apparent stakes”.

What should exist before the pilot?

A baseline for the current process should exist before the pilot, along with an answer to the chapter’s two closing questions. For an agent that acts, the baseline is what makes the value per instance checkable. Chapter 23 says such work comes with numbers you can record (cost per ticket, handle time, resolution rate) and passes on a survey’s advice to “establish baseline metrics for processing time, error rate, and satisfaction so agent performance can be measured accurately,” before deploying.

The break-evens the tool reports are useful at the same stage. One is the highest error rate at which the net value is still positive. The other is the highest checking cost per instance at which it is still positive. With the defaults they are 23% and 0.22 hours. If your measured error rate is close to the first number, or your reviewers are close to the second, the use case has little room.

The chapter closes its accounting with two questions for the budget meeting: “can you afford to check the words, and can you afford the deeds when the checking fails?” The tool gives a number to each. The guide to agents in practice places this chapter among the other application chapters. The full argument, with the research agent and the business agent worked through in detail, is Chapter 23, Research and Business Agents (in the full book).

Questions readers ask

Why subtract the checking cost before reading the chart?
Because checking is a real cost that the usual calculation leaves out. Chapter 23 tells the reader to add a third number, what it costs to check the work, before reading the map. An agent whose output costs as much to verify as to produce by hand returns approximately nothing. The tool subtracts the checking cost from the value first and places your point at what is left.
Why does the book say the safe-looking research agent carries the risk that cannot be capped?
Because a document has no built-in limit on the damage a wrong fact in it can do. Chapter 23 says an agent that acts can be bounded by permissions, consequence tiers and approval gates, so its worst case is a size you chose. A wrong fact in a brief costs nothing when written and an unknown amount when someone believes it weeks later.
What baseline should exist before a business-agent pilot?
The current process should have recorded numbers for processing time, error rate and satisfaction. Chapter 23 quotes that advice from an industry survey and adds that the baseline must exist before deploying. Without it, the value per instance you enter here is an estimate with nothing to measure it against once the agent runs.
Where do the lines that divide the map come from?
They come from you. The book draws a two-by-two and gives no thresholds, and a page with no prices cannot supply them. You enter the smallest value per period that would justify a project and the largest single-error cost your team absorbs as a correction. The tool divides the map there and says that the lines are its own device.
What unit should I use for the amounts?
Use any single unit you already count in, such as hours of staff time. Every amount on the page must be in that same unit: the value of an instance, the cost of checking it, the cost of an error and both dividing lines. The page holds no prices and no currency, and the result is reported in your unit per period.

Sources

  1. Anthropic (2025). How we built our multi-agent research system
  2. Algolia (2026). AI agent use cases: where enterprises are deploying agents today
  3. Raja Parasuraman, Victor Riley (1997). Humans and Automation: Use, Misuse, Disuse, Abuse (Human Factors 39(2):230–253)