AI agent ROI is the value of the work an agent does, minus what it costs to run, minus what it costs to check that work, minus what its errors cost when they get past the check. The sum on most slides stops after the second term. The last two decide whether the project is worth funding.
This post is for a CTO who has to take a number into a budget meeting. It takes one illustrative use case from the first sum to the corrected one, embeds the ROI versus risk map on that case, and lists what would make the answer “no”. Every amount is in hours. There are no prices here, because prices move and your hours are yours.
What is AI agent ROI once the missing terms are in?
AI agent ROI, stated in full, is a net value per period: instances × (value per instance − checking cost per instance − landed error rate × cost of one landed error), with every amount in the same unit. A landed error is one that got past the check and reached a customer, a system or a decision.
The frame comes from Chapter 23 of the book (in the full book), in the section “Framing Return on Investment Against Risk.” It opens with the question this post is about: “the only question that survives contact with a spreadsheet: is this worth it?” The chapter then names the axes: “Plot any candidate use case by how much the work is worth (value per instance times volume) and by what an error costs when it lands.”
The chapter gives the quantities and no equation. The formula above is the one the tool uses, and the tool’s page says so. What I add in this post is the argument for a budget meeting: why each term belongs on the page, what to do when a term cannot be known, and when to walk away.
Why is the usual AI agent ROI sum wrong?
The usual AI agent ROI sum is wrong because it counts the work the agent produces and not the work the agent creates. It reads: hours saved per task, times tasks per month, minus the run cost. Reviewing the output and repairing the errors that slip through are new work, and neither appears.
The people asking for the number already suspect this. An Ask HN thread from December 2024 put the question in the form finance puts it: “What’s the ROI? How many hours or dollars have you saved?” A reply in a 2025 Hacker News discussion quoted an article’s account of where saved time goes: “The 30 minutes you saved on data analysis? You’re using it to manage two more AI tools and review their outputs.”
The pages that rank for “AI agent ROI” are not blind to it either. Of four pages that ranked for the phrase in October 2026, two list review among the costs. One of them, from the agent-platform vendor Pickaxe, says of projections it has seen: “They forget the human-in-the-loop costs.” Another, from StackAI, asks for a “Human review rate” as an input. I’m not claiming nobody mentions review. My claim is narrower: the short sum that reaches the slide drops it, and none of the four pages separates an error whose cost you can cap from one whose cost you cannot.
The hours-saved figure has a second weakness. It is often an estimate made by the people doing the work, and such estimates can be wrong in sign. The baseline section below returns to that.
What does checking cost, and where does it go in the sum?
Checking costs whatever time a competent person, or a calibrated automatic check, spends deciding whether one output is right, and it comes off the value before anything else. Chapter 23 puts the limit case in one sentence: “An agent whose output costs as much to verify as to produce by hand has a return of approximately nothing, however impressive the demo.”
The chapter lists what sits in that bill for a research agent: “the load-bearing spot-checks, the judge’s calibration, the reviewer’s hour.” For an agent that drafts replies, the bill is the reader’s time on each draft. For an agent that acts in a system, it is the approval queue, the audit sample and the reconciliation report.
Two features of checking make it easy to leave out. First, it is paid by different people than the ones who wrote the business case, in minutes that no system records. Second, it does not shrink when the agent’s output gets more fluent. The book calls this the verification gap: “the distance between how convincing research output looks and how cheaply it can be checked”. A commenter on Hacker News described the same thing from the reviewer’s chair in 2025: errors “look correct, which makes it all the more difficult to determine” that they are wrong.
The run cost belongs in the sum too, and it is the term that business cases already include. Chapter 19 (in the full book) explains why it is larger and less even than a per-call estimate suggests, and flags retries as the layer “nobody budgets for.” The guide to what an agent costs to run and the post on AI agent pricing models cover that side. Here I convert the run cost into hours and subtract it from the value of each instance.
What does an error cost when it lands?
A landed error costs the time and damage that follow from a wrong output nobody caught, and the business case needs two facts about it: how often it happens after checking, and how much one costs. The first is a rate you can measure on an eval set and then in production. The second depends on what the agent produces.
Measuring the rate takes volume. A pilot of 100 instances at a true rate of 2% yields about 2 landed errors, which cannot distinguish 2% from 5%. The eval sample size calculator says how many cases a small rate needs.
Why can a wrong action be capped when a wrong sentence cannot?
A wrong action can be capped because its size is set by controls you install before the pilot, and a wrong sentence has no such control once someone believes it. Chapter 23 lists the controls: “permissions bound what can be touched, consequence tiers bound what runs unattended, approval gates bound the worst transaction”. An agent that may issue a credit up to a fixed limit, with an approval gate above it, has a worst case you wrote down. That worst case is its blast radius.
Text has no limit of that kind. The chapter’s sentence is “a wrong fact costs nothing when it is written and an unknown amount when it is believed”. It compresses the contrast into two more: “Deeds fail loudly at a size you chose. Words fail silently at a size the future chooses.”
A public example shows who ends up choosing the size. In February 2024, CBC News reported that a British Columbia tribunal held Air Canada liable for what its website chatbot told a customer about bereavement fares. The tribunal member wrote: “It makes no difference whether the information comes from a static page or a chatbot.” He also found the airline “did not take reasonable care to ensure its chatbot was accurate”. The amount in that case was small. The point for a business case is that the cost of the sentence was set afterwards, by a third party, and nobody at the company had capped it.
So the error cost on your page is one of two things. For a capped action it is a bound. For unreviewed words it is a guess, and the arithmetic built on a guess does not deserve the same trust.
Which baseline has to exist before the pilot?
The baseline that has to exist is a recorded measurement of the current process: how long one instance takes, how often it goes wrong, and how satisfied the people served by it are. Chapter 23 quotes an industry survey by Algolia on this: “establish baseline metrics for processing time, error rate, and satisfaction so agent performance can be measured accurately,” and adds that it must happen before deploying.
Without it, the value per instance is somebody’s impression. A randomized controlled trial published in July 2025 by Becker, Rush, Barnes and Rein shows how far an impression can be from a measurement. It followed 16 experienced open-source developers across 246 tasks in their own projects, with the AI coding tools of early 2025 allowed or disallowed at random. Afterwards the developers “estimate that allowing AI reduced completion time by 20%.” The measured result: allowing AI “actually increases completion time by 19%”.
That study is about coding assistance for experts in familiar code, at one point in time. It says nothing about your support queue. I use it for one thing only: a sincere estimate of hours saved was wrong in direction, and only the measured comparison showed it.
The book treats the baseline as the advantage that operational work has over research work. High-volume, bounded workflows come with “a baseline you can put a number on (cost per ticket, handle time, resolution rate)”, and the agent can then be measured against it “continuously, in production”. The tool asks about the baseline only for agents that act. I’d ask for it in every case, because the first term of the sum rests on it.
A worked AI agent ROI case: from the first sum to the corrected one
Here is a worked example with illustrative numbers. They are not measurements of any system. A software company wants an agent that drafts replies to support tickets, and a support engineer reads each draft before it goes out. The unit is hours per month.
- 1,500 tickets a month.
- Writing a reply unaided takes 0.30 hours.
- The run cost, converted to the same unit, is 0.01 hours per ticket.
- Reading and correcting a draft takes 0.10 hours.
- After that review, 2% of sent replies still carry an error that lands.
- One landed error costs 3 hours: a reopened ticket, a second engineer, a correction to the customer.
What does the first sum say?
The first sum says 435 hours a month. That is 1,500 × 0.30 = 450 hours of writing saved, minus 1,500 × 0.01 = 15 hours of run cost. Put another way, each ticket is worth 0.29 hours, and that is the value per instance used from here on.
What does the corrected sum say?
The corrected sum says 195 hours a month, which is 45% of the first figure (195 ÷ 435).
| Line | Hours per month | How it is computed |
|---|---|---|
| Value of the work, after run cost | 435 | 1,500 × 0.29 |
| Checking | −150 | 1,500 × 0.10 |
| Value after checking | 285 | 435 − 150 |
| Expected landed errors (30 a month) | −90 | 1,500 × 2% × 3 |
| Net value | 195 | 285 − 90 |
Two break-evens show how much room the case has. The net value stays positive while the landed error rate is under (0.29 − 0.10) ÷ 3 = 6.33%, against the 2% assumed. At a 2% rate, checking could cost up to 0.29 − 0.02 × 3 = 0.23 hours per draft, against the 0.10 assumed.
The map needs two lines of your own. Say the team would not start a project for less than 100 hours a month after checking, and absorbs an error of up to 2 hours as a correction. Then 285 is above the first line and 3 is above the second, and the case sits in the figure’s “gate heavily” corner. The embedded tool opens on these numbers.
With JavaScript on, the ROI versus risk map runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the ROI versus risk map on its own page to share a result by link.
One line is still missing, and the tool has no field for it: the fixed cost of having the agent at all. That covers the build spread over its expected life, the eval set, monitoring, and changes to prompts and tools. Suppose it is 60 hours a month. The case then returns 195 − 60 = 135 hours a month, and the value line of 100 was set above 60 on purpose.
What would make the answer no?
Four conditions make the answer no, and two make it “not yet”. The order below is the order in which the tool’s verdicts override each other; the wording and the last two rows are this post’s.
| # | Condition | Verdict | Variant of the worked case that trips it |
|---|---|---|---|
| N1 | Checking one instance costs as much as the instance is worth | No: the return is nothing | The reviewer rewrites every draft: checking 0.29 hours, net −90 hours |
| N2 | The output is words, review before anything irreversible is not guaranteed, and one error costs more than your error line | No: refuse, whatever the arithmetic says | Review is dropped to rescue the number; say 8% of replies then land wrong: the sum still shows +75 hours |
| N3 | The landed error rate is at or above (value − checking) ÷ cost of one error | No: the net value is zero or negative | The landed rate is 7%, above the 6.33% break-even: net −30 hours |
| N4 | The value after checking is under your value line | No agent project: a script, or nothing | 300 tickets a month: 57 hours after checking, against a line of 100 |
| Y1 | No baseline is recorded for the current process | Not yet: measure first | The 0.30 hours per reply is a team lead’s estimate |
| Y2 | The agent acts, and its worst single action has no cap | Not yet: cap first | The same agent is allowed to issue credits with no limit |
N2 is the one that surprises people, because the arithmetic disagrees with it. Dropping the review removes 150 hours of cost, and even at four times the error rate the sum stays positive. The rule comes from the chapter: “The quadrant to refuse is the mirror image: questions where a wrong fact is dangerous and no one will check.” Reading “dangerous” as “above the line your team absorbs” is the tool’s interpretation, and I think it is the right one for a budget meeting.
How do three other cases come out?
Three cases that are not the worked example come out as one “go”, one refusal that turns into a gated “go”, and one “go” with little room. The numbers are illustrative again, in hours per month.
| Case | Inputs | Trips | Result |
|---|---|---|---|
| Internal lookup (the tool’s default) | 2,000 answers worth 0.25; checking 0.02; 3% landed; 1 hour per error; lines 100 and 8 | Nothing | Start here; net 400 |
| Research brief, sent on unreviewed | 20 briefs worth 6; no checking; 10% landed; 40 hours per error; lines 50 and 8 | N2 | Refuse, though the sum shows +40 |
| The same brief with a reviewer | Checking 1.5; 5% landed | Nothing | Gate heavily; net 50; break-even error rate 11.25% |
| Invoice matching (acts; capped; baseline recorded) | 4,000 invoices worth 0.1; checking 0.01; 1% landed; 4 hours per error; lines 100 and 8 | Nothing | Start here; net 200; break-even error rate 2.25% |
The invoice case passes every row and still deserves a warning. Its break-even error rate is 2.25%, and the assumed rate is 1%. A measured rate a little over twice the assumption erases the return.
How do you lower the checking bill without dropping the check?
You lower the checking bill by making each check cheaper or by checking fewer instances where a cap makes that safe. You do not lower it by trusting the output more. Three moves are available, and each has a condition.
The first is to shape the output so that review is fast: a draft that cites the record it relied on takes less time to verify than one that does not. The post on review theater covers what a reviewer can check and what only looks like review.
The second is a cheap automatic check in front of the person. Chapter 19 describes a classifier that “runs one cheap call on a tight scope” with a measured misroute rate. It passes the clear cases and sends the doubtful ones to a person. Its own error rate then becomes part of the landed error rate, so it has to be measured. The comparison of an LLM classifier and a fine-tuned classifier covers the choice.
The third is sampling: review a share of the output and not all of it. This fits capped actions with cheap errors. It does not fit costly words, because sampled review is not review “before anything irreversible depends on it”, and the case falls under N2.
What goes on the one page for the budget meeting?
The page carries the corrected AI agent ROI sum, the two break-evens, an answer on review and caps, and the conditions that would stop the project. Fill it in one unit throughout. The six rows at the bottom are the table above.
AGENT BUSINESS CASE (one page)
Use case: ______________________ Unit: hours per ______
Output is: [ ] words (a draft, an answer, a brief) [ ] deeds (actions in real systems)
BASELINE (current process, recorded on ______)
Time per instance: ____ Error rate: ____ Satisfaction: ____
Source of these numbers: [ ] measured [ ] estimated (then: not yet)
THE SUM, per period
Instances per period (n): ____
Value per instance, after run cost (v): ____
Checking cost per instance (c): ____
Landed error rate after checking (p): ____ %
Cost of one landed error (x): ____ [ ] a cap [ ] a guess
Value of the work = n × v = ____
Checking = n × c = ____
Expected landed errors = n × p × x = ____
NET VALUE = n × (v − c − p × x) = ____
Fixed cost per period (build, evals, monitoring) = ____
NET AFTER FIXED COST = ____
ROOM
Break-even error rate = (v − c) ÷ x = ____ % (assumed: ____ %)
Break-even checking cost = v − p × x = ____ (assumed: ____)
OUR TWO LINES
Value line (smallest value after checking that justifies a project,
at least the fixed cost): ____
Error line (largest single error we absorb as a correction): ____
CONTROLS
Who reviews, and before what irreversible step: ______________________
If deeds, the cap on the worst single action: ______________________
STOP CONDITIONS (any "yes" stops the project or delays it)
N1 c is equal to or above v [ ] yes [ ] no
N2 words, review not guaranteed, x above error line [ ] yes [ ] no
N3 p is at or above the break-even error rate [ ] yes [ ] no
N4 n × (v − c) is under the value line [ ] yes [ ] no
Y1 no recorded baseline [ ] yes [ ] no
Y2 deeds with no cap on the worst action [ ] yes [ ] no
Numbers to re-measure in the pilot, and when: ______________________
Before the meeting, check that each number on the page has an owner and a source.
- A recorded baseline for the current process: time per instance, error rate, satisfaction
- Instances per period, from a system count and not from memory
- Value per instance, with the run cost already subtracted in the same unit
- Checking cost per instance, timed with the people who will do the checking
- Landed error rate after checking, from enough instances to trust it
- Cost of one landed error, marked as a cap or as a guess
- The fixed cost per period of building and maintaining the agent
- A value line and an error line, both set before the result is known
- The name of the reviewer and the irreversible step the review precedes
- For an agent that acts, the cap on its worst single action
- Both break-evens, next to the rates assumed
Which questions expose a one-sided business case?
Five questions expose a one-sided AI agent ROI case, whether it is your team’s slide or a vendor’s calculator. Each one asks for a term of the corrected sum.
- Who checks the output, how long does one check take, and whose budget holds that time?
- What share of outputs is still wrong after the check, and how many instances is that figure based on?
- What does one such error cost, and is that number a cap enforced by a control or an estimate?
- Against which recorded baseline were the hours saved measured, and when was it recorded?
- At what error rate does the return reach zero?
A case that answers the fifth question has done the arithmetic. A case that cannot answer the first has reported the gross value and called it a return. For a purchased agent, the fifteen vendor questions on reliability go further, including what changes in regulated industries.
The shape of the use case predicts which question bites. The catalog of AI agent use cases for startups sorts ideas by whether they hand a person a draft or act on a real system. Drafts fail on questions 1 and 3. Actions fail on question 4 when nobody recorded the process, and on question 3 when nobody set a cap.
Where does this frame stop working?
This way of computing AI agent ROI stops working where its inputs stop being numbers. I see four such places.
Value that is not saved time. The sum prices an agent as a substitute for existing work. An agent that does work nobody did before, such as reading every contract where people sampled a few, has no baseline to beat. Its value has to be argued another way, and the checking and error terms still apply.
The link between checking and errors. More checking normally lowers the landed error rate, and neither this post nor the tool models that link. You enter both numbers as measured. If you change the review process, measure both again.
Costs outside the unit. Converting run cost into hours needs a rate, and that conversion is yours to make. A regulatory penalty or a lost customer does not convert well into hours at all. When the cost of an error is of that kind, treat it as above your error line and rely on the controls, not on the multiplication.
Expensive machinery. Some designs cost far more to run than the worked case assumes. Anthropic reported in 2025 that in its data “multi-agent systems use about 15× more tokens than chats”, and drew the conclusion that such systems “require tasks where the value of the task is high enough to pay for the increased performance.” That is one team’s measurement at one date. For designs like that, the run cost is no longer the small term.
The number to take into the room
Take the corrected figure, its two break-evens and the stop conditions into the room, and leave the first sum out. In the worked case that means saying 195 hours a month before fixed costs, 135 after, positive while fewer than 6.33% of replies land wrong, and conditional on a named person reading every draft. That is a smaller number than 435. It is also one that will still be true after the pilot.
Chapter 23 ends its accounting with two questions that do not depend on which model is on the desk: “can you afford to check the words, and can you afford the deeds when the checking fails?” An honest AI agent ROI figure is the answer to both, written down before the pilot.
The full argument, with the research agent and the business agent worked through, is in Chapter 23, “Research and Business Agents”, in the full book; the cost model is in Chapter 19, also in the full book. The Agents at work guide collects the related posts and tools, and you can see the formats.
Questions readers ask
- How do you calculate AI agent ROI?
- Count everything in one unit, such as hours per month. Take the value of the work after run cost, subtract the cost of checking it, then subtract the landed error rate times the cost of one landed error. Subtract the fixed cost of keeping the agent running. For a percentage, divide what is left by the total cost over the same period.
- What is a good ROI for an AI agent?
- No published multiple can answer that for your process. The figures that are yours are the net value per period and a line you set in advance for the value after checking: the smallest amount that would justify the project, which should be at least the fixed cost of building and maintaining the agent. Multiples quoted on vendor pages describe other organizations and other error costs.
- Does AI agent ROI improve as models get cheaper?
- Only the run-cost term improves, and it is often the smallest. In this post's illustrative case the run cost is 15 of the 450 hours of writing saved, while checking and landed errors take 240. Those two terms fall when the output gets easier to check or the error rate drops, which is a matter of evaluation and task design.
- What if nobody can estimate what an error costs?
- Then treat the cost as above the level your team can absorb. For an agent that acts, put a cap on the worst single action so that the cost has a ceiling you chose. For an agent that writes, keep a person reviewing the output before anything irreversible depends on it, or do not build that one.
- How long should a pilot run before the numbers mean anything?
- Until enough errors have landed to estimate their rate, which depends on volume and not on the calendar. At a 2% landed error rate, 1,500 instances produce about 30 errors; 100 instances produce about 2, too few to tell 2% from 5%. Record the baseline for the current process first, or there is nothing to compare against.
Sources
- Joel Becker, Nate Rush, Elizabeth Barnes, David Rein (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089)
- Anthropic (2025). How we built our multi-agent research system
- Algolia (2026). AI agent use cases: where enterprises are deploying agents today
- Jason Proctor, CBC News (2024). Air Canada found liable for chatbot's bad advice on plane tickets
- Pickaxe (retrieved 2026). AI agent ROI: metrics and formulas
- StackAI (retrieved 2026). AI agent ROI calculator: measure, calculate and maximize the business impact of AI automation
- kyrilku (Hacker News) (2024). Ask HN: How are AI agents and LLMs delivering real value in your company?