Home / Blog / Agents at work / AI Agents for Product Managers: Which Decisions…

Agents at work

AI Agents for Product Managers: Which Decisions Are Yours

AI agents for product managers: the six decisions that are product calls, the artifact that records each, and a requirements page to copy before kickoff.

By Enrique Gutiérrez · Published · 21 min read

AI agents for product managers come down to a division of decisions. Six calls on an agent product belong to product: what counts as done, which errors are tolerable, how much autonomy each task starts with, what the user sees when the agent is unsure or wrong, what the user can undo, and how reliance is calibrated.

I wrote this for a product manager or founder who is about to write requirements for a feature in which an agent acts for a user. It isn’t about the agents a PM can use for their own work, which is what the three pages I fetched from the top results for this phrase in October 2026 cover. When nobody makes the six calls, they still get made: by a default in the prototype, by whoever wrote the prompt, or by the first incident.

AI agents for product managers: which decisions are product calls?

A decision is a product call when its answer changes what a user is promised or what a user can lose. It is an engineering call when two different answers keep the same promise and differ in cost, speed or upkeep. It is shared when product sets the target and engineering says what is feasible and how it will be enforced.

That sorting rule is this post’s own. It rests on a sentence from Chapter 12, Oversight and Autonomy (in the full book), which spends a chapter on gates and autonomy settings and closes the subject this way: “What has not changed, at any position of the dial, is whose signature it is.” Somebody answers for what the agent did. The product decisions are the ones that fix what that person, and the user, signed up for.

Two terms before the map. An agent here is software in which a language model decides the next step and takes actions through tools, such as sending a message or changing a record. A requirements page is whatever your team calls the document that engineering builds against.

Three of these calls have a head start in the book. Chapter 14, Choosing Your Approach: When Not to Build an Agent (in the full book) says an agent is worth building when the goal is open-ended, judgment is needed at branches nobody can list in advance, and three conditions hold: “A success criterion a machine can evaluate,” “A feedback signal at each step,” and “A sensible place for a human to look” (the gates and the dial). The first is decision 1 below, and the third is decisions 3 and 6. A PM who has not made those calls has not yet established that the agent should exist.

What does the decision-rights map look like?

The map for AI agents for product managers has sixteen decisions: six that are product’s, five that are engineering’s, and five that are shared. Each row gives the question, the artifact that records the answer, and what goes wrong when the owner doesn’t make the call. The buttons narrow the table to one owner.

For the engineering rows, the question is the one a PM should ask, and the failure is what happens when product makes the call for engineering or nobody asks.

Decision The question Artifact that records it What goes wrong when the owner doesn’t make it Owner
1. Done What result counts as done for each task, and who or what checks it? An acceptance set: real cases, each with a written pass rule The agent’s own “done” becomes the definition product
2. Tolerable errors Which wrong results can users absorb, at what rate, and which must never happen? An error budget: error classes, a rate for each, a must-never list Engineering ships against a bar nobody else saw product
3. Starting autonomy For each task, how much may the agent do before a person looks? One autonomy record per task The prototype’s setting becomes the product’s product
4. Unsure or wrong What does the user see when the agent can’t finish, isn’t sure, or turns out to be wrong? A failure-state spec: each state, what the user sees, what happens next Only the happy path is designed, and a guess is shown as a result product
5. Undo Which actions can the user take back, for how long, and what are they told about the rest? An undo map: action, way back, window Users learn what is permanent by losing something product
6. Reliance What evidence arrives with each result, and how will we know if users check too little or too much? An evidence spec and two reliance measures Rubber-stamping or redoing by hand, both invisible in usage numbers product
7. Model and routing Which class of model runs each step, and what would make you change it? Architecture note Product picks by reputation; cost and speed go unexamined engineering
8. Tools and their limits What can each tool do, and how does it refuse a call outside its limits? Tool spec Limits are written as prose instructions the model can ignore engineering
9. Where each block and gate is enforced Which code holds each must-never and each signature? Guardrail spec The must-never list exists only as a sentence in a prompt engineering
10. Grader and test harness How is each case graded, and how was the grader checked against people? Evaluation harness notes A pass rate nobody can reproduce engineering
11. Traces, budgets and the stop What is recorded per run, what caps a run, and who can stop every run? Operations runbook An incident with no record and no off switch engineering
12. Agent at all Is this an agent, a fixed workflow or one model call? A one-page proposal against the three conditions The word “agent” is chosen for the roadmap slide shared
13. Consequence of each action Which actions are read-only, reversible, visible to outsiders, or irreversible? Tier list, one line per action An outward-facing action is treated as an internal one shared
14. Contents and size of the acceptance set Which cases are in it, how many, and who labels them? A named, versioned set The set holds only what was easy to collect shared
15. Cost and wait per accepted result What may one accepted result cost, and how long may the user wait? A budget line A design that passes the bar and can’t be afforded shared
16. Rollout and change control Which numbers widen, pause or tighten the rollout, and what is re-run before a prompt or model change ships? Rollout plan with stop numbers A model upgrade changes behavior with no re-test shared

What are the six product decisions?

They are the six rows marked product, and each needs a number or a sentence that only someone who knows the user can supply. The sections below take them in the table’s order.

1. What counts as done, and who checks it?

Done is a rule a second party can apply to a finished task without asking the agent. Write it per task: the state of the world that must be true afterward, and who or what confirms it. One practitioner put the difficulty in a question to other builders (Ask HN, December 2025): “Demos are easy: a task finishes once the happy path works.”

The artifact is an acceptance set: a collection of real cases, each with its pass rule. The field calls it an eval set. An evaluation vendor’s blog states the product side of it this way: “The PM defines what”good” looks like through structured, repeatable tests” (Braintrust, March 2026). I agree with the sentence and not with that post’s title, “Evals are the new PRD”: evaluation records decisions 1 and 2 and says nothing about the other four.

2. Which errors are tolerable, and at what rate?

An error budget lists the kinds of wrong result, the rate of each that users can absorb, and the outcomes that must never happen. It is a product call because the cost of each error lands on a user, and only someone who knows the user can price it.

Name the classes from real failures. A consultant’s account of one client (Husain, March 2025) describes annotating “dozens of conversations” and finding an assistant “failing 66% of the time when users said things like”let’s schedule a tour two weeks from now.”” After targeted fixes, “Their date handling success rate improved from 33% to 95%.” Those are one client’s figures as the consultant reports them. The point for a PM is that “date phrases” was a class nobody had named until someone read the transcripts.

The rate must be written before the build. One engineer’s essay on requirements for AI features (Pan, May 2026) describes the alternative: “The trap most teams fall into is leaving the threshold for the engineer to set after the fact, in a Slack message that nobody archives.” His remedy is “a number in the doc, owned by the PM, debated up front, and changeable only through an explicit amendment with a paper trail.”

3. How much autonomy does each task start with?

Each task starts at a chosen setting between a person signing every consequential action and the agent running unattended with its work sampled afterward. The book calls the instrument the autonomy dial and defines it as “how much an agent may do between moments of your attention, set per task rather than built into the system.”

It is a product call because the setting decides what the user is asked to do: approve each step, approve a plan, or review a sample. The four positions, the evidence that moves a task between them and a record to fill per task are in the post on the autonomy slider for AI agents. Product chooses the starting position and owns the record. One line of it isn’t negotiable, in the book’s words: “Irreversible or high-stakes actions wait for a signature, every time.”

4. What does the user see when the agent is unsure or wrong?

The user should see a named state with a next step: handed to a person, needs one answer from you, couldn’t verify, or was wrong and here is the repair. Chapter 24, Agent UX and Human Trust (in the full book) gives the standard: “Honest uncertainty, surfaced early, with a repair path, is what graceful failure means.”

The artifact is a failure-state spec. Pan’s essay gives its shape for refusals, timeouts and tool failures: “For each row the spec names two things: what the user sees, and what the system does next.” The first half is product’s, and it includes the words on the screen.

A confidence percentage is a poor substitute for these states. Chapter 24 says such a score is wired to “the model’s self-report, which correlates poorly with correctness”. If you display one, treat it as one signal among several.

5. What can the user undo?

The undo map lists every action the agent can take, the way back from it, and how long that way stays open. It is a product call because it sets the user’s real exposure. A task whose every action can be reversed for a day is a different product from the same task without the window.

Engineering knows what can be reversed technically. Product decides which reversals the user is offered, what the user is told about the actions with none, and whether such an action is offered at all. Chapter 24 puts “the undo story” on every item a person is asked to approve, next to “the action in plain language.”

6. How is reliance calibrated?

Reliance is calibrated when users check the agent’s work about as much as its reliability warrants. Chapter 24 defines trust calibration as “the process of adjusting how much you rely on an automated system until the reliance matches the system’s actual reliability,” and appropriate reliance is the result. It fails in two directions: over-trust, where wrong results are accepted unread, and under-trust, where good results are redone by hand.

This is the decision a product team is most likely to get backwards. The chapter describes the instinct: “A product team told to make users trust the agent will reach for exactly the moves that work: confident tone, polished summaries, stumbles smoothed over, the green check prominent.” Its verdict follows: “Every one of those moves raises trust without touching reliability, which is to say every one of them manufactures over-trust.”

The first artifact is the evidence spec: what arrives with each result. The book’s figure has four cards and sets the raw transcript aside.

Show your work as a curated pack, not a transcript.
Figure 24.2 Show your work as a curated pack, not a transcript. The record a reviewer needs is four cards—what the agent read, what it did, what changed, and (in accent) what proves the result. The raw reasoning stream, set aside at lower left, is longer and less load-bearing; transparency is this curation, not volume. Reuse this diagram

The reason is a cost the chapter states in two sentences: “A claim with its evidence attached costs a glance to check. A claim without it costs an investigation”, and the investigation is what a busy user skips.

The second artifact is a pair of measures, one per direction. For over-trust, the share of results accepted with the evidence never opened. For under-trust, the share redone or re-entered by hand after the agent finished. Both names are mine.

An always-on reliability display does not do this job. In a simulated inspection experiment, Okamura and Yamada (PLOS ONE, 2020) report: “The participants of the group without TCC did not change their behavior when they were in over-trust status, even if the system reliability information was continuously presented.” TCC is their term for a cue shown at the moment over-trust is detected. It is one study in a simulator.

Which calls are engineering’s, and what should a PM ask?

Rows 7 to 11 are engineering’s: the model, the tools, where each block is enforced, the grader, and the operational plumbing. A PM should ask the question in each row and should not answer it.

The test is the sorting rule. If engineering swaps the model for a cheaper one and the acceptance set still clears the bar, the user was promised nothing different. Product’s lever is the bar and the re-test, which is row 16.

Row 9 deserves a second look, because it is where a product decision becomes real. A must-never outcome on the error budget is a wish until some code refuses it. Chapter 12 says of approval gates that “the gate must live in your code, never in the agent’s judgment.” Ask which code, and ask to see the test.

Engineering has its own version of a decision-rights page for the case where agents write the software. It is in the post on an agentic coding workflow for teams, which records who signs what at each stage.

Which decisions are shared?

Rows 12 to 16 are shared: each needs a fact from product and a fact from engineering, and neither side can settle it alone.

Whether to build an agent at all comes first. Product knows the task’s value and who would look at the output; engineering knows whether the steps can be drawn in advance. The should this be an agent tool walks the questions, and the agent verifiability scorecard scores how cheaply the work can be checked. For candidate products, the eight recipes in AI agent use cases for startups each come with a suggested gate.

Consequence is the second. Product knows which actions a customer or a third party will see; engineering knows which can be reversed. Together they sort each action into the four tiers described under approval gates keyed to consequence, and those tiers feed decisions 3 and 5.

The remaining three follow the same pattern. Product supplies the cases and the labels for the acceptance set, and engineering says how many are needed for the bar. Product says what an accepted result is worth, and engineering says what it costs. Product names the numbers that pause the rollout, and engineering wires them.

How does a requirements page for an agent differ?

It differs in one structural way: acceptance is a rate on a set of cases, where a deterministic feature has scenarios that each pass or fail. A deterministic feature is one in which the same input always produces the same output. Pan’s essay lists what the usual template lacks: “It also has no field for”what the model gets wrong five percent of the time.””

Section Deterministic feature Agent feature
Acceptance Each scenario passes or fails, once A rate on a named, versioned set, with an interval
Outcomes per case Pass, fail Done unaided, handed to a person, wrong
Errors Bugs, to be fixed A budget: classes, rates, and a must-never list
Failure screens Error messages Designed states with a next step
Autonomy Not applicable A starting position per task, and what moves it
After launch Regression tests The same set re-run on every prompt or model change

A single demo can’t stand in for the set, and the arithmetic shows why. Nine passes in ten runs is 90%, with a 95% interval from 59.6% to 98.2%. That is consistent with an agent that fails one time in fifty and with one that fails four times in ten. The post on why an agent works in the demo and fails in production follows one such case through 200 real tickets.

Three outcomes per case matter because a handoff is neither a pass nor a failure. An agent that hands every hard case to a person is rarely wrong and saves little. An agent that never hands off looks capable and hides its errors. Product has to set a bar on both numbers.

The must-never list needs separate treatment. Zero failures in 20 cases is a 0% observed rate whose 95% interval reaches 16.1%. No affordable test set shows that an outcome never happens, so a must-never is enforced by code that blocks it, and the set only confirms the block is wired in. The post on how many eval examples you need has the sizes for other bars.

What goes on the agent requirements page?

One page per agent feature. The case for an agent sits at the top (row 12 of the map, numbered 0 on the page), then the six product decisions, the other four shared ones, and the questions for engineering last. Copy it and treat each empty bracket as a decision still open.

AGENT REQUIREMENTS: [feature], [product owner], [date]

0. Why an agent (shared)
   - The steps can't be drawn in advance because: [...]
   - Machine-checkable success criterion: [...]
   - Feedback at each step comes from: [...]
   - A person looks at: [...]

PRODUCT DECISIONS

1. Done
   - Task: [one line per task the agent performs]
   - Done means: [the state that must be true afterward]
   - Checked by: [a system of record | a test | a named role]
   - Acceptance set: [name, version], [n] cases, labeled by [who]

2. Tolerable errors
   - Per case, three outcomes: done unaided | handed to a person | wrong
   - Bar: done unaided at least [x of n]; wrong at most [y of n]
   - The bar applies to: [the observed rate | the interval's worse end]
   - Error classes and what each costs the user: [...]
   - Must never happen (blocked in code, see 9): [...]

3. Starting autonomy (one line per task)
   - [task]: starts at [every consequential action signed |
     standing permissions | plan-level approval | monitored autonomy]
   - Moves looser when: [...]   Moves tighter when: [...]
   - Autonomy record kept by: [name]

4. Unsure or wrong
   - Handed to a person: user sees [...]; then [...]
   - Needs an answer from the user: user sees [...]
   - Could not verify: user sees [...]
   - Found wrong afterward: user is told [...] by [whom] within [...]

5. Undo
   - [action]: way back [...], open for [...]
   - Actions with no way back: [...]; the user is told [...] before each

6. Reliance
   - Shown with every result: what it read | what it did |
     what changed | what proves it
   - Over-trust measure: results accepted with evidence unopened [%]
   - Under-trust measure: results redone by hand afterward [%]
   - Reviewed by [name] every [interval]; action if either exceeds [...]

SHARED DECISIONS

13. Consequence tier of each action: [action: tier]
14. Acceptance set: sources of cases [...]; size agreed with engineering [n]
15. One accepted result may cost [...] and take [...]
16. Rollout: widen when [...]; pause when [...];
    re-run the acceptance set before any prompt or model change

QUESTIONS FOR ENGINEERING (theirs to answer)

7.  Which class of model runs each step, and what would change it?
8.  What can each tool do, and how does it refuse a call outside its limits?
9.  Which code blocks each must-never, and which test proves it?
10. How is each case graded, and how was the grader checked against people?
11. What is recorded per run, what caps a run, and who can stop every run?

A worked example: how does a clinic-booking agent fill the page?

Here is a worked example with illustrative numbers. A clinic chain wants an agent that handles patients’ messages asking to move or cancel an appointment. About 2,000 such messages arrive a week.

Done. The appointment in the scheduling system matches what the patient asked for, and the patient has a confirmation with the new time. The scheduling system is the check, since the agent’s reply doesn’t count as proof.

Tolerable errors. The set has 200 past messages. The bar: at least 170 done unaided (85%), at most 6 wrong (3%), and the remaining 24 (12%) handed to the front desk. At 2,000 messages a week that is 1,700 done, 240 handed off and 60 wrong. If a wrong booking takes staff 15 minutes to repair, the 60 cost 15 hours a week; if a handoff takes 4 minutes, the 240 cost 16 hours.

Then the interval. A run that lands exactly on 6 wrong of 200 has a 95% interval from 1.4% to 6.4%, so the PM must say whether 3% is a bar on the observed rate or on the worse end. Holding the worse end near 4% means observing about 3 wrong in 200 (interval 0.5% to 4.3%), or building a larger set. The must-never list: giving clinical advice, and booking a patient with a clinician outside their referral.

Starting autonomy. Looking up slots and proposing times run freely. Moving an appointment runs under standing permission. Cancelling keeps a signature from the front desk, because the freed slot can be taken by someone else within minutes.

Unsure or wrong. Three states: “I’ve passed this to the front desk, who will reply by [time]”; “Which of these two appointments do you mean?”; and, for a wrong booking found later, a call from a person.

Undo. A move can be reversed by the patient from the confirmation message until the clinic’s cutoff. A cancellation has no guaranteed way back, and the message says so before it happens.

Reliance. The front desk sees four things with each handoff or signature request: the patient’s message, the slots checked, the change proposed, and the scheduling system’s record. The two measures are signatures given with the record unopened, and bookings the staff re-enter by hand.

How do two other agents fill the page?

The same six lines hold for an agent that takes no action and for one that writes to a system of record, and what changes is which lines carry the weight. Both sketches are illustrative.

Contract-summary drafts for a legal team Supplier invoices entered into a ledger
Done Every clause on the firm’s list is summarized with a citation that opens to the clause The ledger entry matches the invoice’s supplier, amount, date and tax
Tolerable errors A missed clause is the costly class; an invented clause is a must-never A wrong amount is the costly class; paying is a must-never
Starting autonomy A lawyer reads every draft Entries under standing permission; anything outside the supplier list signed
Unsure or wrong “I couldn’t find a termination clause” is a state, shown as such Unreadable invoice: handed to accounts with the page attached
Undo Nothing leaves the product without a person, so the draft is discarded A reversing entry, available until the period closes
Reliance Citations beside each claim; measure drafts accepted with no citation opened The invoice image beside the entry; measure entries re-keyed by hand

The drafting agent shows where the page thins out. With no action taken, decisions 3 and 5 take one line each, and decision 6 becomes the product: the gap between how convincing a summary looks and how cheaply it can be checked is what the book calls the verification gap.

What should be settled before kickoff?

Nine items. Seven of them are yours: the six product decisions, with the must-never list on a line of its own. Work through the list with the engineering lead in the room, and treat an unticked item as a decision that will be made by default.

  • Done: each task has a written pass rule and a named check that doesn’t rely on the agent’s own report.
  • Tolerable errors: the bar is a number on a named set, with three outcomes per case, and it says whether it applies to the observed rate or the interval.
  • Must-never list: each item is written down, and engineering has named the code that blocks it.
  • Starting autonomy: every task has a starting position and an owner for its record.
  • Unsure or wrong: each failure state has its screen text and its next step.
  • Undo: every action has a way back and a window, or a warning shown before it.
  • Reliance: the evidence shown with each result is listed, and both reliance measures have an owner.
  • Shared decisions: the action tiers, the set’s contents and size, the cost per accepted result and the rollout’s stop numbers are agreed by both sides.
  • Engineering’s five questions: each has been asked and answered in writing.

Where does this map break?

It breaks in three places. First, in a company of five, one person holds every column. The map still helps there, as a list of decisions to make deliberately, and it stops helping as a division of labor.

Second, some rows move with the product. When engineers are the users, as with an internal tool, the engineering lead is the person who knows the user, and rows 1 to 6 are theirs. The sorting rule follows the knowledge of the user and ignores the job title.

Third, the map is my arrangement of the book’s material and the sources above. I found no study comparing teams that assign these decisions with teams that don’t, so I can’t tell you how much the page is worth. Pan’s essay argues from postmortems it doesn’t cite, and the Husain figures are one client’s.

A caveat on the measures in decision 6: opening the evidence isn’t the same as reading it. The measure catches results accepted in under a glance and misses the reviewer who opens everything and reads nothing.

Six lines to fill first

The short version of AI agents for product managers: if the kickoff is tomorrow and the page is blank, fill the six product lines and leave the rest marked open: done, tolerable errors, starting autonomy, unsure or wrong, undo, reliance. An agent with those six written can be built against, tested against and argued with. An agent without them will still ship with answers to all six, chosen by whichever default was nearest.

Chapter 24, in a line it quotes from a published essay on supervision interfaces, gives the test for the result: “A supervision surface becomes worthy when it reduces work without reducing responsibility.”

Chapter 24, “Agent UX and Human Trust,” develops trust calibration, the evidence pack and graceful failure (in the full book); Chapter 12, Oversight and Autonomy (in the full book) covers gates and the autonomy dial, and Chapter 14, Choosing Your Approach: When Not to Build an Agent (in the full book) the case for not building an agent. The agents at work guide places this post beside its neighbors, or you can see the formats.

Questions readers ask

What decisions does a product manager own on an AI agent product?
Six, in this guide's map: what counts as done and who checks it, which errors are tolerable and at what rate, how much autonomy each task starts with, what the user sees when the agent is unsure or wrong, what the user can undo, and how reliance is calibrated. Each one changes what a user is promised or can lose.
Should a product manager choose the model for an AI agent?
No. Model choice, prompts, tool design and where each block is enforced are engineering calls, because two answers can keep the same promise to the user and differ only in cost, speed or upkeep. The product manager asks what would make engineering change the model and what is re-tested when it changes.
How is a requirements document for an AI agent different from a normal PRD?
Acceptance becomes a rate on a named, versioned set of cases with an interval, where a deterministic feature has scenarios that pass or fail. The page also carries an error budget with a must-never list, a starting autonomy per task, the states a user sees when the agent fails, an undo map and the evidence shown with each result.
What is trust calibration in an AI product?
The book defines trust calibration as adjusting how much you rely on an automated system until the reliance matches the system's actual reliability. It fails in two directions: over-trust, where wrong results are accepted, and under-trust, where good results are redone by hand. A product should measure both.
Do product managers need to write evals for AI agents?
They need to own what the evaluation set says is correct: the cases, the written pass rule and the bar. Engineering builds the grader and the harness that runs it. Evaluation covers two of the six product decisions, done and tolerable errors, and leaves autonomy, failure states, undo and reliance to be decided elsewhere.

Sources

  1. Kazuo Okamura and Seiji Yamada, PLOS ONE 15(2) (2020). Adaptive trust calibration for human-AI collaboration
  2. Tian Pan (2026). The PRD for an AI Feature: Why Your Old Template Misses the Cliff
  3. Braintrust Team (2026). Evals are the new PRD
  4. Hamel Husain (2025). A Field Guide to Rapidly Improving AI Products
  5. IntelliAvatar, Hacker News (2025). Ask HN: How do you define "done" for long-running AI agents?