What are the main use cases for AI agents?
The main use cases for AI agents are a small number of shapes of work that repeat across industries, and Chapter 25 of the book names eight of them. The chapter calls each one a transversal recipe. A litigator assembling a case file, a biologist surveying a protein family and an analyst sizing a market use different words for what they do, and all three are running the same loop: plan the sub-questions, search widely, read, synthesize, cite, verify.
Most lists of use cases are lists of tasks, such as drafting one reply or summarizing one document. The chapter calls each of those a point-task, “a single, bounded request with a known output”, and is blunt about its limits: “none of them teaches you anything reusable.” A recipe sits one level up. It is “a reusable shape of work with three properties.” It recurs across industries, it composes into an ongoing capability, and it maps onto a few design patterns from the rest of the book.
The eight recipes are autoresearch, the self-maintaining personal knowledge base, the deep-research analyst, the ambient watcher, unstructured-to-structured at scale, the queue triage and router, the digital coworker, and the premortem and red-team simulator. The tool above helps you decide which of them your work is, and what the chapter says must be in place before it runs.
How do I tell a reusable recipe from a one-off task?
You tell a reusable recipe from a one-off task by swapping the industry and checking what changes. The chapter states the test in one sentence: “swap the industry, and if only the nouns change, the recipe is transversal.” If the sources, the tools and the jargon change and the loop stays the same, you have a recipe.
The first part of the tool asks two questions built from those definitions. Their wording is the tool’s own. If you say the work is one bounded request with a known output, the tool tells you it is a point-task and suggests you build the single call. The should-this-be-an-agent tool covers that earlier choice between a script, a model call, a workflow and an agent.
Passing the test does not finish the job. The chapter compares itself to a builder’s pattern book and marks where the comparison stops: “no book of plans knows your soil.” Adapting a recipe to your domain is still engineering work.
Which recipe fits my work?
The recipe that fits is the one whose defining trait matches the work, and the second part of the tool turns each trait into a statement you can tick. The eight statements and the recipe each one opens are the tool’s own. Each card restates Chapter 25 and quotes it.
| If this describes the work | The recipe | Band in the book’s figure |
|---|---|---|
| A number a script can compute tells you whether a change helped, and everything touched is disposable | Autoresearch | Sandbox only |
| You keep a private pile of sources and want it kept synthesized | Self-maintaining knowledge base | Deeds, at its loose edge |
| The deliverable is a cited report on an open-ended question | Deep-research analyst | Drafts |
| Nobody presses go, and events arrive on a stream | Ambient watcher | Deeds |
| Messy documents go in and records with named fields come out | Unstructured-to-structured | Deeds |
| A backlog needs classifying, prioritizing, enriching and routing | Queue triage and router | Deeds |
| A process is documented well enough for a new hire to run | Digital coworker | Deeds |
| You want a plan or an artifact attacked before it launches | Premortem and red-team simulator | Drafts |
You can tick more than one. Each card answers the five questions the chapter asks of every recipe: what it is, why it recurs everywhere, which of the book’s machinery it assembles, where its slider naturally sits, and how it most often fails. The card also gives the chapter’s maturity note in its own words, because the eight are not equally proven.
The chapter explains why the names are worth having: “What the catalog adds is names, and names are operational”. Once you can say that a problem is a triage problem, you know which chapter’s machinery to assemble and which safeguards it needs.
What is the autonomy slider?
The autonomy slider is the answer to one question every recipe has to settle first: how much may this system do between moments of a human’s attention? It runs from a person in the loop, approving as the work goes, to a person on the loop, reviewing what an autonomous run produced. It is the same instrument the book builds in Chapter 12 as the autonomy dial, seen from the product side.
The figure places the eight recipes along the slider in three bands. Recipes that commit deeds on real systems sit toward the tight end, behind gates. Recipes that hand a person a draft sit further along. One recipe sits alone at the far end. The tool redraws that order as a strip and marks the recipes you ticked.
A position on the slider is where a recipe starts. Of any use case, the chapter says: “It starts tight, and it earns its way loose”. The figure’s caption repeats the point: “The positions are starting points; each deployment moves along the slider on evidence”. The chapter also names what looser settings cost: “the speed of verification is what the slider’s loose settings are purchased with”. The consequence tier classifier sets the gate for a single action, and the pass@k calculator shows the gap between a run that works once and a run that works every time.
Why can only one recipe run unattended?
Only autoresearch can run unattended, and the reason is that everything it touches can be thrown away. The agent changes one thing, runs an experiment, compares a number, keeps or discards the change, and repeats through the night with no person in the cycle.
The chapter picks out three design choices that make the published example trustworthy: a single editable artifact, a scalar metric that needs no human judgment, and a time-boxed cycle. A fourth condition holds the others together. The loop works in a scratch environment where the worst outcome is a wasted night. The chapter draws the general lesson from it: “full autonomy is a fact about the sandbox, whatever the model’s competence.”
The tool’s checklist for this recipe has those four items. If you answer that something the loop can touch is not discardable, the verdict changes to say the recipe does not belong at the far end of the slider. That rule is the tool’s own reading of the chapter’s sentence.
The recipe’s weakness is the metric itself. “The loop is exactly as trustworthy as its metric.” A metric that can be gamed invites reward hacking, overnight, with nobody watching.
What safeguards does each recipe need?
Each recipe needs the specific safeguards its chapter section names, and each card the tool opens lists them as a checklist. You answer yes, no or not yet known. The tool counts what is in place and names what is missing. It does not invent a passing score, and it does not say a recipe is ready.
Every item on a checklist is something Chapter 25 states. Which items appear, and the count over them, are the tool’s own arrangement. Three examples show the range.
For unstructured-to-structured extraction the chapter calls two features non-negotiable: field-level citations, and a confidence-gated lane that sends uncertain records to a person. Its rule for the second is absolute: “high-stakes fields never auto-post without a confidence gate, however good the accuracy number looks, because the accuracy number is an average and the ledger does not experience averages.” The cascade cost estimator prices a cheap first pass in front of a careful second one.
For queue triage the safe default is escalation. The chapter says “in triage, escalate to a human must be the failure mode.” It also warns that accuracy alone hides the dangerous errors: “you must monitor the false-confident rate (how often the system was sure and wrong) rather than the accuracy number alone”.
For the ambient watcher the standing rule is that actions which change state are proposed by default. The chapter’s reason is that a watcher reads other people’s content all day, and “an injection payload does not need to lure the agent anywhere, because the stream delivers it.” The lethal trifecta audit checks what such an agent can read, hold and send.
Can recipes be combined?
Recipes can be combined, and the chapter says real systems do combine them. Its example is a support pipeline in which a watcher feeds the triage and the triage feeds a coworker, so one pipeline spans three recipes. When you tick two recipes that the chapter joins, the tool shows a composition note.
The chapter treats the digital coworker as what two other recipes grow into: “an extraction pipeline that starts posting its records and a triage agent that starts resolving its tickets have each crossed the line from analysis to deeds”. That crossing, and not the model, is what calls for governance. The coworker’s own premise is “a process documented well enough for a new hire is documented well enough for an agent.”
A second pair runs the other way, toward reading. The deep-research analyst feeds the knowledge base, and the chapter’s advice is short: “if you are going to pay for the analyst, give it a filing cabinet.”
Combining recipes raises the stakes. In the chapter’s words, “composition compounds the stakes: the gate belongs at the consequential act, wherever in the chain it lands”. The multi-agent failure explorer covers what goes wrong at the handoffs once several agents are involved, and the ROI versus risk map helps decide which candidate to build first.
What should I do with the result?
You should take the card for your recipe to the chapter it points to, and treat the missing safeguards as the work to do before the first run. The Copy as Markdown button gives you the cards, the checklists with your answers, the counts and any composition notes in one block for a design document.
The tool cannot tell you whether your metric is honest, whether your reviewers really inspect, or whether your process is as well documented as you believe. Those are judgments about your own system. The guide to agents in practice places this chapter among the other application chapters. The full catalog, with the sources and the examples behind each recipe, is Chapter 25, Transversal Recipes: Big Patterns That Cut Across Industries (in the full book).
Questions readers ask
- Why is autoresearch the only recipe allowed to run unattended?
- Because everything it touches is disposable. Chapter 25 describes a scratch repository, a throwaway model and experiments whose worst outcome is a wasted night. The chapter's conclusion is that full autonomy is a fact about the sandbox and not about how capable the model is. Every other recipe either hands a person a draft or acts on real systems behind gates.
- What is the difference between a point-task and a recipe?
- A point-task is a single, bounded request with a known output, such as one reply or one extraction. A transversal recipe is a reusable shape of work that recurs across industries, composes into an ongoing capability, and maps onto a few design patterns. The chapter's test is to swap the industry and see whether only the nouns change.
- Which of the eight recipes are production-grade today?
- Chapter 25 says unstructured-to-structured extraction ships broadly and queue triage ships in production, and that the deep-research loop is widely shipped while its verification stays a human cost. The ambient watcher and autoresearch ship in narrow forms. The knowledge base ships as an assisted memory. The digital coworker is real and growing, limited by governance and integration. The premortem and the red-team panel are available to anyone today, while the deployment simulator needs recorded traffic worth replaying.
- Can one system use more than one recipe?
- Yes, and the chapter says real systems do. Its example is a support pipeline that spans three recipes: a watcher feeds the triage and the triage feeds a coworker. The chapter adds that composition raises the stakes, and that the approval gate belongs at the consequential act, wherever in the chain that act lands.
- Does a full checklist mean the recipe is safe to run?
- No. The checklist collects the safeguards Chapter 25 names for a recipe, and the choice of which ones to list is the tool's. A full checklist means the starting conditions are met. The chapter says a use case starts tight and earns its way loose on evidence, and each card still names how that recipe most often fails.
Sources
- Andrej Karpathy (2026). llm-wiki: a pattern for building personal knowledge bases using LLMs
- Andrej Karpathy (2026). autoresearch (repository README)
- Harrison Chase (2025). Introducing ambient agents
- Gary Klein (2007). Performing a Project Premortem