The AI agent use cases for startups that hold up in production share a shape more than an industry. Chapter 25 of the book catalogs eight such shapes, which it calls transversal recipes. Each comes with a verification signal, a starting level of autonomy and a known way of failing, and those three facts decide what to build first.
This post prints the eight as one table, walks through each with dated examples where I could verify one, and gives a seven-step test for placing an idea of your own. You should finish able to do three things: name the shape of your idea, name the check it depends on, and choose how much version one may do unattended.
Why do AI agent use cases for startups sort better by shape than by industry?
Shapes sort better because the loop an agent runs stays the same when the industry changes, and the loop is the thing you build, test and govern. The industry supplies the nouns: the documents, the systems, the jargon. A list of industries tells you where agents are being sold and very little about what will make yours work.
Chapter 25 opens with three professionals who will never meet: a litigator assembling a case file, a biologist surveying a protein family, an analyst sizing a market. An agent built for one of them would need new sources, new tools and new jargon to serve the other two. The loop would stay put: plan the sub-questions, search widely, read, synthesize, cite, verify.
The chapter names two units. A point-task is “a single, bounded request with a known output,” such as extracting the total from one invoice. A transversal recipe is a reusable shape of work that recurs across industries and roles, composes into an ongoing capability, and maps onto a few design patterns. Its test fits in a line: “swap the industry, and if only the nouns change, the recipe is transversal.”
Most pages that answer this search list tasks or departments: meeting prep, invoice chasing, support, investor updates. I read five of the top results on 6 October 2026. None attached a check to each use case, and none set a starting level of autonomy for each one.
Of the three facts, the check matters most to the chapter. In every one of its eight plans, it says, “the load-bearing element is the same: the verification signal”: the thing that tells you a run worked. One poster asked a forum in November 2025 for agents that scale in production outside coding and wrote, “I am struggling to see anything” (spacemnstr42069, Hacker News). The honest reply is a short list of shapes, each with its own state of proof.
What are the eight transversal recipes?
The eight transversal recipes are the Digital Coworker, Unstructured-to-Structured at Scale, the Queue Triage and Router, the Ambient Watcher, the Self-Maintaining Personal Knowledge Base, the Deep-Research Analyst, the Premortem and Red-Team Simulator, and Autoresearch. The names are the chapter’s. The table gives each one’s check, its starting autonomy and its usual failure.
The fourth column uses the autonomy slider for AI agents, which the book’s glossary files under the autonomy dial. It runs from a person approving each step (a human in the loop) to a person reviewing what a finished run produced (a human on the loop). The chapter’s rule is that a use case starts tight and earns its way loose, on evidence, and re-earns it on every model change.
Six columns come from the chapter’s text. The last borrows its three words from the brackets in the chapter’s figure, shown below, and I filed one row differently from that figure. It sorts the recipes by what the product hands over: a draft a person decides on, a deed in a real system, or work in a sandbox where every result is disposable. The buttons narrow the table by that label, which is the quickest way to shortlist AI agent use cases for startups that can ship as a draft first.
| Recipe | What it does | Verification signal | Where to start on autonomy | How it usually fails | Book chapters | Hands over |
|---|---|---|---|---|---|---|
| The Digital Coworker | Runs one documented, repeatable process end to end against real systems, and escalates edge cases to a person | An audit trail; escalation paths designed in advance; a closed scope | Per action: routine, reversible steps run free and consequential ones wait for a signature. Start internal and low-stakes | Deed-level autonomy on money, records or customers before the plumbing has proven itself | 25, drawing on 12, 14, 23 | deed |
| Unstructured-to-Structured at Scale | Turns messy documents into clean records, in bulk and continuously | Field-level citations; a confidence-gated lane to a person, with the threshold set per field by consequence | Acting, middle of the range: autonomy granted per field and earned per format | Silence at volume: a small misread, multiplied and posted into a system of record | 25, drawing on 2, 10, 12, 14, 15, 23, 26 | deed |
| The Queue Triage and Router | Classifies, prioritizes, enriches and routes each item in a queue | A sealed, held-out test set that gives the misroute rate; an abstain class; the false-confident rate against a baseline recorded before deployment | Enrichment and routing go loose early; auto-resolution is gated and earned | Miscalibrated confidence: the item that was wrongly closed, or routed where nobody looks, stays invisible | 25, drawing on 10, 23, 26 | deed |
| The Ambient Watcher | Sits on an event stream and wakes itself when something is worth acting on | Notify, question or review; proposals collect on a supervision board | Human on the loop: state-changing actions are proposed by default and run only from pre-approved tiers | It acts wrongly with nobody watching; its input is other people’s content | 25, drawing on 3, 12, 17, 18, 19, 20, 24 | deed |
| The Self-Maintaining Personal Knowledge Base | Keeps an interlinked wiki between its owner and the raw sources: ingest, query, lint | Every claim linked to a source file; a scheduled lint pass; spot-checks of the load-bearing pages; plain files with history | Assisted: the owner stays involved, one source at a time | Confident mis-integration, including “resolving” a contradiction that was the finding | 25, drawing on 8, 9, 10, 13, 23 | draft |
| The Deep-Research Analyst | Plans sub-questions, searches many sources in parallel, synthesizes, and delivers a cited report | Citations a reader can open, and a person who opens them | Advisory: the agent reads and drafts, a human decides | The report arrives convincing whether or not it is right | 25, drawing on 11, 23 | draft |
| The Premortem and Red-Team Simulator | Narrates a stipulated failure of a plan; a panel attacks an artifact; recorded traces are replayed against a candidate system | Findings deduplicated and ranked for a human decision; replays graded by judges | Advisory end of the slider | The premortem always finds a failure story; a clean simulator run is read as a guarantee | 25, drawing on 2, 10, 11, 15, 16, 17 | draft |
| Autoresearch | Changes one thing, runs the experiment, compares a number, keeps or reverts, and repeats overnight | A scalar metric computed without human judgment; one editable artifact; a fixed time box | Far loose end, unattended, and only where everything it can touch is disposable | Reward hacking: it optimizes the number and leaves the goal behind | 25, drawing on 6, 9, 10, 13, 16, 17, 22 | sandbox |
One label needs a note. I filed the knowledge base under draft because what its owner receives is pages to read, and that filing is mine. The chapter’s text describes it as an assisted tool with the owner involved at each step, which is the position the table uses.
The figure is the chapter’s own overview of the catalog. It draws the knowledge base at the loose edge of the deeds bracket, where the table says draft. On that one row I went by the chapter’s text. The chapter also warns, before it starts, that “These eight recipes are not equally proven,” so each section below ends with its maturity note.
Which recipes hand a person a draft?
Three recipes hand a person a draft: the Deep-Research Analyst, the Premortem and Red-Team Simulator, and (by this post’s filing) the Self-Maintaining Personal Knowledge Base. Nothing they produce changes a customer’s system. Their risk sits in whether anyone checks the words before acting on them.
The Deep-Research Analyst
The analyst answers an open question by planning sub-questions, searching in parallel, reading, and writing a cited report. Its position is advisory: “The agent reads and drafts; a human decides.” The report “arrives convincing whether or not it is right,” the chapter warns, so the label protects you only while a person really inspects.
Maturity splits in two. The chapter calls the loop “production-grade and widely shipped, while the verification remains the human-priced step,” and compresses the advice into one line: “Budget the inspector before you enjoy the analyst.” Any estimate of an AI agent’s ROI has to subtract that checking bill.
One dated example. Anthropic’s engineering write-up of 13 June 2025 describes a lead agent directing subagents, and reports that “multi-agent systems use about 15× more tokens than chats” (Anthropic, 2025). That is one team’s measurement at one date. I verified no second write-up for this recipe.
The Premortem and Red-Team Simulator
This recipe points the machinery at your own plan. A premortem stipulates that the launch failed and asks for the story of why; a red-team panel attacks the artifact through several lenses; a simulator replays recorded traces against a candidate system. All three stay at the advisory end of the slider, as input to a human decision.
The chapter names one failure per direction. The premortem “tilts pessimistic by construction,” because you asked for a failure story and will always get one. The simulator’s clean run tempts a misreading: “A clean bill from the simulator is evidence. It was never a guarantee.”
The technique is older than language models: Gary Klein described it in Harvard Business Review in September 2007. As one dated product example, a Show HN of 1 October 2026 offered agents that red-team a startup idea. A commenter, isawczuk, reported the next day: “I’ve tested on few fast growing startups with track record. Red-team still didn’t like the pitch.” That is the pessimistic tilt, described by a user.
The Self-Maintaining Personal Knowledge Base
Here an agent maintains an interlinked wiki that sits between its owner and the raw sources. It has three verbs: ingest a source, query the wiki, and lint it for contradictions and stale claims. The checks are source links on every claim, a scheduled lint pass, and spot-checks of the pages that carry weight.
The failure is structural. The owner now reads the wiki in place of the sources, so a confident mis-integration is read straight past. The chapter borrows a sentence from an earlier one: “a system that learns while you sleep can mislearn while you sleep.”
One example, and it is a pattern file with no product behind it: Andrej Karpathy’s “llm-wiki” gist, retrieved 6 October 2026. It says “Humans abandon wikis because the maintenance burden grows faster than the value.” I verified no second example and no startup product. The chapter’s maturity note agrees with that thin record: assisted memory ships, and “the unattended librarian is still a direction of travel.”
Which recipes act on real systems, and behind what gates?
Four recipes act on real systems: Unstructured-to-Structured at Scale, the Queue Triage and Router, the Ambient Watcher and the Digital Coworker. Each buys its autonomy with a gate placed where a wrong action would cost something. They are also where most of the verified production write-ups are.
Unstructured-to-Structured at Scale
This recipe is a pipeline that turns invoices, claims, contracts and transcripts into records with named fields. The chapter says the machinery “is barely an agent at all”: a chain that classifies the document, extracts the fields, validates them and routes the record. It means that as praise, since enumerable steps make a flowchart you can draw.
Two checks are required in practice. Every extracted value carries a pointer to its place in the source, and records the pipeline is unsure of queue for a person. The failure is “silence at volume,” and the rule against it is absolute: “high-stakes fields never auto-post without a confidence gate, however good the accuracy number looks, because the accuracy number is an average and the ledger does not experience averages.”
Two dated examples. In a Launch HN of 9 October 2025, the founders of Extend said they launched with APIs “to parse, classify, split, and extract documents” and built a layer that reviews “low confidence OCR errors” (kbyatnal, Hacker News). Uber’s engineering blog of 17 April 2025 describes a review screen offering “a side-by-side comparison of the PDF data versus the data extracted from the models,” and reports “an impressive overall accuracy rate of 90%” (Uber Engineering, 2025).
Both are the companies’ own statements, and the second number is exactly the kind of average the chapter warns about. The chapter’s maturity note for this recipe is its shortest: “shipping, broadly, today.”
The Queue Triage and Router
The triage agent sits in front of a queue of tickets, alerts, leads or forms. For each item it classifies, prioritizes, enriches with context, and routes. The chapter’s summary: “The dispatcher, that is, and deliberately not the doer.”
Its instrument is a classifier with an abstain class, tested on a held-out set that stays sealed until scoring day. That gives you a misroute rate before the queue depends on it. The failure is asymmetric: a wrong escalation costs a specialist a shrug, and a wrong auto-close is found weeks later, if ever. So the chapter splits the autonomy: “Enrichment and routing are reversible and go loose early, while auto-resolution is a deed to be gated and earned”; resolution waits for a measured record.
Two dated examples, both from incident response inside large engineering organizations. Microsoft’s Azure blog of 6 March 2025 says its local triage system had been “in production in Azure since mid-2024” and that “As of Jan 2025, 6 teams are in production” (Microsoft Azure, 2025). Kiro’s post of 21 August 2026 reports that “96.9% of the agent’s tool calls are reads (safe, reversible, parallelizable); the rest can be gated” (Kiro, 2026). The chapter says this recipe “ships in production today” and counts it, with extraction, as the workhorse of enterprise deployments.
The Ambient Watcher
A watcher removes the person from the trigger. It listens to a stream of tickets, alerts, filings or mail and wakes itself. The person moves from in the loop to on it, and the chapter adopts three ways for the agent to reach that person: notify, question, and review.
The failure is the premise turned around: “A system that acts without being asked can act wrongly without being watched,” in the chapter’s words. The stream is also other people’s content, so a watcher holding private data and tools assembles the lethal trifecta in routine operation; the lethal trifecta audit checks a design for it. The standing posture is conservative: “state-changing actions are proposed by default” and run only from pre-approved tiers that are bounded, reversible and low-stakes.
Two dated sources, one a definition and one a deployment. LangChain’s post of 14 January 2025 defines ambient agents as ones that “listen to an event stream and act on it accordingly” (Chase, 2025). LaunchDarkly’s post of 3 September 2026 describes a triage agent that “watches for incoming alerts” and a validator agent that “either approves it or rejects it and kicks it to a human” (LaunchDarkly, 2026). The chapter rates the recipe “production-grade in narrow forms,” naming support front doors, alert enrichment and inbox assistants.
The Digital Coworker
The coworker runs one documented process end to end against real systems: it reads the inputs, applies the policy, takes the actions, keeps the audit trail and escalates the odd cases. The premise is that “a process documented well enough for a new hire is documented well enough for an agent.” The scope is strict, “A process, mind, and never a department,” as the chapter puts it.
Two of the earlier acting recipes grow into this one. An extraction pipeline that starts posting its records, or a triage agent that starts resolving its tickets, has crossed from analysis to deeds, and the crossing is what demands governance. Autonomy is set per action: “the password reset runs free while the account closure waits for a signature, inside one and the same coworker.” The consequence tier classifier sorts a list of actions that way.
The failure the chapter names is “a pilot that impressed in a demo being promoted straight to the general ledger” with its escalation paths undesigned. Its maturity note is the one I would show an investor: “real and growing, gated by governance and integration readiness rather than by anything a better model would fix.”
One dated example, a large company’s account of its own first month. Klarna’s press release of 27 February 2024 said its assistant “has had 2.3 million conversations, two-thirds of Klarna’s customer service chats” (Klarna, 2024). I did not verify how that deployment went afterwards, and I verified no second example.
Which recipe may run unattended?
Only Autoresearch is designed to run unattended, and only inside a sandbox. The agent changes one thing, runs the experiment, compares a number, keeps or reverts, and repeats through the night. The chapter grants that freedom on one condition: “everything it can touch is discardable.”
Three design choices carry over to any field: one editable artifact, a scalar metric computed without human judgment, and a fixed time box per experiment. The whole weight rests on the metric. “The loop is exactly as trustworthy as its metric,” the chapter says, and a weak proxy invites reward hacking, where the system optimizes the number and abandons the goal.
A dated example: the README of Andrej Karpathy’s autoresearch repository (retrieved 6 October 2026), in which the agent “modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.” The chapter lists where else the shape fits: prompt search, copy tested against conversion, query plans tested against latency, and your own agent’s prompts once an eval set exists to serve as the metric.
The maturity note is blunt: the recipe “ships today in narrow, well-instrumented forms and is aspirational everywhere else.” My reading for a startup is that this one is usually an internal tool. The chapter’s lesson travels further than the recipe: “full autonomy is a fact about the sandbox, whatever the model’s competence.”
How do you place your own idea on the table?
Place an idea by running it through seven questions, in order, and stopping at the first one it fails. The test below is this post’s construction for sorting AI agent use cases for startups one idea at a time. Each step leans on a line from Chapter 25 or Chapter 14, and the wording of the steps is mine.
- Task or department? State the idea as one process with an input and an output. Chapter 14 calls “Automate our customer support” a “department wearing the grammar of a task”; split a department into tasks and test each one.
- Can you draw the flowchart first? If every branch can be written without judgment, build plain code and stop. If the steps are enumerable and a few need judgment, you have a workflow, which may still be one of the eight.
- Swap the industry. If only the nouns change, the shape is transversal and may be one of the eight. If the loop itself changes, it is a custom build: go to step 6 and find its signal yourself. A single, bounded request with a known output stays a point-task however the swap goes.
- Draft, deed or sandbox? Decide whether the product hands a person something to decide on, changes a real system, or works only on things you can throw away. Mixed products get an answer per action. The table’s last column shows how far each recipe reaches: a watcher that only notifies, or a router that only routes, is a draft-first version of a deed row.
- Which recipe, or which chain? Name the row in the table. If the idea chains two rows, name both and mark the consequential act.
- Does your domain supply the signal? Write the recipe’s verification signal in your customer’s terms. If the domain cannot supply it, the idea is not yet a product.
- Start tight, and write down what loosens it. Take the row’s starting autonomy for version one, and name the measured evidence that would move it. If the row’s condition fails in your domain (for Autoresearch, anything it touches that is not disposable), start at the gated setting: propose and wait.
Here is one idea run through it. The pitch is “an agent that handles maintenance for property managers.” Step 1 stops that sentence, since it names a department. The task inside is narrower: read each tenant request, judge its urgency, and get the right trade booked.
Step 2: I can draw the flow, and two boxes need judgment (how urgent, which trade), so it is a workflow and stays on the table. Step 3: swap tenants for patients or drivers and only the nouns change. Steps 4 and 5: it is a chain, a Queue Triage and Router feeding a Digital Coworker, and the booking is the consequential act because it spends an owner’s money.
Step 6: the triage half needs a sealed set of past requests with known outcomes, and a property manager’s ticket history can supply one. Step 7: version one classifies, enriches and routes by itself, and proposes each booking for one click. A measured misroute rate, a false-confident rate on urgent requests and an audit trail of approved bookings are what would loosen it.
That is one answer: a triage product first, with the coworker half gated. The chapter’s rule for chains says the same thing: “the gate belongs at the consequential act, wherever in the chain it lands.” I ran two more ideas while drafting; a supplier price-sheet reader landed cleanly on Unstructured-to-Structured, and a contract-renewal reminder stopped at step 2 as plain code.
The test is also a short form of what AI agents ask of product managers: a named check and a named gate before a launch date. Three tools go deeper on single steps. The Should this be an agent? decision aid covers steps 1 and 2, the agent verifiability scorecard covers step 6, and the consequence tier classifier linked above covers the gates in step 7.
What does the evidence say about agents in production?
The evidence I could verify says deployed agents are mostly short-leashed and checked by people, and that each number describes a particular sample at a particular date. Three sources are worth a founder’s time. None of them is a census.
An academic study. Pan and colleagues ran 20 case studies and surveyed 86 practitioners with deployed systems across 26 domains. Version 4 of the paper (June 2026) reports that “68% execute at most 10 steps before human intervention” and “74% depend primarily on human evaluation” (Pan et al., 2025). In that sample, a tight first version is the norm.
A vendor’s survey of its own audience. LangChain collected 1,340 responses between 18 November and 2 December 2025. Of those respondents, 57.3% said they had agents in production; as the primary use, 26.5% named customer service, 24.4% research and data analysis, and 18% internal workflow automation (LangChain, 2025). The respondents chose to answer, 63% worked in technology, and 49% came from organizations under 100 people.
One model developer’s usage study. Anthropic’s analysis of 18 February 2026 covers tool calls on its own public API. Software engineering made up “nearly 50%” of them, and the study reports that “73% appear to have a human in the loop in some way, and only 0.8% of actions appear to be irreversible” (Anthropic, 2026). It is one provider’s traffic, and the study notes that its method tends to overestimate human involvement.
Forum testimony points the same way, and it is testimony. One builder wrote in July 2025 that the agents still working in production “do one specific thing” (rudderdev, Hacker News). Another wrote in July 2026 that “the reason demos die in production isnt capability, its trust” (jacksonxly, Reddit).
Taken together, this is weaker than it looks. The samples are self-selected or come from one provider, the dates are 2025 and 2026, and none of it measures startups as a group. What it supports is modest: the AI agent use cases for startups with the most public proof are the tightly checked ones.
When is the answer “no agent”?
The answer is “no agent” in four cases: the idea is a department, the logic can be written out in full, the domain lacks the signal, or checking costs as much as doing. Each one shows up early in the seven-step test, and each one saves a quarter.
A department comes first. Until it is decomposed, Chapter 14 says, any architecture is a guess about an undefined problem. Fully specified logic comes second: a threshold on a database column is an if statement, and the chapter’s litmus is “can you draw the flowchart before the request arrives?”
That litmus is the core of the AI agent vs automation question. Two of the eight recipes, extraction and triage, are a workflow inside, and Chapter 25 calls that a compliment in extraction’s case. The sibling post on the agent-or-workflow decision rule works through it with an example, and what agentic AI means covers the label.
Both remaining cases concern the check. A recipe whose signal is missing in your domain is a plan for different ground, and the chapter concedes that “no book of plans knows your soil.” If checking is as dear as doing, Chapter 23 gives the verdict: “An agent whose output costs as much to verify as to produce by hand has a return of approximately nothing, however impressive the demo.” The glossary calls that distance the verification gap.
Is an AI agent startup just a wrapper?
The book cannot answer that, because it offers no theory of moats or defensibility, and I will not invent one for it. What Chapter 25 says is narrower. Building a recipe “is mostly configuration, plus the unglamorous integration and evaluation work where the real effort always lands.” That sentence tells you where the work is and makes no promise about who can copy it.
Where does this catalog stop being useful?
The catalog stops being useful at four edges: ideas outside the eight, questions of money, the passage of time, and the thin spots in the examples. I would rather name them than have you find them.
Coverage. Eight shapes are one chapter’s pattern book. An idea that fits none of them may still be sound; the test then leaves you with the signal question and no row to borrow from. Software engineering is the largest category in the usage study above, and coding agents have their own chapters in the book, outside this table.
Money. A recipe tells you what must be true for the product to work and nothing about the bill. The sibling post on AI agent pricing models covers cost per task and how to charge. The catalog is also silent on whether to build or buy an AI agent, since a recipe describes what the system must contain whoever assembles it.
Time. The maturity notes are the chapter’s judgment when it was written, and the examples carry dates for a reason. The slider positions are starting points that move on evidence.
Examples. The analyst, the coworker, the knowledge base and Autoresearch each rest on a single verified example here, and the knowledge base’s example is a pattern file. Several examples come from large companies, written by the teams that built them. For agentic AI for business leaders the usable part is the question each example answers: what did they check, and where was the gate?
The one thing to keep
Among AI agent use cases for startups, pick the one whose check you can name. A commenter on an agent builders’ forum put the bar well in August 2026: “the agent is ready when its worst failure mode is bounded, not when the average demo looks good” (CODE_HEIST, Reddit). The chapter’s own version is shorter: “A demo needs one run to succeed; a product needs every run to.”
When to build an AI agent product, then, comes down to three written lines: the recipe, its signal in your domain, and the tight setting version one starts at. A good AI agent MVP is often the draft half or the reversible half of the product you eventually want. Chapter 25 closes on the sentence that explains why: “A recipe is only as transversal as that signal can be found in every industry it visits.”
The eight recipes, with their machinery and figures, are in Chapter 25, “Transversal Recipes”, in the full book. Chapter 1, “What Is an Agent?” is free to read online and introduces the verification question the whole catalog rests on. The Agents at work guide collects the related posts and tools, and you can see the formats.
Questions readers ask
- Which AI agent use cases work in production for startups?
- The shapes with the most verified, self-described production write-ups are document extraction into structured records, queue triage and routing, and narrow event-triggered watchers. Research agents ship widely with a person checking the report. The book calls the end-to-end digital coworker real and growing, gated by governance and integration readiness.
- What is a good first AI agent MVP?
- A recipe whose output is a draft a person decides on, or the reversible half of an acting recipe: enrich and route a queue before resolving anything in it. The book reports the field's rule as start internal and low-stakes, and prove the governance where a bad week costs apologies.
- Is my AI agent startup just a wrapper?
- The book has no moat theory, so it cannot settle that. What Chapter 25 does say is that building a recipe is mostly configuration plus integration and evaluation work, and that the verification signal has to be found again in each domain. That is where the work is.
- How much autonomy should version one of an agent product have?
- Start at the recipe's tight setting and write down the evidence that would loosen it. The chapter sets autonomy per action and per recipe, moves it on evidence in both directions, and requires it to be re-earned on every model change.
- AI agent vs automation: when is plain automation enough?
- When you can draw the flowchart before the request arrives and write every branch without judgment, plain code is enough. When the steps are enumerable and a few need judgment, a workflow is enough. Two of the eight recipes are workflows inside by the chapter's own account.
Sources
- Melissa Z. Pan et al. (2025; v4 2026). Measuring Agents in Production (arXiv:2512.04123, v4)
- LangChain (2025). State of Agent Engineering
- Anthropic (2026). Measuring AI agent autonomy in practice
- Anthropic (2025). How we built our multi-agent research system
- Harrison Chase (LangChain) (2025). Introducing ambient agents
- LaunchDarkly (2026). Building a self-driving ops triage loop
- Microsoft Azure (2025). Optimizing incident management with AIOps using the Triangle System
- Kiro (2026). How We Learned to Trust an AI Agent to Triage Production Incidents
- Uber Engineering (2025). Advancing Invoice Document Processing at Uber using GenAI
- kbyatnal (Hacker News) (2025). Launch HN: Extend (YC W23) – Turn your messiest documents into data
- Klarna (2024). Klarna AI assistant handles two-thirds of customer service chats in its first month (press release)
- Andrej Karpathy (retrieved 2026). autoresearch (repository README)
- Andrej Karpathy (retrieved 2026). llm-wiki: a pattern for building personal knowledge bases using LLMs (gist)
- Gary Klein (Harvard Business Review) (2007). Performing a Project Premortem
- ahoskins; comment by isawczuk (Hacker News) (2026). Show HN: Premortem – AI agents that red-team your startup idea
- spacemnstr42069 (Hacker News) (2025). Ask HN: Are Agents Just Hype?
- comment by rudderdev (Hacker News) (2025). The current hype around autonomous agents, and what actually works in production (discussion)
- MagicitePower; comment by jacksonxly (Reddit) (2026). How to create an ai agent that actually does something useful, not just a demo? (r/AI_Agents thread)
- Consistent_Dress3064; comment by CODE_HEIST (Reddit) (2026). Agent builders: what actually got your agent from cool demo to reliable in production? (r/AI_Agents thread)