For a backend developer, an AI agent is a distributed system with one nondeterministic component in its control path. Queues, retries, idempotency keys, checkpoints and tracing carry over. Three habits change (how you test, debug and audit), and one trust boundary is new. This guide to AI agents for backend developers sorts all of it, row by row.
The subject here is building a system with a model in the loop. Prompting an assistant to write backend code is a different subject. When I ran a web search for this phrase in October 2026, the results included a product directory, a ranked list of tools and an agency’s page on that other subject. After this page you can sort an agent ticket into three piles: work your existing habits cover, work where one habit changes, and a boundary you have not defended before.
You need no machine-learning background to follow it. The two sentences I lean on hardest come from Chapters 1 and 2 of AI Agents, Engineered, which are free to read online. Chapters 3 and 15 to 20, quoted in the table, are in the full book.
What is new when a backend developer builds an AI agent?
One component is new: a model whose output is a sample from a distribution, sitting where your code used to choose the next call. Everything around it is code you write, and the book’s name for that code is the harness. The page on what an agent harness is lists its parts, so I name only the split here.
Chapter 3 (in the full book) states the division of labor in a sentence it asks you to carry out of the chapter: “the model supplies the judgment—which tool, with what arguments, and whether the job is done—and your code supplies everything else: the hands, the memory, and the stop button”. The same chapter is blunt about how little machinery that takes: “An agent has exactly four parts, and you already understand every one of them.”
The four parts are a model client, a set of tools, a message history and a loop. The post on what an agent loop is walks one pass through them. One person asking why agent tooling is not written in backend languages described the work in December 2025 in terms you will recognize: “the core work seems to be API orchestration, state management, retries, tool calling, and I/O—not numerical computing or data science”.
How nondeterministic is the new component?
It varies between identical calls, and one published experiment found variation even at the setting meant to remove it. In September 2025 a research lab sent one prompt 1,000 times to an open-weights model at temperature zero. It got 80 distinct completions, and the most common appeared 78 times.
That is one model, one prompt and one lab, at one date. The same write-up reports all 1,000 completions coming out identical once the team changed the inference kernels, which places the cause in the serving stack. Chapter 2 retells the experiment and then says what a builder should do about it.
Which camp is right, “it’s a distributed system” or “you can’t test it”?
Both are right, about opposite sides of one line. The first camp is stated at full strength in a comment from April 2026: “A distributed systems problem with non-deterministic run lengths.” A 2026 essay on multi-agent systems says of agent failures, “These are 30-year-old bugs wearing a new hat.” The book agrees for the code you wrote, and in Chapter 19 calls an agent’s chain “a distributed system whose remote calls happen to be inferences”.
The second camp has a 2026 essay on releasing model-backed features: “merge, deploy, watch. That works for code. It fails for LLMs.” The same author wrote an essay on long-running agent tasks that the book quotes for the first camp’s view: applying those patterns to agents “is less about AI and more about taking distributed systems seriously”.
The line between the two is the book’s own term, from Chapter 15: the deterministic–nondeterministic seam. On one side is everything you wrote: tools, parsers, loop control. On the other is the model’s output and everything derived from it. For AI agents for backend developers, that seam is the whole sort: the mapping table covers the first side, and the three broken habits live on the second.
Is the thing on your ticket an agent at all?
Often it is a workflow: model calls wired along a path your code chose in advance. The author of the 12-Factor Agents notes describes products billed as agents that way: “A lot of them are mostly deterministic code, with LLM steps sprinkled in at just the right points”. Nobody has measured that share, and I put no number on it.
In a workflow the control path is yours, so the state, retry and recovery rows hold with no adjustment. The testing rows still change, because the model’s output still varies. The rule for telling the two apart is in agents vs workflows, and the page on what agentic AI means covers how the label gets stretched.
AI agents for backend developers: which habits carry over?
Thirteen habits in the table below carry over, some with nothing to add and some with one adjustment, and the last three rows are the habits that break. The table is my construction. Each row is tied to one sentence of the book, quoted exactly, and to the page that goes deeper.
Chapters 1 and 2 are free to read, and the other chapters in the table are in the full book. Pick a ticket type to narrow the rows; all of them stay in the page.
| Backend habit | What carries over | What changes | The book’s sentence | Chapter | Read next | Ticket |
|---|---|---|---|---|---|---|
| A stateless service, with state in your own store | The model call is stateless and you own the state | The transcript is the component’s whole memory, and every pass sends it again in full, so state design is also cost design | “the conversation belongs to you, not to the provider” | Ch. 3 | What an agent loop is | long-running job, on-call |
| Retry with backoff and jitter, bounded | The whole policy | Classify before you retry. A transient read failure is retried; a model’s malformed output goes back to the model as an error | “retry only the transient class, since retrying a bad credential or a model’s malformed output burns money on a certainty” | Ch. 18 | Stop conditions for agent loops, on which errors go back to the model | tool integration |
| A deadline on every call | All of it | Nothing. A timed-out read is retried, and a timed-out write is as ambiguous as it always was | “a timeout on a write is the ambiguous case—the work may have happened—and the next section’s idempotency discipline is the only thing that makes retrying it safe” | Ch. 18 | Run the loop, a tool with both timeouts in it | tool integration |
| Idempotency keys on mutating calls | The key, the stored first result, the downstream contract | The key comes from the logical action (run, step, tool, arguments), with an intent record before the call and a receipt after it | “A checkpoint records position. It does not, and cannot, record whether the outside world did the work.” | Ch. 18 | Idempotent tools and safe retries | tool integration, long-running job, on-call |
| A circuit breaker around a flaky dependency | Closed, open, half-open | It trips on failures only. A lookup that finds nothing has answered | “a pattern borrowed from electrical panels by way of the distributed-systems literature” | Ch. 18 | Circuit breaker simulator | tool integration, on-call |
| Checkpoints, journals and sagas | A checkpoint keyed by run; compensations run in reverse | Replay needs every model result and tool result recorded, plus the versions each step depended on | “The code replays; the world does not.” | Ch. 18 | The AI agent as a state machine, on resuming after a crash | long-running job |
| Accept the request, queue the work, return a ticket | All of it | Nothing | “the patterns are the ones distributed-systems engineers have used for decades” | Ch. 20 | AI agent architecture guide | long-running job |
| Distributed tracing | Spans and trace trees | A span also carries what a replay needs: the full prompt, the sampling parameters, the model’s identity and the tokens returned | “its two nouns map onto agents almost without alteration” | Ch. 15 | What an agent span should carry | on-call |
| Validate input at the boundary | Parse, schema and range checks, now applied to the model’s output too | A valid shape says nothing about a correct value | “passing the schema proves the output is well-formed and nothing more” | Ch. 18 | What tool calling is, step by step | tool integration, tests and CI |
| Catch the exception, log it, return a default | The catch and the log | The failure goes back to the model as a structured result. An empty default is the worst choice, because the model reasons over the hole | “Nothing crashed. Every status code on your dashboard reads as success. The failure has been laundered into an answer.” | Ch. 18 | AI agent failure modes, under the swallowed error | tool integration, on-call |
| An explicit state machine for a lifecycle | Named states and legal transitions | Fixed states cost some of the adaptability you wanted a model for, so production designs sit between the two | “a deterministic skeleton of states with model judgment doing the deciding inside each one” | Ch. 18 | The AI agent as a state machine | long-running job |
| Least privilege | All of it | It matters more, because text the model reads can ask for the privilege | “grant the minimum access the task requires, and nothing on standing” | Ch. 17 | Sandboxing agent tool execution | security review, tool integration |
| Keep data out of the command channel | The instinct | No structural fix exists for text a model reads; see the trust boundary section below | “The difference is that SQL injection had a complete fix waiting to be adopted.” | Ch. 17 | The lethal trifecta, explained | security review |
| Breaks: tests that assert exact output | Exact assertions on everything you wrote, against a scripted model | Live model output gets property assertions and a pass rate over several runs | “That side is ordinary software, and it deserves ordinary software’s tests, exact assertions and all.” | Ch. 15 | How does testing change? | tests and CI |
| Breaks: reproduce a bug by running it again | Stepping through your own code, once a run is recorded | A rerun takes a different path, so model and tool results are recorded as they happen and replayed later | “expect the failure you saw once to resist reproduction on demand” | Ch. 2 | How does debugging change? | on-call, tests and CI |
| Breaks: read a success status, or the component’s own “done”, as success | Health checks on the code you wrote | A program checks the outcome, and a release is judged by comparing distributions | “The model’s “done” is testimony; a green test is evidence.” | Ch. 3 | How does auditing change? | on-call, long-running job |
Why is retrying without an idempotency key missing from the breaks?
It is missing because it was unsafe before agents existed, and the book treats the cure as long-established practice that carries over. A 2017 post on a payment company’s engineering blog, one example of the category, defined the goal as endpoints that “can be called any number of times while guaranteeing that side effects only occur once”. Chapter 18 asks for the same idempotency key, built from the logical action.
What an agent adds is more places for a repeat to start: the transport, the model asking again, a resumed run. The idempotency post in the table counts them and works one crash through. I found no first-person public report of an agent’s retry charging a customer twice; essays and library authors name it as the failure they guard against.
What breaks when one component is nondeterministic?
Three habits break, and Chapter 1 (free to read) names them in one sentence about the costs an agent adds: “Nondeterminism, third: the same input can take a different path and produce a different answer on each run, which quietly breaks your habits for testing, debugging, and auditing.”
Each gets the same treatment below: the old habit, why it fails, the replacement and a first-week exercise. The exercises are mine; the replacements are the book’s.
How does testing change?
Testing splits in two along the seam. The old habit is to run the code and assert that the output equals an expected value, then read one green run as proof. Chapter 16 names the assumption underneath: “The practice of running fixed tests over deterministic code assumes the code is the behavior.”
With a model in the path, the behavior is a distribution, and an equality assertion samples it once. A 2026 essay, Stop Asserting Equality, describes the result in two sentences that Chapter 15 quotes: “It passes on Tuesday. It fails on Wednesday because the model reworded one sentence.”
The same failure reaches CI gates. One issue from September 2026 describes a build that failed whenever a model-reported count of critical findings was above zero: “That integer returned 1, then 0, then 3 across three consecutive runs on a single branch whose production code barely changed.”
Chapter 15 gives the diagnosis for teams that delete such tests and call agents untestable: “The conclusion is wrong; the assertion was simply on the wrong side of the seam. Exact-match belongs to the deterministic layer.” The replacement is two kinds of test, shown here in pseudocode with invented names and illustrative numbers.
# Your side of the seam: a scripted model, exact assertions, no network.
test "a write that times out is escalated and never sent twice":
model = scripted([ call(send_invoice, order="A-1042") ])
tools = { send_invoice: times out on every call }
run = run_agent(task, model, tools)
assert run.exit == ESCALATED # exact
assert tools.send_invoice.calls == 1 # exact
# The model's side: the live model, properties, and a pass rate.
test "the refund summary is acceptable", trials = 10:
passes = 0
repeat 10 times:
out = run_agent(task, live_model, sandbox_tools)
if parses(out) and out.customer == "C-77" and checker(out):
passes = passes + 1
assert passes >= 8 # illustrative gate; set yours from measured runs
The first kind is where retries, budgets and stop rules get proven. Chapter 2 words the second kind: “write tests that assert properties of the output—it parses, it names the right customer, it passes the checker”. The file in build an AI agent from scratch uses the same scripted stand-in model and no framework.
Why does a CI job with ten agent tests go red with no bug?
It goes red because every one-shot test is a draw, and a job needs all of its draws to come up green. Take an illustrative case: ten one-shot agent tests in one job, each passing 90% of the time, independently. All ten pass together with probability 0.910 = 34.9%. The job is green about one run in three with no defect anywhere.
The calculator below opens on those numbers. Read its “attempts” as the tests in the job. It prints 34.9% for all ten passing, “> 99.9%” for at least one passing, and a 65-point gap between the two.
With JavaScript on, the pass@k and passk calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the pass@k and passk calculator on its own page to share a result by link.
A pass-rate gate changes the odds without removing them. The gate in the pseudocode, at least 8 passes in 10 trials, is green 93.0% of the time at that 90% rate, so it still fails about one run in fourteen. If the true rate fell to 70%, the same gate would still pass 38.3% of the time. Independence is assumed throughout, and real runs of one task are often correlated.
Those numbers are why sizing matters, and two pages own it: how many eval examples you need and pass@k versus passk. Live trials also cost model calls. One unanswered question from February 2026 put the bind well: “Mocking at the HTTP layer felt brittle. Hardcoded fixtures were unrealistic. Calling live models felt unnecessary just to validate application logic.”
Chapter 16’s accommodation for the cost is “a smoke subset per commit and the full battery nightly and before release”, wired in as a regression gate. Keep the live suite, though. One issue from July 2026 reports two defects, one in output parsing and one in how run status was derived, that passed a scripted end-to-end test: “Both were invisible to the scripted E2E and only surfaced by running a real agent.”
First-week exercise. Find every test that compares model output to a string. Move each to one side of the seam: rewrite it as a property, or script the model and keep the exact assertion.
How does debugging change?
Debugging moves from recreating a failure to recording it. The old habit is to run the failing input again and watch. Chapter 2’s advice is to “expect the failure you saw once to resist reproduction on demand”, because the second run draws a different sample and takes a different path.
The replacement is record and replay, which Chapter 15 (in the full book) says is borrowed from an older trade. In record mode, an interception layer writes down every model call and every tool call as it crosses the seam. That means the full prompt, the sampling parameters, the model’s identity and the returned tokens, then each tool’s arguments and complete result, errors included.
In replay mode your own code runs again, live and debuggable, while every external call is served from the recording. The clock is the input people forget, since a date stamped into a prompt makes tomorrow’s replay a different prompt. The chapter’s summary: “Record once, and the snowflake becomes a specimen: a failure you can step through as many times as the diagnosis takes.”
Your tracing habit transfers directly. A trace of spans already has the right shape, and the post on what an agent span should carry lists the fields, with the ones that must never be recorded.
First-week exercise. Take one failed run and try to replay it from what you stored. List what is missing: the rendered prompt, the parameters, the model’s identity, a tool result, the clock.
How does auditing change?
Auditing moves from reading a status to checking an outcome. The old habit has three forms: a 200 means it worked, the component’s own completion message means it finished, and a quiet dashboard after a deploy means the release is fine. All three trust a signal that the nondeterministic side can produce while being wrong.
Chapter 18 describes the first form, a tool wrapper that catches an exception and returns an empty result, after which the model writes a fluent summary of data it never reached: “Nothing crashed. Every status code on your dashboard reads as success. The failure has been laundered into an answer.” Chapter 3 covers the second in one line: “The model’s “done” is testimony; a green test is evidence.”
Chapter 20 covers the third with a scene it labels a composite of incident reports. A one-line prompt edit merges on a green suite, and support tickets rise two days later with nothing paging. Its verdict is that outputs vary from run to run, “so any single green run proves little”.
The replacements follow from each form. Return failures to the model as structured results that say what failed. Put a check a program can evaluate between the model’s last message and the line that records a run as done. Judge a release by comparing pass rates across versions.
Those three are the first moves in the larger subject of how to make AI agents more reliable, which also takes in budgets, fallbacks and escalation.
First-week exercises. Make one tool wrapper fail on purpose and read what the model receives. Then find the line where a run ends, and see whether the model’s final message can reach it unchecked.
The first-week checklist
Work through these against your own agent, in any order.
- Type a minimal agent yourself, with a scripted model in place of a real one, and break it on purpose once.
- Every test that compares model output to a string is now a property test, or runs against a scripted model.
- At least one live case runs several times in CI and gates on a pass rate that you measured.
- One failed run has been replayed from stored data, and the missing fields are written down.
- One tool wrapper has been made to fail, and the model received a structured error instead of an empty result.
- A check that a program evaluates stands between the model’s final message and the record that says done.
- Every tool that creates, charges, sends or deletes carries an idempotency key and leaves a receipt.
- Every source of text the agent reads that an outsider can write to is on a list.
What is the one new trust boundary?
The new boundary is text: anything the model reads can act as an instruction. Chapter 2 lists it as a property of the engine, separate from nondeterminism: “the model reads instructions and data in one undifferentiated stream, and it cannot reliably tell them apart”. Chapter 17 (in the full book) draws the consequence for an agent with tools: “you turn every scrap of text it reads into a potential set of orders”.
You have defended a boundary like this before. The person who named prompt injection explained in 2024 that he chose the name because “it was analogous to SQL injection, where untrusted user input is concatenated with trusted SQL code”. Parameterized queries closed that hole with structure, and Chapter 17 marks where the analogy ends: “The difference is that SQL injection had a complete fix waiting to be adopted.”
No equivalent exists for a prompt. So the defense moves from the channel to the consequences: checks that run outside the model, least privilege on every credential, and removing one leg of the lethal trifecta. The posts on the lethal trifecta and how to prevent prompt injection in AI agents own that argument, and the lethal trifecta audit runs it against your tool list.
First-week exercise. List every source of text your agent reads that someone outside the team can write to: inbound email, fetched pages, ticket bodies, file contents, another system’s tool results.
Which post should you read for the ticket you were handed?
Read the post that matches the ticket; a second general introduction to AI agents for backend developers will not help. The table is my routing through pages that already exist, and each owns its topic in more depth than this map does.
One ticket, walked through both tables
Take “the agent sent a customer the same email twice”, an invented ticket. In the mapping table, the on-call filter keeps the idempotency row, whose book sentence is the diagnosis: a checkpoint records position and cannot record whether the outside world did the work. The fix in that row is a key built from the logical action, plus a receipt checked before any retry.
The reading path agrees. Its third row sends you to the invoice-email example, which is this crash, and then to the resume table for what a restarted process may repeat. Neither table points at the model.
What do you still need to learn about the model?
You need a working picture of the engine, and Chapter 2 (free to read) is that picture in one chapter with no math. Skipping model training is reasonable. A comment from July 2024 made the case that “deep knowledge of how to train models isn’t actually that useful when working with generative AI”. Treating the model as an opaque box that needs no understanding is the opposite mistake.
Six facts from that chapter change design decisions a backend engineer makes:
- It predicts the next token. The model completes text, so a missing fact gets filled with a plausible string.
- Its limits are counted in tokens. What it can see, what a call costs and how long it takes share one unit.
- It samples. Output is drawn from a distribution, the source of all three broken habits.
- It keeps nothing between calls. The history you send is all the model knows.
- Programs read its output. Structured output and tool calls let ordinary code branch on a response.
- It fails in characteristic ways. Chapter 2 lists six, from hallucination to instructions and data sharing one stream.
Small per-step error rates multiply along a loop, and why agent errors compound does that arithmetic.
Where does this map stop being useful?
The map stops being useful once you start designing, and it has four limits I can name.
It is an orientation. Each row hides a post’s worth of detail, and the linked page is where the decision gets made.
The count of three is the book’s. Other writers divide the changes differently, and Chapter 2 has its own list of three habits, which are the replacements. Prompt injection is a fourth change by any count.
No share is measured. The book says of the harness that “building it is most of the work” and attaches no figure. I found no study of how much of an agent is ordinary engineering.
The evidence is a harvest. The forum posts and issues quoted here are examples of what people ask. They were collected by hand and support no claim about how common each problem is.
The sort to do this week
Open the agent ticket and mark each task with one of three labels: build it the way you always would, same pattern with one adjustment, or old habit fails here. A guide to AI agents for backend developers earns its keep if you apply the first label without apology and stop at the third. The book’s thesis is the test for all of them: “an agent is only as trustworthy as the signal you can use to verify it”.
Chapter 1, What Is an Agent? holds the sentence about the three habits and is free to read, as is the engine chapter. The loop and its four parts are in Chapter 3, The Agent Loop (in the full book), and retries, state and recovery are in Chapter 18, Reliability, State, and the Harness (in the full book). The free guide to agent fundamentals collects the related posts and tools, and you can see the formats.
Questions readers ask
- Do I need a machine-learning background to build AI agents?
- No for the engineering: the work is orchestration, state, retries, tool calls and I/O, which a backend engineer already does. You do need a working picture of how the model behaves: that it samples, that it keeps nothing between calls, and how it fails. Chapter 2 of the book covers that in one chapter with no math, and it is free to read.
- Is an AI agent just a distributed system?
- For the code you wrote, yes: delivery, retries, duplicate side effects, state, recovery, tracing and capacity all follow the patterns you know. The model's output is a sample from a distribution, and that changes how you test, debug and audit. Text the model reads is also a trust boundary that classical systems closed with structure and agents have not.
- How do you unit test an AI agent?
- Split the tests along one line. Tools, parsers and loop control are ordinary code: test them with exact assertions against a scripted model that returns a fixed sequence of responses. Live model output gets property assertions (it parses, it names the right customer, it passes a checker) and a pass rate over several runs. A sampled live suite is still needed, because a scripted model only emits what the test tells it to.
- Why are my agent tests flaky in CI?
- A one-shot test of a component that is right nine times in ten is red one run in ten by construction. Put ten such independent tests in one job and, in this post's illustrative arithmetic, the job is green 34.9% of the time with no bug. Move exact assertions to the code you wrote, and gate live cases on a pass rate.
- Can I build an AI agent without a framework?
- Yes. The book counts four parts (a model client, tools, a message history and a loop) and describes what a framework adds as plumbing: persistence, retries, tracing, memory management and orchestration. Its verdict is that a framework can make an agent more dependable and cannot make it smarter, so write the small version first and adopt when you can name the need.
Sources
- Saurav Bhattacharya (2026). Stop Asserting Equality: How to Test Agents When Every Run Is Different
- Horace He, in collaboration with others at Thinking Machines (2025). Defeating Nondeterminism in LLM Inference (one model, one prompt, September 2025)
- Tian Pan (2026). Async Agent Workflows: Designing for Long-Running Tasks
- Tian Pan (2026). Releasing AI Features Without Breaking Production: Shadow Mode, Canary Deployments, and A/B Testing for LLMs
- Veera Ravindra Divi (2026). Your Multi-Agent System Is a Distributed System. Treat It Like One
- Brandur Leach (Stripe) (2017). Designing robust and predictable APIs with idempotency (one example of a payment API's engineering blog)
- Simon Willison (2024). Prompt injection and jailbreaking are not the same thing
- Dex Horthy (HumanLayer). 12-Factor Agents (README, read October 2026)
- k3ntaki (Hacker News) (2025). Ask HN post on why agent tooling is not written in backend languages (December 2025)
- rafterydj (Hacker News) (2026). Hacker News comment calling agentic systems a distributed systems problem (April 2026)
- simonw (Hacker News) (2024). Hacker News comment on what transfers to work with generative models (July 2024)
- akarshc (Hacker News) (2026). Ask HN post on a CI pipeline that called live models in integration tests (February 2026; no replies)
- laurigates (ForumViriumHelsinki/.github issue tracker) (2026). Issue #123: a CI gate on a model-reported count of critical findings (September 2026)
- aguil (aguil/agents issue tracker) (2026). Issue #79: two defects a scripted end-to-end test did not show (July 2026)