Durable execution for AI agents means recording the result of every model call and tool call in an append-only journal, then recovering from a crash by re-running the agent’s code against that journal. Finished steps return their recorded results. The one step in flight when the process died runs again, so that step must be safe to repeat.
Most explanations blur that last sentence. A 2026 sandbox study ran 25,930 agent episodes (Li, 2026, version 1 abstract; a single-author preprint). When a request was still in flight or had been delivered twice, frontier models duplicated the write in 56% and 74% of episodes. A journal on its own does nothing about that window.
After this post you can read your own agent loop as a journal, name the one step a crash can repeat, and decide in writing between a durable runtime, a checkpoint row, or neither. The model is the book’s. The trace, the tables and the checklist are mine, and I say so where each appears.
What is durable execution for AI agents?
Durable execution is a way of running agent code so that a crash costs one step instead of the whole run. The book’s glossary, which is free to read, defines it this way: “Writing agent code as if failure did not exist and letting a runtime make the fiction true: the result of every model call and tool call is recorded in an append-only journal, and on recovery the code re-executes from the top, with journaled steps returning their recorded values instead of running again.”
Chapter 18 introduces it as the second of two checkpoint styles. The first is the snapshot: serialize the working state every so often and reload it on restart. The second is the journal, where the chapter’s instruction is to “record the result of every non-deterministic step (each model call, each tool call) in an append-only journal, and on recovery re-execute the agent’s code from the top, except that any step whose result is already journaled returns the recorded value instead of running again.”
Five terms carry the model, and I use them as the book does:
- Step. One unit whose result gets recorded: a model call or a tool call.
- Journal. The append-only record of one run’s step results, in storage no process owns.
- Replay. Re-running the code from the top, with journaled steps answering from the record.
- Determinism. The same journal always leads the code through the same steps.
- The gap. The interval between a side effect and the record of it.
The book compresses the idea into one line: “The code replays; the world does not.”
The figure is from Chapter 20 and shows what both styles buy: state that lives outside the process.
What does replay do, line by line?
Replay runs every line of your code again and lets only the journaled steps answer from the record. The trace below is this post’s own artifact. The book states the rule and does not walk it. The run, the order number and the amount are illustrative.
The loop
The agent is a refund assistant with two tools: one reads an order, one sends a refund. The loop is pseudocode.
run(task, run_id):
1 attempt = random_id() # outside a step: the planted mistake
2 messages = [instructions, task]
3 n = 0
4 loop:
5 n = n + 1
6 reply = step(n, "model", do: ask_model(messages))
7 messages.append(reply)
8 if reply.is_final: return reply.text
9 n = n + 1
10 result = step(n, "tool", do: run_tool(reply.call, key: attempt + ":" + n))
11 messages.append(result)
step is the only piece a runtime supplies. If a result is recorded under the step number, it returns that result and never evaluates the do: body. Otherwise it records that the step started, evaluates the body, records the result and returns it.
The journal at the crash
The first process got through three steps and died inside the fourth. The payment provider had accepted the refund, and a deploy killed the worker before the result reached the journal.
entry step kind status recorded
1 1 model result call lookup_order(order 8812)
2 2 tool result order 8812: paid 40.00, no refund yet
3 3 model result call send_refund(order 8812, 40.00)
4 4 tool started (no result)
Entry 4 says the step began. Nothing in the journal says whether the money moved.
What does each line do on replay?
A new process starts the same run with the same task and run id, and every line executes again. Only the calls to step behave differently.
| Pass | Line | On replay | Why |
|---|---|---|---|
| start | 1 | Executes for real | Outside a step, so it draws a new random value |
| start | 2–3 | Execute for real | Plain code; same result if the instructions text is unchanged |
| 1 | 6 | Returns the recorded value | Step 1 has a result. The model is not asked |
| 1 | 7–8 | Execute for real | The recorded reply is appended; it is a tool call |
| 1 | 10 | Returns the recorded value | Step 2 has a result. The lookup tool is not called |
| 2 | 6 | Returns the recorded value | Step 3 has a result. The model is not asked |
| 2 | 10 | Executes for real | Step 4 has no result. The refund request goes out again |
| 3 | 6 | Executes for real | Step 5 was never journaled: the first live model call |
Two model calls and one tool call were answered from the journal, so the recovery bought no inference for them. By pass 2, line 10, the messages list in memory matches what the dead process held, although nobody saved a transcript. The code rebuilt it from the journal.
What does the random id on line 1 break?
Line 1 breaks the protection on step 4, and it shows no symptom before that point. The first process drew one random value and built the refund’s key from it. The replaying process draws a different value, because nothing recorded the first one.
Steps 1 to 3 still answer from the journal, since step finds them by number. Then line 10 sends the refund under a key the provider has never seen, the provider treats it as a new request, and the customer is refunded 40.00 twice.
What a runtime does at that moment varies, and here I am reasoning from the model. A runtime that checks each replayed request against its record can stop the run: one runtime’s documentation says a request that does not match its recorded event “returns a non-deterministic error.” The check that page describes covers the step’s kind and name, which a changed key would pass. A runtime that matches by position alone lets the new key through as well.
Either of two repairs satisfies the journal. Build the key from inputs to the run, such as run_id + ":" + n, which replay reproduces exactly. Or draw the random value inside a step, so it is recorded once and returned on every replay. A clock read between steps fails the same way, since a deadline check can branch differently on replay.
To meet the unprotected failure first, Run the loop lets you step a small support agent by hand, duplicate side effect included.
How can replay be deterministic when the model is not?
Replay is deterministic because the model’s answer is a recorded step result: the runtime returns it from the journal and the model is not asked again. Determinism is required only of the code between steps. A reader put the confusion precisely in a launch thread in January 2025: “What is the determinism constraint? I noticed it mentioned several times in blog posts, but one of the use-cases mentioned here is for use with LLMs, which produce non-deterministic outputs.”
The answer is in the trace: line 6 never reached the model in passes 1 and 2. Chapter 20 states what that buys once runs are numerous: “each step’s result is persisted once and retrieved on replay, so a resumed run reads its finished steps from the journal instead of re-running them, and since the expensive steps are model calls, recovery stops re-buying inference.”
One consequence is my reasoning from the model. A model call that was in flight at the crash has no recorded result, so it is asked again and may answer differently. Nothing after it was recorded, so the journal stays consistent, but the recovered run may take a different path.
What if the prompt or the model changed between crash and replay?
The journaled answers come back unchanged, and every live step after them runs under the new prompt and the new model. Chapter 18 states the risk and the remedy in one sentence: “If the prompt text, the model, or a tool’s schema changed between crash and recovery, the replay can diverge from the original run; if you depend on replay, record the versions each step depended on.”
Line 2 executes for real, so instructions holds whatever text the new build carries. Replies 1 and 3 still come from the journal, written by the old model under the old instructions. Step 5 then asks the new model to continue a history it did not produce.
Nothing raises an error in that sequence, which is my reading of the trace and goes one step past the book’s sentence. So store the prompt version, the model identifier and the tool schema version with each journal entry, or pin all three for the life of the run. Then decide in advance what a mismatch does: finish on the old versions, or stop with an error that names it.
What does replay demand of your code?
Replay demands that the code between steps be deterministic given the journal, and that everything else happen inside a recorded step. The book’s wording is “Replay has a price of admission: the replayed code must be deterministic enough that the same journal always reproduces the same sequence of steps, and the world must hold still.”
The book stops at “deterministic enough”. The list and the table below are this post’s own.
Where does non-determinism come from?
Non-determinism enters through anything the code reads that the journal did not record. In the code between steps, I look for seven sources:
- Clock reads, including deadline checks and elapsed-time branches.
- Random values and generated ids, including ids used as keys.
- Iteration order over a set or a map whose order the language does not fix.
- The environment: configuration, feature flags and files that can change.
- I/O outside a step: a network or database call made directly from the loop.
- Concurrency: branches whose completion order decides what happens next.
- The code and the prompt, which change at deploy time under live runs.
What does each requirement cost you if you break it?
Each requirement has a specific failure and a check you can run.
| Requirement | What breaks if you violate it | How to check |
|---|---|---|
| Code between steps is deterministic given the journal | Replay takes a different branch or sends a different request | Search the loop for the seven sources; replay a recorded journal in CI |
| Every side effect and non-deterministic read sits inside a step | The effect repeats on every replay | List each external call and confirm a step wraps it |
| Step inputs and results are serializable and small | A result cannot be recorded, or the journal hits its limit | Round-trip each result type through storage; measure the largest |
| The in-flight step is safe to repeat | A second charge, email or order | Kill the process between the effect and the record; count effects |
| Concurrent steps start in a fixed order and join at a fixed point | Replay resumes branches in a different order | Replay a run with concurrent branches many times |
| Prompt, model, tool schema and code versions are pinned or recorded | Mixed versions in one run, or errors after a deploy | Deploy to staging while a run is mid-flight |
| The journal is bounded | The runtime ends the run at its history limit | Count entries and bytes for your longest run |
What do public bug reports show about these rules?
Two reports on public issue trackers show the rules biting people who were following them. Each is one user’s report, described as the thread stood when I read it to its end on 2026-10-07.
Call order is part of determinism. In a report from September 2025, a user added one call to the runtime’s replay-safe random generator ahead of an existing call and redeployed. Running workflows then failed with a non-determinism error “after they woke up from the sleep”. A contributor replied: “This is expected.” The generator is seeded, so an inserted call shifts every later value, and the issue was closed as not planned.
Concurrent completion order. In a report from June 2026, a user describes replay failing for concurrent branches. On their reading, which they say “might be wrong”, results “arrive in completion order (wall-clock timing)” at first and “in sequence number order” on replay. A contributor answered that a change meant to address it “is in progress”. The issue was open, with no fix confirmed in the thread.
Chapter 18 gives the general warning: “parallel branches finish in nondeterministic order, which naive deterministic replay does not forgive.”
Which step can run twice?
In a one-branch loop, exactly one step can run twice after a crash: the step that was in flight when the process died, in the case where its side effect landed and its result was never journaled. Steps with a recorded result are never run again. With concurrent branches, it is one step per branch.
The book reaches this window from the checkpoint side. Chapter 18 says, “A checkpoint records position. It does not, and cannot, record whether the outside world did the work.” It then names the interval: “The dangerous window is the gap between performing a side effect and recording that you performed it”.
Runtimes say the same in their own documentation. One states: “If a failure occurs after the activity completes but before the result is recorded, the runtime might rerun it.” (Microsoft Learn, page dated 2026-04-22). Another: “Steps are tried at least once but are never re-executed after they complete.” (DBOS documentation).
Why can’t the runtime close the window for you?
The runtime cannot close it because the journal and the outside service are two separate systems, and no single write lands in both. A reader made the argument in a thread from October 2025: “you can crash in-between completing and recording and when you restore you run the step twice.”
So durable execution for AI agents still requires idempotency under any runtime. The re-executed step must carry the same key as the first attempt, and the service must answer a repeated key with the stored result. In the trace that means a key of run_id + ":" + n, which gives step 4 the same key in both processes.
The sandbox study from the opening suggests how much the key matters: offering an idempotency key on every write “lowers the duplicate rate from 28% to 4%” (Li, 2026, version 1 abstract). That is one author’s preprint and a sandbox of six services, so read the figures as direction.
How to build that key, where the receipt lives and which retry policy suits which tool are covered in idempotent tools and safe retries.
Keys stop a step from repeating itself and say nothing about a sequence that must be undone, which is the saga pattern’s job in Chapter 18. And whether a failed step should be retried by the runtime or handed back to the model depends on which of the AI agent failure modes produced the failure.
What happens to runs in flight when you deploy?
A deploy puts new code under every live run, and replay then executes that new code against a journal the old code wrote. Two remedies recur in the runtime documentation I read, under names that are mine.
Pin and drain. Old runs finish on the old version while new runs start on the new one. One runtime’s documentation says requests “start and end on the same version” (Restate documentation, read 2026-10-07).
One practitioner running travel bookings wrote in June 2024 that they keep the old version until its runs complete: “Usually this can be done within a few minutes but sometimes we need to wait days.” A fix reaches those runs only as a hotfix to the old version.
Branch on a recorded version. The new code keeps the old path and chooses between paths by a version marker in the journal. One runtime’s documentation warns that a definition “can change in very limited ways once there is a Workflow Execution depending on it” and recommends its versioning feature (Temporal documentation). Old branches then stay in the code until the last old run ends.
With durable execution for AI agents, the code that changes includes the prompt, the model identifier and the tool schemas. As far as replay is concerned, a prompt edit is a deploy. The test I would add before either remedy is a replay of recorded journals against the new build in CI.
How big can a journal get on a long run?
A journal grows by at least one entry per step, and runtimes cap it. The book does not treat journal growth, so this section rests on vendor documentation and my own reasoning, each marked.
Two documented limits serve as dated examples. One workflow engine says a run’s history is “limited to 51,200 Events or 50 MB and will warn you after 10,240 Events or 10 MB” (Temporal documentation), and its documented remedy is to close the run and continue as a new one. An edge workflow service caps a step’s result at 1 MiB on both its plans, and a workflow’s steps at 1,024 on the free plan and 10,000 by default on the paid plan, configurable up to 25,000 (Cloudflare documentation).
For agents, bytes can grow faster than entries, and this part is my inference with no source behind it. Each model call takes the lengthening transcript as its input, so a runtime that records step inputs stores many overlapping copies of one history.
Two general habits keep a journal bounded. Keep large tool results in a separate store and journal a reference to them. Give long runs a rollover point, where a successor run starts from a compact summary of state.
Journaling also costs time on every step. Chapter 20 puts it this way: “The engines’ price is latency, and it is a fair price honestly disclosed: persisting every step adds a round trip to storage”. I have no sourced figure for that cost, so I print none.
What do durable runtimes say about themselves?
Durable runtimes form a category with several independent members, and their documentation agrees on the points that matter here. The table lists five as labeled examples, each described in its documentation’s own words, read on 2026-10-07. None is a recommendation, and provider behavior is a dated snapshot.
| Example and kind | What it says is recorded | Its determinism rule | Side-effecting steps | Code changed under a live run | Documented size limit |
|---|---|---|---|---|---|
| Temporal, a workflow engine | An event history for each execution | “Workflow code must be deterministic to support replay.” | “Temporal recommends that Activities be idempotent.” | Definitions change “in very limited ways”; a versioning feature | 51,200 events or 50 MB |
| Azure Durable Functions and Durable Task, a cloud service | Execution history, appended “much like an append-only log” | “orchestrator functions must be deterministic” | Each runs “at least once” | Versioning, side-by-side deployment, or stopping in-flight runs | Not read |
| Restate, a journal-based runtime | “Restate tracks every step of your code execution in a journal.” | Not read | A recorded step “won’t repeat on retry”; nothing found on the in-flight step | Requests “start and end on the same version” | Not read |
| DBOS, a library on a relational database | Completed steps | “a workflow function must be deterministic” | “Steps are tried at least once” | Not read in the documentation | Not read |
| Cloudflare Workflows, an edge workflow service | Not read | Steps “should be named deterministically” | “a step might be retried multiple times” | Not read | 1 MiB per step result; a step count cap |
“Not read” means I did not open a page covering the point.
The agreement across the rows is my synthesis: finished steps answer from the record, the code between steps must be deterministic, and a side-effecting step can run more than once. Three of the pages use the words “exactly once” for something narrower: an invocation run to completion (Restate), an activity “observed as completed” that “may be executed multiple times” (Temporal), a database transaction (DBOS). None promises it for an external side effect.
When is a checkpoint enough, and when do you need neither?
A checkpoint row is enough when a run outlives its process, never parks on a person, writes from one branch only and keeps its irreversible effects few and keyed. You need neither when the run finishes in one sitting and a restart from zero is cheap.
The book’s guidance is prose, and it starts from the small end. Chapter 20: “A tasks table and a worker that checkpoints to it is honest engineering, and for runs measured in minutes it is usually all the durability you need.” The same paragraph names the signals for buying more: “The engine earns its keep when runs span hours or days, when human approvals or external events park them mid-flight”, and when a multi-agent fan-out gives one task many concurrent branches to keep consistent.
A softer signal follows: a queue that has grown a scheduler, then retry bookkeeping, then a resume path. “The moment you catch yourself building the machine is the right moment to buy it.” Chapter 18 reduces the judgment to two questions: “what here is irreversible, and how long does the run live?”
The decision table
The table is mine. It turns that prose into rows, adds the rows where the answer is nothing, and splits hours-long runs by where their irreversible effects sit. That split is my refinement, since the book names run length alone as its first signal.
Take the first row that matches, reading from the top. “One sitting” means starting again from zero costs less than building a resume path.
| Row | Your run | What you take on | Verdict |
|---|---|---|---|
| 1 | Parks mid-run on a person or an external event, for hours or days | Determinism rules, version discipline, journal limits, a storage round trip per step, one more system to operate | durable runtime |
| 2 | Fans out into concurrent branches that each change the outside world | The same, plus a fixed start order and a fixed join point | durable runtime |
| 3 | Lives hours or days with irreversible effects spread through the middle | The same as row 1 | durable runtime |
| 4 | Already sits on a scheduler, retry bookkeeping and a resume path you wrote | A migration, and the rules in row 1 | durable runtime |
| 5 | Too long to restart from zero cheaply; no parked waits; one branch, or branches that only read; irreversible effects none or few, at known points | A tasks row, your own dead-run detection and resume logic, crash drills | checkpoint row |
| 6 | Finishes in one sitting with one or two writes | An idempotency key on each write, then restart from zero | neither |
| 7 | Finishes in one sitting, read-only | Restart from zero and an honest error to the user | neither |
In every row, a write that can be in flight at a crash still needs its key. The verdict changes who records position and leaves that requirement alone.
Is a framework checkpointer the same as a durable runtime?
A framework checkpointer is the snapshot style, and what it leaves to you is the argument for a runtime. A vendor’s CTO wrote in February 2026 that such checkpointers give you “a snapshot of state that you, the developer, are responsible for detecting the need to use, manually triggering, and coordinating at scale to avoid duplicate work” (Schneider, 2026, listed in the sources).
That is a fair description of row 5, from someone who sells the alternative.
Recovery is one late step in how to make AI agents more reliable: error surfacing, bounded retries and budgets come earlier and cost less.
Worked example: three agents through the table
Three illustrative agents go through the decision table and then the checklist in the next section, and each ends with one verdict.
A 40-second support agent with read-only tools. Rows 1 to 5 do not match: nothing parks, nothing fans out, and the run lasts seconds. Row 6 needs a write, and there is none. Row 7 matches, so the verdict is neither. No checklist item applies; the one test is that a killed run gives the user an error or a clean retry.
A 3-hour coding agent that opens one pull request at the end. Rows 1 to 4 do not match, provided it runs as one branch and edits only a disposable workspace. Row 5 matches, so the verdict is a checkpoint row: save the transcript and plan every few units of work.
Its in-flight risk is the pull request, so key that step by the branch name. Checklist items 1 to 4 apply.
A multi-day procurement agent that waits on human approvals and places orders. Row 1 matches at once. The verdict is a durable runtime, every checklist item applies, and each order needs a key that survives a wait of days.
Chapter 18 adds one requirement for this agent. The approval itself must be durable state, “recorded with who granted it, when, and a fingerprint of the exact artifact they saw”. A replayed run that reaches a live step should also pass the same AI agent guardrails as a fresh run, including the approval gate in front of each order.
How do you prove recovery works before you depend on it?
You prove it by killing the process at the worst moments and watching what the resumed run does. Chapter 18 calls the habit crashing on purpose and gives the reason in one clause: “a checkpoint you have never restored from is a guess.”
The checklist is this post’s own. Items 1 to 4 apply to any run that keeps state. Items 5 to 9 apply only when you depend on replay.
- 1. Kill the process right after a model call returns. Under a journal the resumed run does not ask that call again; under a snapshot it re-asks only calls made after the last checkpoint.
- 2. If the run has a write tool, kill the process after the tool succeeded and before its result was recorded. The outside system shows one effect.
- 3. Resume a run on a different machine from the one that started it. It continues from storage alone.
- 4. Deploy a new build while a run is mid-flight. The run finishes correctly, or stops with an error that names the version mismatch.
- 5. Search the code between steps for clock reads, random values, environment reads and direct I/O. Each one is inside a step or removed.
- 6. Replay recorded journals against the new build in CI. The sequence of steps matches the recording.
- 7. Confirm each journal entry carries the prompt version, model identifier and tool schema version it depended on.
- 8. If the run has concurrent branches, replay it many times. Every replay resumes the branches in the same order.
- 9. Measure entries and bytes for the longest run against the documented limit. A rollover point is defined and tested.
Item 2 is the only drill that exercises the gap.
What does a replay audit of your own loop look like?
A replay audit is one page that names your steps, the code between them, the write that can repeat, and the verdict.
REPLAY AUDIT run type: ______________ date: __________
1. STEPS
One step is: ____________ Results are recorded in: ________
Recorded before the next step begins? yes / no
2. CODE BETWEEN STEPS (one row per line that is not a step)
line | what it reads | same value on replay? | fix
____ | _____________ | yes / no | ____________
Look for: clock, random ids, iteration order, environment, I/O
3. THE STEP THAT CAN RUN TWICE, ONE PER BRANCH (row per write tool)
tool | key built from | same key when re-executed? |
| what the service does with a repeated key
____ | ______________ | yes / no | ____________________
4. VERSIONS recorded per step: prompt [ ] model [ ] tool schema [ ]
On a mismatch the run: finishes on old / stops with error
5. RUNS IN FLIGHT AT DEPLOY
Longest run: ________ Time between deploys: ________
Remedy: pin and drain / branch on recorded version / none
6. JOURNAL SIZE
Steps in the longest run: ______ Documented limit: ______
7. VERDICT: neither / checkpoint row / durable runtime
Because (two sentences): __________________________________
What does this post not cover, and where is its evidence thin?
This account of durable execution for AI agents rests on documentation, two user reports and one sandbox preprint, and it has four limits worth stating.
No first-hand production account. I found no primary write-up of a production agent repeating a side effect after a crash, and I claim no rate for real systems.
A dated snapshot of five products. Every documentation quote was read on 2026-10-07, and limits, guarantees and versioning features change.
No latency or cost numbers. The book names latency as the price, and I found no figure with a source, a date and stated conditions. Measure the per-step overhead on your own storage.
The decision table is a judgment. Its split of hours-long runs is mine, and a reasonable engineer could put the 3-hour coding agent on a runtime for deploy survival alone.
The line to find this week
Open your agent loop and find the line that sits between a side effect and the record of it. Then find every line between steps that reads a clock, draws a random value or loads a prompt. Those two searches show what durable execution for AI agents would do to your own code on replay.
The model in this post comes from Chapter 18, Reliability, State, and the Harness and Chapter 20, Deploying and Scaling, both in the full book. The glossary entries are free to read. The security, reliability and cost guide collects the neighboring posts and tools, or you can see the formats.
Questions readers ask
- Does durable execution give exactly-once tool calls?
- No. A step whose result is in the journal is never run again, but a step that was in flight when the process died, one per concurrent branch, can run a second time. Four of the five runtimes read for this post say in their own documentation that such a step runs at least once, may be retried, or should be idempotent. An idempotency key on the write closes the gap.
- How can replay be deterministic if the model is not?
- The model's answer is recorded as a step result. On replay the runtime returns that recorded answer and does not ask the model again. Determinism is required only of the code between steps: given the same journal, it must request the same steps in the same order.
- What happens to running agents when I deploy new code?
- Replay runs the new code against a journal the old code wrote. Either the old runs finish on the old version while new runs start on the new one, or the new code keeps the old path and chooses by a version recorded in the journal. For agents, a prompt edit, a model change and a tool schema change all count as a deploy.
- Is a framework checkpointer the same as durable execution?
- It is the other of the two checkpoint styles in Chapter 18: a periodic snapshot of working state that a fresh process reloads. It saves position. Detecting the dead run, resuming it and protecting the write that was in flight remain your code unless the framework says otherwise.
- Do I need durable execution for a short-running agent?
- Usually not. Chapter 18 says the thirty-second read-only summary needs none of the state machinery, and Chapter 20 says a tasks table with a checkpointing worker is usually all the durability a run of minutes needs. Any write still needs its idempotency key.
Sources
- Jiapeng Li (2026). Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents (preprint, version 1, abstract)
- Temporal documentation (2026). Workflow Definition (documentation, read 2026-10-07)
- Temporal documentation (2026). Events and Event History (documentation, read 2026-10-07)
- Temporal documentation (2026). Workflow Execution limits (documentation, read 2026-10-07)
- Temporal documentation (2026). Activity Definition (documentation, read 2026-10-07)
- Microsoft Learn (2026). Durable orchestrations (documentation, read 2026-10-07)
- Microsoft Learn (2026). Orchestrator function code constraints (documentation, read 2026-10-07)
- Microsoft Learn (2026). Versioning in Durable Functions (documentation, read 2026-10-07)
- Microsoft Learn (2026). Durable Task programming model overview (documentation, page dated 2026-04-22)
- Restate documentation (2026). Durable Execution (documentation, read 2026-10-07)
- Restate documentation (2026). Versioning (documentation, read 2026-10-07)
- DBOS documentation (2026). Workflows tutorial (documentation, read 2026-10-07)
- Cloudflare documentation (2026). Rules of Workflows (documentation, read 2026-10-07)
- Cloudflare documentation (2026). Workflows limits (documentation, read 2026-10-07)
- Yaron Schneider (Diagrid) (2026). Why Checkpoints Aren't Durable Execution (vendor blog)
- dpr-synth (GitHub) (2025). Issue 1109: a call to the seeded random generator added before an existing one breaks replay (user report, closed)
- paulcacheux (GitHub) (2026). Issue 1578: non-determinism with concurrent local steps (user report, open on 2026-10-07)
- darkteflon (Hacker News) (2025). Hacker News comment asking what the determinism constraint is
- chc4 (Hacker News) (2025). Hacker News comment on exactly-once claims
- rockostrich (Hacker News) (2024). Hacker News comment on keeping old workflow versions until their runs finish