Home / Blog / Agent fundamentals / What Is the ReAct Agent Pattern, and When Does …

Agent fundamentals

What Is the ReAct Agent Pattern, and When Does It Help?

What is the ReAct agent pattern? See one task run three ways, where the original paper won and lost, and what a written thought costs. Pick your shape.

By Enrique Gutiérrez · Published · 23 min read

The ReAct agent pattern is a schedule for a model with tools: write a short thought, take one action, read the observation, then think again with that result in view. So what is the ReAct agent pattern in practice today? It is the rhythm inside most tool-calling loops, with the original text format largely gone.

By the end of this post you should be able to do three things. You can say what ReAct adds to a plain tool loop. You can predict from a few task properties whether interleaved written reasoning will earn its tokens. And you can choose between chain of thought, plan-then-execute and ReAct with the arithmetic on the page.

The material draws on two chapters of AI Agents, Engineered: Chapter 3, “The Agent Loop” and Chapter 4, “Planning, Reasoning, and Self-Correction”, both in the full book. The glossary, which is free to read online, carries the short definitions.

What is the ReAct agent pattern, in the book’s words?

ReAct, short for Reason + Act, is an interleaved trajectory. Chapter 3 describes it as “a short written thought, one action, one observation, then the next thought, conditioned on whatever just came back”. Shunyu Yao and colleagues named it in a paper first posted in 2022 (Yao et al., 2022).

The chapter arrives at the pattern through two older lines of work, each with a typical failure. One taught models to reason by writing intermediate steps, a technique called chain of thought. The other taught models to emit actions in an environment. On the first the book is blunt: “Reasoning alone is a sealed room.”

A model that only reasons can rearrange what it already holds, so a missing fact gets a plausible substitute. A model that only acts keeps no account of where it is heading. Interleaving gives each a turn, and the book’s name for the benefit is grounding: reasoning kept tied, step by step, to facts checked against the world.

Three schedules for a model with tools.
Figure 3.3 Three schedules for a model with tools. Reasoning alone (top) drifts ever farther from the world and ends by elaborating a fabrication. Acting alone (middle) leaves no account of where it is headed, so the first surprise leaves it repeating a gesture that no longer fits. Interleaving (bottom) alternates a thought, an action, and an observation, so every step is re-anchored to the world before the next—the schedule that reaches the goal. Reuse this diagram

One distinction keeps the two halves straight. In the chapter’s words, “a thought changes nothing outside the model”, while an action touches the world or fetches a piece of it. A thought is one more entry in the history, and no tool ever executes it.

What does ReAct add to a plain tool loop?

ReAct adds one thing to a plain tool loop: a written reasoning step between each observation and the next action, kept in the history. The loop already acts and observes. The post on what an agent loop is and how its four beats work covers that machinery.

The written step does two jobs. It steers the next action, because text produced before a decision can shape the decision, and it leaves a log that a person can read later. Both jobs cost tokens, so a useful answer to what is the ReAct agent pattern has to say when that purchase pays.

Is ReAct the same as native tool calling in a loop?

ReAct and native tool calling answer different questions, and the name “ReAct” hides two layers. The format is the original recipe: literal Thought:, Action: and Observation: lines, parsed out of raw text. The schedule is the order of events: reason a little, act once, observe, repeat.

A tool call is the interface for a single action. The model emits a structured request, and your code runs it. ReAct is the schedule around that interface. A loop that calls the model again after every tool result is running the ReAct schedule, whether or not anyone typed the word “Thought”.

The book’s glossary entry says what happened to each layer: “The rigid textual format of the original has dissolved into model training; the schedule survives”. Chapter 3 adds that current models are typically trained to reason before acting, “so you will rarely type the word Thought: yourself”.

This matters when a framework offers two constructors, one labeled “ReAct agent” and one labeled “tool-calling agent”. Often the difference is the format layer: who parses the action, and whether a prompt asks for a visible thought. The schedule is the same in both. The guide to AI agent architecture and its fifteen decisions lists this free loop as the default shape.

Is the ReAct pattern obsolete?

The format is close to obsolete, and the schedule is the default behavior of tool-calling models. That two-part answer covers most of the confusion in the forum threads I read for this post.

Here is one dated example of how far the schedule has moved into the model. One provider’s documentation pages, titled “Thinking” and “Extended thinking” and read on 2026-10-07, describe reasoning that can occur between tool calls within a single turn. The pages say these tokens “are billed as output tokens, even when the thinking text isn’t returned to you”. They also say the visible text of such reasoning is a summary.

Product documentation changes, so treat the page as a snapshot. The transferable idea is that thought, action and observation can now alternate with no prompt asking for it, and that the thought may be hidden or summarized.

From there I draw an inference, and I want to label it as mine. The book is silent on this point. When a model already reasons between tool calls, an instruction to write a thought before each call asks for part of that reasoning a second time. The visible thought then earns its tokens mainly as a log you can read and grade.

One task, three reasoning shapes

Answers to what is the ReAct agent pattern tend to stay abstract, so here is one small task run three ways. Each transcript below is my own illustration, with an invented order and invented stock. All three follow the plain-text pseudocode style of the book and no particular framework.

The goal: “Can order A-1042 ship today, and from which of our two EU warehouses?” Two read-only tools exist: get_order(order_id) and get_stock(warehouse, sku). I chose the task because the first observation changes what should happen next.

Chain of thought only

model (one call, no tools):
  The order is most likely a keyboard order, SKU KB-200.
  Rotterdam is the larger site and normally carries KB-200.
  Three units is a small quantity, so stock should cover it.
  answer: Yes. Ship from Rotterdam, which has KB-200 in stock.

Every line reads as careful, and every fact in it is invented. The model had no order record and no stock count, so it supplied plausible ones. One call, zero tool runs, zero facts checked.

Plan-then-execute

model (call 1, the planner):
  plan:
    #1 = get_order("A-1042")
    #2 = get_stock("rotterdam", #1.sku)
    #3 = get_stock("lyon", #1.sku)
    answer from #1, #2, #3

code (no model call between steps):
  #1 -> { sku: "KB-200", qty: 3, substituted_by: "KB-210" }
  #2 -> { warehouse: "rotterdam", sku: "KB-200", on_hand: 0 }
  #3 -> { warehouse: "lyon", sku: "KB-200", on_hand: 0 }

model (call 2, the solver):
  answer: Unknown. The order now ships as KB-210, and no step
          fetched stock for KB-210.

The plan was sensible and it was written before anyone had seen the order. The record carried a substitution the planner could not know about, and the executor filled the placeholders as instructed. A careful solver reports the gap, as above. A careless one reads two zeros and answers “no”.

Add one replanning check after step #1 and the design recovers. A third model call sees the substitution, rewrites steps #2 and #3 for KB-210, and the solver answers correctly.

ReAct

thought: I need the order's items before I can check stock.
action:  get_order("A-1042")
observation: { sku: "KB-200", qty: 3, substituted_by: "KB-210" }

thought: The order now ships as KB-210, so that is the item to check.
action:  get_stock("rotterdam", "KB-210")
observation: { warehouse: "rotterdam", sku: "KB-210", on_hand: 1 }

thought: One unit is short of three. Check Lyon.
action:  get_stock("lyon", "KB-210")
observation: { warehouse: "lyon", sku: "KB-210", on_hand: 5 }

thought: Lyon covers all three units.
answer: Yes. Ship from Lyon, which holds 5 of KB-210.

What changed between the three runs?

The three runs differ in how often the model was made to look before deciding again. The table counts the calls in the transcripts above.

Shape Model calls Tool runs Outcome in this example
Chain of thought only 1 0 Fluent answer built on invented facts
Plan-then-execute, no replan 2 3 Correct lookups for the wrong item
Plan-then-execute, one replan 3 3 Correct
ReAct 4 3 Correct

Interleaving paid the most model calls and needed no special provision for the surprise. The plan paid the fewest calls among the tool-using shapes and needed a replan step to survive. On a task with no surprise, the two-call plan would have been right and cheapest.

What did the original ReAct paper measure?

The original paper measured prompted ReAct against reasoning-only and acting-only prompts on two knowledge benchmarks and two interactive ones, with mixed results. The knowledge tasks were HotpotQA (multi-step question answering) and FEVER (fact verification). The tool was a simple encyclopedia search interface.

The baselines came from the same worked examples with parts removed. “Standard” had neither thoughts nor actions, “CoT” (chain of thought) kept the thoughts only, and “Act” kept the actions only. “CoT-SC” sampled 21 chain-of-thought answers and took the majority.

Prompting method (Yao et al., 2022, Table 1) HotpotQA, exact match FEVER, accuracy
Standard 28.7 57.1
CoT 29.4 56.3
CoT-SC 33.4 60.4
Act 25.7 58.9
ReAct 27.4 60.9
CoT-SC, backing off to ReAct 34.2 64.6
ReAct, backing off to CoT-SC 35.1 62.0

ReAct beat the acting-only prompt on both benchmarks. Against chain of thought it split: ahead on FEVER by 4.6 points, and behind on HotpotQA by 2.0 points. The paper says so itself, writing that ReAct “slightly lags behind CoT on HotpotQA”. My own reading of the table adds that plain ReAct also sat below the Standard prompt there.

On the two interactive benchmarks the comparison was against acting alone. In ALFWorld, a simulated household, the best of six ReAct trials reached 71% success against 45% for the best acting-only trial. In WebShop, a shopping simulation, the success rate was 40.0 against 30.1.

On the knowledge tasks the highest scores were combinations. One started with ReAct and fell back to the chain-of-thought vote when no answer arrived within the step cap. The other started with the vote and fell back to ReAct when fewer than half the samples agreed. The paper reports that these reached “CoT-SC performance with 21 samples using merely 3-5 samples”.

Where did interleaving lose, and why?

Interleaving lost on HotpotQA, and the paper’s failure analysis shows a trade of one failure for another. The authors labeled 50 correct and 50 incorrect HotpotQA runs for each of ReAct and chain of thought by hand.

Failure analysis (Yao et al., 2022, Table 2) ReAct CoT
Correct answers with sound reasoning and facts 94% 86%
Correct answers with hallucinated reasoning or facts 6% 14%
Failures from a reasoning error, including repeated steps 47% 16%
Failures from an empty or unhelpful search result 23% n/a
Failures from hallucination 0% 56%
Failures from label ambiguity 29% 28%

Read the two middle rows together. More than half of the chain-of-thought failures were hallucinations, and none of the ReAct failures were. Meanwhile a far larger share of ReAct’s failures were reasoning errors. The paper names “one frequent error pattern specific to ReAct, in which the model repetitively generates the previous thoughts and actions”.

So the complaint that an interleaved agent circles on the same thought and call is as old as the pattern.

What are the limits of that setup?

The setup was one frozen large model of its time, steered by a handful of worked examples in the prompt. It predates native tool-calling interfaces and built-in reasoning modes. The paper reports no token or cost figures.

Several details narrow the result further. The failure table rests on 50-example samples. The HotpotQA metric is exact match, and label ambiguity accounts for nearly three in ten failures on both sides. Step caps were 7 on HotpotQA and 5 on FEVER.

One more detail corrects a common picture. For the interactive tasks the paper says thoughts “appear sparsely in the most relevant positions of a trajectory”, with the model deciding when to think. The original did not demand one thought per action everywhere.

What did later work find?

Later work complicated the picture in two directions: one study questioned why the prompt helped, and a plan-first lineage measured large savings in tokens and latency. Neither settles the question for current models, and I report each with its own limits.

Does the content of the thought matter?

A 2024 preprint by Verma, Bhambri and Kambhampati argues that it mattered little in one benchmark family (arXiv:2405.13966). The authors varied the worked examples in a ReAct prompt on ALFWorld, a simulated household task. They moved all the thinking to the front, anonymized it, or replaced it with a placebo.

Their abstract states the finding directly:

the performance is minimally influenced by the “interleaving reasoning trace with action execution” or the content of the generated reasoning traces in ReAct, contrary to original claims and common usage.

By their account, similarity between the example tasks and the query drove the results. In their Table 1, three of four models scored higher with the reasoning moved to the front, and one scored lower.

The limits are stated by the authors: “we restricted our discussion to sequential decision making problem of AlfWorld”. The study varies where thoughts sit in few-shot examples. In every variant the agent still acts and observes at each step. It leaves untouched the schedule under native tool calling with no examples at all.

Is planning first cheaper than interleaving?

Planning first was measured as cheaper in two papers, under a specific condition: code executes the plan’s steps with no model call between them. The lineage starts with Plan-and-Solve prompting (Wang et al., 2023), a single-call prompt with no tools. In its 100-example error analysis, missing-step errors fell from 12% to 7% with the fuller plan prompt. Semantic errors stayed at 26% to 27%.

ReWOO (Xu et al., 2023, arXiv:2305.18323) turned the idea into an agent design. A planner writes the whole plan with placeholders for evidence, workers run the tools, and a solver answers. The abstract claims “5× token efficiency and 4% accuracy improvement on HotpotQA”.

HotpotQA, 1,000 questions (Xu et al., 2023, Table 2) Accuracy Tokens per question Steps
Direct prompting, no tools 37.8 55.5 1.00
Chain of thought, no tools 41.6 481.9 1.79
ReAct 40.8 9,795.1 4.97
ReWOO 42.4 1,986.2 4.45

The interleaved agent used 4.9 times the tokens of the plan-first agent in that table. It also used about 20 times the tokens of chain of thought, for 0.8 points less accuracy. Across six benchmarks the paper reports that ReWOO could “reduce token usage by 64% with an absolute accuracy gain of 4.4%”.

The same paper contains a result against both tool designs. Reading its own table, it writes that “Direct Prompting and CoT, where we don’t provide any external tool, outperform both ALM paradigms”. ALM is its abbreviation for a tool-augmented language model, and the remark holds on some of its sets.

On its trivia set, direct prompting scored 80.6 against 59.4 for ReAct and 66.6 for ReWOO.

LLMCompiler (Kim et al., 2023) writes the plan as a dependency graph and runs independent calls in parallel. Its abstract reports “latency speedup of up to 3.7×, cost savings of up to 6.7×” against ReAct. The “up to” figures come from a task built of independent lookups.

On HotpotQA the same paper shows accuracy level at 62.47% against 62.00%, with a 1.80 times speedup. Its token table shows where the saving lives: 2,900 input and 120 output tokens for ReAct, against 1,300 and 80. Input dominates in every row. The authors also write that they “frequently observe looping and early stopping” in the ReAct baseline, and added prompting to reduce it.

Two ways to plan.
Figure 4.2 Two ways to plan. Interleaved (left) is a tight cycle: decide one step, act, observe, and decide again—maximally adaptive, structurally myopic. Plan-then-execute (right) has a planner draft the whole plan and an executor work through it. What keeps the second design honest is the replan path in accent: when the executor meets a surprise, control returns to the planner to revise. A plan with no replan arrow is just a blind executor. Reuse this diagram

Where do the sources disagree?

The sources disagree in three places, and each side holds under its own conditions.

Question One side The other side
Does interleaving beat reasoning alone? Yes on FEVER (Yao et al., 2022) No on HotpotQA in the same paper; no on sets where no-tool prompts beat both tool designs (Xu et al., 2023)
Is the thought itself what helps? Thoughts guide acting; ReAct beat acting-only on all four benchmarks with its main model (Yao et al., 2022) On ALFWorld, moving or replacing thoughts did not reliably lower scores (Verma et al., 2024)
Is plan-first cheaper? 64% fewer tokens over six benchmarks (Xu et al., 2023); up to 6.7 times lower cost (Kim et al., 2023) The book: with a model-driven executor the design can cost more calls than the loop

Chapter 4 says of the planner and executor split that “run naively, the design spends more total model calls than the loop it replaces”. That describes an executor that calls a model at every step, while the measured savings come from executing steps with code. Both statements hold, and the next section prices each variant.

Xu and colleagues also state where their own design stops working. When little is known about the environment in advance, they write, relying fully on reasoning without observations becomes impractical. A plan needs a knowable list of steps.

What does a written thought cost?

A written thought costs its own tokens once as output and then again as input on every later pass, because the loop re-reads the whole history. The book’s version of the bill: “every action now costs a full model call, with the entire desk re-read each time”.

Every number below is arithmetic of my own, computed from illustrative inputs. Nobody measured them on a running system. I assume a fixed prefix of 3,200 tokens (instructions, tool definitions and the task), tool results of 700 tokens, and ten passes.

The formula is the one the book’s Chapter 19 uses and the site’s estimator implements. Input over n passes is F·n + g·n(n−1)/2, where F is the prefix and g is what each pass adds to the history. Output is what the model writes per pass, times n.

Shape, ten tool steps, illustrative inputs Model calls Input tokens Output tokens Total
Chain of thought only, one 400-token reply, no facts fetched 1 3,200 400 3,600
Tool loop, no written thought (100 tokens out per pass) 10 68,000 1,000 69,000
ReAct, a 150-token written thought per pass (250 out) 10 74,750 2,500 77,250
Tool loop, 150 hidden reasoning tokens per pass, never re-read 10 68,000 2,500 70,500
Plan executed by code: one planner call, one solver call 2 13,800 700 14,500
The same plan with one replan after step five 3 21,300 1,100 22,400
Planner plus a model call per step as executor 11 75,200 1,400 76,600

For the plan rows, the planner writes a 400-token plan and the solver a 300-token answer. The solver reads the prefix, the plan and all ten results. The replan call reads the prefix, the plan and five results, then writes a revised plan of 400 tokens, which the solver reads beside the first. The loop rows count one tool result per pass, as the estimator does, and leave out a final answering call; an eleventh pass would add 12,950 tokens to the ReAct row.

With JavaScript on, the Agent cost-per-task estimator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Agent cost-per-task estimator on its own page to share a result by link.

The estimator above opens on the ReAct row: 77,250 tokens per run, 74,750 in and 2,500 out. Set the output per step to 100 for the bare loop. Then set hidden reasoning to 150 for the fourth row. The tool has no setting for the plan rows, which I computed separately with the same formula.

What do the numbers say?

Read together, the rows say that the loop is the large term and the thought is a surcharge on it. Four readings follow, all of them tied to these inputs.

  • The written thought adds 8,250 tokens, or 12%, to the bare loop. Of those, 6,750 are input: each thought is re-read by every later pass.
  • In the ReAct row, 57% of the input is history being re-read. Tool results account for most of that history.
  • The code-executed plan uses 14,500 tokens against 77,250, a ratio of 5.3. The measured HotpotQA ratio in Xu’s table was 4.9. The two point in the same direction, and a toy sum confirms nothing.
  • The last row is the book’s “run naively” case: eleven calls against ten, and 76,600 tokens against 69,000 for the bare loop.

The hidden-reasoning row is a lower bound. The provider pages cited earlier say that, when you return tool results, “you must pass the thinking blocks from the assistant message back to the API, complete and unmodified.” If those blocks are re-read like any other history, the row climbs toward the ReAct row. Chapter 4 says of hidden reasoning in general: “A model’s private deliberation is billed whether or not you ever display it”.

What is a written thought evidence of?

A written thought is evidence of what the model wrote, which makes it a good pointer and a weak proof. Chapter 3 gives the instruction for an agent’s narration: “treat it as a debugging aid rather than a sworn statement of what the computation did”.

Two studies support the caution, and I read both from the abstract only. Turpin and colleagues (arXiv:2305.04388) report that chain-of-thought explanations “can systematically misrepresent the true reason for a model’s prediction”. Their abstract adds that biasing features cut accuracy by as much as 36% on a suite of 13 tasks while going unmentioned in the explanations.

Lanham and co-authors (arXiv:2307.13702) report from their abstract that models vary in how strongly they condition on their stated reasoning, “sometimes relying heavily on the CoT and other times primarily ignoring it”. Both papers study a single call on question-answering tasks. Neither studies thoughts inside a tool loop.

A tool loop gives you something a single call lacks: the action is observable. One issue thread shows what to compare.

In a 2026 GitHub issue comment, a commenter, srinivasycplf8, described the looping agent in the issue’s log. The log says “Let me fix that by using the correct parameter name” before “sending the exact same index_pattern argument” (issue thread). One user’s anecdote proves little, and it still names the check.

So grade the pair. When a thought claims a change, compare the arguments of the next call with the previous ones. When a thought claims a fact, find the observation it came from.

Chapter 4 compresses the habit: “read a trace as a map of places to dig, never as a proof of correctness”. The glossary entry for a trace defines the record you are reading.

Which reasoning shape fits which task?

Four task properties decide the shape: whether a needed fact is missing, whether the steps are knowable up front, whether a result decides the next step, and what a wrong step costs to undo. Chapter 4 reduces the choice to two questions. The first is “what does it cost to take a wrong step back?” The second is “do you already know the steps?”

The book then gives a ladder: “no explicit plan for short tasks with cheap steps, interleaved for medium tasks whose right next move genuinely depends on the last result, plan-plus-replan for long tasks with dependencies and expensive mistakes, parallel graphs when the subtasks are independent and latency matters”.

My own arrangement of that ladder is the table below. The book supplies the ladder, the two questions and the first row. The last row and most of the “what breaks” column come from the studies above. Read it top to bottom and stop at the first row whose condition holds.

Shape Choose it when What it costs What breaks Anchor Your task
No written reasoning One step, and the model already holds the answer or one routine call fetches it One call Added deliberation “wastes money” on easy inputs Chapter 4’s reasoning dial one easy step
Chain of thought in one call Several dependent reasoning steps, and no fact is missing One longer call A missing fact gets invented: 56% of chain-of-thought failures in Yao’s Table 2 Wei et al., 2022; Chapter 4 reasoning with no lookup
Plan with placeholders or a graph, run by code Steps are knowable up front, cheap to undo, and linked only by values passed forward Fewest tokens of the tool-using shapes in Xu’s and Kim’s tables Blind to a surprise; Xu’s own limit for unknown environments Xu et al., 2023; Kim et al., 2023 steps known and cheap to undo
Plan, execute, replan Steps are mostly knowable, and a wrong step is expensive to undo A planner call plus one call per replan The replan check that nobody built Chapter 4’s ladder; Wang et al., 2023 costly to undo
Interleaved, the ReAct schedule Each result decides the next step, or the environment is unseen A model call per action and a growing re-read Repetition: 47% of ReAct failures in Yao’s Table 2 were reasoning errors Yao et al., 2022; Chapter 3 result decides next step
Combine and back off You cannot tell which regime a question is in Two mechanisms to maintain Heuristics tested on two benchmarks in one paper Yao’s back-off heuristics can’t tell

If your steps are fully known and you can draw the flowchart, you may want no agent shape at all. The post on agentic workflow patterns such as chains, routers and parallel calls covers that territory.

How does a debugging task land in the table?

An exploratory debugging task lands on the interleaved row, and walking it shows how the table is meant to be used. Take “find out why last night’s export job failed”, with read-only tools for logs, files and queries.

Row one fails: the task needs more than one step. Row two fails: the model lacks the log, so a fact is missing. Row three fails: nobody can list the steps before the first error message is read. Row four fails: reading a log costs nothing to undo.

Row five holds, since each result decides the next step. If the fix that follows is a risky change, treat it as a second task. That task lands on row four. I walked two other tasks before finishing: a lookup answerable from memory lands on row one, and five independent supplier lookups land on row three.

Why did the text parser break, and what replaced it?

The text parser broke because it asked a model to follow a grammar it had seen in a few examples only. Your code then had to recover structure from free text. A model could write prose where an action name belonged, or keep going and invent the observation itself.

A Hacker News commenter, zby, described the burden in 2024: without function calling “you have to write your own grammar for the communication with the LLM, a parser for it, and also you need to teach the LLM to use that grammar” (comment).

What replaced it is a division of labor. The model is trained to emit an action as a structured object with a tool name and arguments that fit a schema. The harness validates that object before running anything. The reasoning, when present, travels in its own channel instead of sharing a string with the action.

Nothing about the schedule changed in that move. The grammar became the provider’s problem, and three questions stayed with you: how often to look, what to write down, and what to pay.

What does none of this evidence cover?

None of the sources read for this post measured written thoughts against none under native tool calling with a built-in reasoning mode. Every number above comes from models and prompt formats of 2022 to 2024. I did not search exhaustively, so read this as a statement about these sources.

Three more gaps deserve a name. No source here measured the faithfulness of thoughts inside a tool loop. No source isolated the token cost of the thought from the cost of re-reading observations, which is why my table is arithmetic. And the plan-first comparisons ran on short question-answering tasks and games, with no report of what the saving shrinks to once a replanning step is included.

The book’s own summary of interleaved planning carries both sides in one sentence: “Maximally adaptive, because no commitment outlives one observation; and structurally myopic, because no pass ever forces the model to hold the whole task in view.” Its verdict on the alternative is just as short: “A plan with no replanning step is a guess with a schedule.”

My position is modest. Start from the task properties, price the shapes with your own token counts, and test the choice on your own traces. The benchmarks above tell you which failures to expect from each shape. They cannot tell you the rates on your backlog.

If someone on your team asks what is the ReAct agent pattern, a fair short answer has three parts. It is a schedule of thought, action and observation. Its text format has mostly retired. And its value on a given task depends on whether the next step waits on the last result.

Chapters 1 and 2 and the glossary, including the entry on ReAct, are free to read online. The interleaving argument is in Chapter 3 and the planning ladder in Chapter 4, both in the full book. The agent fundamentals guide places this post beside its neighbors.

The book is sold as a Kindle edition and a paperback on Amazon, and the guide to choosing between the Kindle and paperback editions compares the two. You can also see the formats.

Questions readers ask

Is ReAct the same as tool calling?
No. Tool calling is the interface for one action: the model emits a structured request and your code runs it. ReAct is the schedule around that interface: think a little, act once, read the result, think again. The original parsed text format has mostly disappeared, and the schedule is what a tool-calling loop runs by default.
Is the ReAct pattern obsolete?
The Thought/Action/Observation text format largely is, because current models are typically trained to reason before acting and to emit structured tool calls. The schedule is alive in every loop that calls a model again after each tool result. Whether an extra written thought still earns its tokens when a model has a built-in reasoning mode was measured by none of the sources read for this post.
What is the difference between chain of thought and ReAct prompting?
Chain of thought writes intermediate reasoning inside one call and never looks anything up. ReAct alternates reasoning with tool calls, so each thought can use a fresh observation. In the original ReAct paper each approach won one of two knowledge benchmarks, and combinations of the two scored highest on both (Yao et al., 2022).
Is plan-and-execute cheaper than ReAct?
It can be, when plan steps are executed by code with no model call between them. One paper reported 64% fewer tokens averaged over six benchmarks (Xu et al., 2023) and another up to 6.7 times lower cost on tasks built from independent lookups (Kim et al., 2023). A plan with a model-driven executor per step can cost more calls than the loop it replaces.
Can I trust the thoughts in an agent's trace?
Use them as a pointer to where to look. Two studies of single-call chain of thought, read here from their abstracts, found that stated reasoning can omit the real cause of an answer and is sometimes ignored by the model (Turpin et al., 2023; Lanham et al., 2023). Check a thought against the action and the observation beside it.

Sources

  1. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao (2022). ReAct: Synergizing Reasoning and Acting in Language Models
  2. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
  3. Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, Ee-Peng Lim (2023). Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
  4. Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, Dongkuan Xu (2023). ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models (arXiv:2305.18323)
  5. Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W. Mahoney, Kurt Keutzer, Amir Gholami (2023). An LLM Compiler for Parallel Function Calling
  6. Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati (2024). On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models (arXiv:2405.13966)
  7. Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting (arXiv:2305.04388; abstract read)
  8. Tamera Lanham and co-authors (2023). Measuring Faithfulness in Chain-of-Thought Reasoning (arXiv:2307.13702; abstract read)
  9. Anthropic (2026). Thinking (provider documentation, read 2026-10-07; a dated example of a built-in reasoning mode)
  10. zby (Hacker News handle) (2024). Hacker News comment on function calling and the original ReAct text format
  11. srinivasycplf8 (GitHub handle) (2026). GitHub issue comment, 2026-04-14: an agent's thought claims a fix while the call repeats