Phase 4: AI Agents & Autonomous Systems

ReAct pattern: reasoning & acting in an interleaved loop

Advanced ~13 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you’re trying to bake your favorite cookies, but instead of following a recipe, you just try to guess all the steps at once. You might end up with something completely different, or even a burnt mess, right? That’s a bit like how some early computer programs worked – they’d try to give you an answer all in one go, even if they didn't really know how to get there or couldn't use any tools, like an oven or a mixer. But what if your computer program could think and act more like a super-smart baker?

That’s where something called the "Reasoning and Acting" pattern comes in. Think of it like this: your computer brain acts like a chef who's making a complex cake. First, the chef has a Thought: "Okay, I need to make a delicious chocolate cake. What's step one?" Then, the chef takes an Action: "I'll open the recipe book to see what it says." Next, the chef makes an Observation: "Ah, the recipe says to preheat the oven to 350 degrees."

But it doesn't stop there! The chef doesn't just give you the cake now. They use that observation to fuel their next Thought: "Right, oven needs preheating. I should go set the temperature." Then they take another Action (turn on the oven), make another Observation (oven light comes on, it starts heating up), and so on. It’s a constant back-and-forth between thinking, doing, and checking. This systematic, step-by-step way of working is much better than just guessing.

This thoughtful way of working is super powerful for building AI programs. Instead of an AI just saying it baked a cake (which might just be a made-up answer!), it actually goes through the process. It knows how to use "tools" like searching the internet for ingredients, checking a calendar for delivery times, or even doing math. If something goes wrong – like the oven isn't working – the AI's "observation" step tells it, and it can "think" about what to do next, maybe find a different recipe or ask for help.

So, when you build computer programs that use this "Reasoning and Acting" idea, it means your programs aren't just guessing or making up answers. They become like truly capable, helpful assistants. This means you can build AI that can figure out problems, use all sorts of digital tools reliably, and give you much better, more trustworthy answers because they've actually done the work, step by step, just like a master baker follows a recipe to perfection.

The ReAct pattern comes from a 2022 paper by Yao et al., but the idea is straightforward enough to re-derive yourself. A plain LLM call is stateless: you give it a prompt, it generates tokens, done. If the answer requires fetching live data, running code, or chaining multiple lookups, a single call either hallucinates or admits ignorance. ReAct solves this by turning one LLM call into a loop. Each iteration appends the previous Thought, Action, and Observation to the context window, giving the model a running scratchpad. The model keeps generating until it produces a special "Final Answer" token sequence instead of another tool call. Your orchestration code parses each generation, routes it to the right tool, and appends the result before looping.

Concretely, the prompt template defines four labelled tokens: Thought:, Action:, Action Input:, and Observation:. The model generates the first three; your code generates the fourth by actually calling the tool. A minimal hand-rolled loop looks like: build initial prompt, loop { call LLM, parse output, if Final Answer return, else execute tool, append observation, continue }. That's it. LangChain's AgentExecutor, LlamaIndex's ReActAgent, and LangGraph's graph nodes all implement exactly this, just with better error handling, streaming support, and callback hooks bolted on top. When you understand the raw loop, framework magic becomes readable source code.

Here is a real-world scenario: a customer support agent needs to answer billing questions. The user asks: "Why did my invoice jump 40% this month?" A single LLM call would hallucinate a reason. A ReAct agent instead reasons ("I need to look up this customer's usage delta"), calls a billing API tool, gets back raw JSON, reasons again ("the spike is in vector DB queries, not compute"), and only then drafts a natural-language explanation. A senior engineer would structure the tools as narrow, single-responsibility functions, each returning structured JSON, not prose. Wide tools that return blobs of text make the model's job harder and your parsing logic fragile.

Tradeoffs versus alternatives: Plan-and-Execute (or ReWOO) generates the full plan first, then executes all steps without re-querying the LLM between steps. This is faster and cheaper when the task is well-defined and tools are reliable, because you pay for one planning call instead of N reasoning calls. ReAct wins when tools can fail, return unexpected results, or when downstream steps depend on upstream results in ways you can't predict upfront. Function-calling-only agents (no explicit Thought step) reduce token usage but sacrifice the interpretable reasoning trace that makes debugging possible. For production agents where you need audit logs, the Thought trace is not overhead, it is the artifact you ship to your compliance team.

Scale changes the calculus significantly. At 10 users, a synchronous Python loop is fine and verbose=True helps you debug. At 10,000 users you need async execution (asyncio + aiohttp for tool calls), streaming responses so the UI doesn't time out, and per-request token budgets enforced before the loop starts. At 10 million users, the individual ReAct loop is not your bottleneck, tool latency is. A web search call that takes 800ms inside a 6-iteration loop adds nearly 5 seconds of latency. You cache tool results aggressively (Redis with a short TTL), run independent tool calls in parallel when the Thought step indicates multiple lookups, and precompute common retrieval paths. You also instrument each loop iteration with structured logs so you can identify which tool is the p99 latency culprit without reproducing the full conversation.

Key Takeaways

  • ReAct loops Thought, Action, and Observation until the model emits a final answer, not a tool call.
  • Every framework (LangChain, LlamaIndex, LangGraph) is a thin layer on top of this same loop.
  • Always enforce a max-iteration cap; uncapped loops burn tokens and money.
  • Parse tool calls from structured output, not free text, to survive prompt injection and format drift.

Pro tips

  • Set max_iterations based on your tool call graph depth, not an arbitrary number. If a task needs at most 3 tool calls, cap at 5 so you have one retry budget per tool but don't spiral. A cap of 20 on a simple agent is a silent money leak.
  • Use handle_parsing_errors=True (LangChain) or equivalent in every framework. Without it, a single malformed Thought/Action block kills the entire request. With it, the model gets the parse error as an Observation and usually self-corrects.
  • The Thought text is a first-class debugging artifact. In production, store it in your trace store (LangSmith, Langfuse, or a plain Postgres JSONB column) alongside the session ID. You will need it when a user escalates a wrong answer.
  • When a tool consistently returns noisy or long responses, add a summarization step inside the tool wrapper rather than passing the raw blob to the model. Keeping Observations under 300 tokens per call noticeably improves reasoning quality and cuts costs on long sessions.

Common pitfalls

  • Mistake: No iteration cap, so a confused agent loops until the context window fills or budget runs out. Fix: Always set max_iterations and max_execution_time; treat them as required, not optional parameters.
  • Mistake: Parsing the Action from free text with regex, which breaks on any format variation. Fix: Use structured output / function calling for the Action step so the tool name and input are typed, not scraped.
  • Mistake: Returning raw HTTP error bodies as Observations, which confuses the model and leaks internal URLs. Fix: Wrap every tool to return a normalized error string like {"error": "service unavailable"} on failure.
  • Mistake: Running tool calls synchronously inside a blocking loop, causing timeouts at load. Fix: Use async tool wrappers with asyncio.gather when the Thought step implies independent parallel lookups.

When to use ReAct vs alternative agent execution patterns

Option Use when Avoid when
ReAct (interleaved reasoning) Steps are interdependent, tools can fail, or you need an audit trail of the model's reasoning. Task is well-scoped with reliable tools and you need minimum token cost.
Plan-and-Execute (ReWOO) You can enumerate all needed tool calls upfront and tools are stable; saves N-1 LLM calls. Later steps depend on the exact output of earlier ones in unpredictable ways.
Single function-calling call One tool call is sufficient; no multi-step reasoning needed; latency is critical. Task requires iterative refinement or the model needs to react to intermediate results.
Multi-agent handoff (CrewAI / LangGraph) Subtasks are parallel and specialist agents outperform a generalist on each domain. Overhead of agent-to-agent communication exceeds the benefit; single ReAct loop suffices.

Code Example

python
# langchain >= 0.2, openai >= 1.0
from langchain.agents import create_react_agent, AgentExecutor
from langchain_openai import ChatOpenAI
from langchain_community.tools import DuckDuckGoSearchRun
from langchain import hub

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
tools = [DuckDuckGoSearchRun()]

# Pull the canonical ReAct prompt from LangChain Hub
prompt = hub.pull("hwchase17/react")

agent = create_react_agent(llm, tools, prompt)
executor = AgentExecutor(agent=agent, tools=tools, max_iterations=6, verbose=True)

result = executor.invoke({"input": "What is the current Python version and when was it released?"})
print(result["output"])

How this code works

This code showcases the ReAct pattern, enabling an AI agent to intelligently reason and take actions to answer complex questions. Its job is to set up an agent that can dynamically use tools, specifically an internet search engine, to find information like the current Python version and its release date, by breaking down the problem and executing steps.

The setup begins by importing components such as create_react_agent, AgentExecutor, ChatOpenAI for the language model, and DuckDuckGoSearchRun for web searching. An llm (large language model) based on gpt-4o-mini provides the agent's "brain," and tools defines its available actions, like searching. Crucially, the ReAct prompt is fetched from langchain hub; this isn't just an instruction but a structured template that teaches the LLM its "Thought, Action, Observation" reasoning loop—a subtle yet vital detail for beginners to understand how the agent thinks. The create_react_agent function combines these elements into an agent. Finally, AgentExecutor runs this agent, managing its interleaved reasoning and tool use, with max_iterations limiting its steps and verbose=True revealing its internal thought process as it arrives at the result["output"].

Production-grade example

Adds retry with backoff, per-call token logging, wall-clock timeout, graceful degradation, and structured logs.

python
# langchain >= 0.2, openai >= 1.0, tenacity >= 8.0
import os
import logging
import time
from typing import Any
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
from langchain.agents import create_react_agent, AgentExecutor
from langchain_openai import ChatOpenAI
from langchain_community.tools import DuckDuckGoSearchRun
from langchain.callbacks.base import BaseCallbackHandler
from langchain import hub
import openai

logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

class TokenLogger(BaseCallbackHandler):
    """Log token usage and latency per LLM call."""
    def on_llm_end(self, response, **kwargs: Any) -> None:
        usage = getattr(response, "llm_output", {}).get("token_usage", {})
        logger.info("llm_call tokens=%s", usage)

@retry(
    retry=retry_if_exception_type((openai.RateLimitError, openai.APITimeoutError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
    reraise=True,
)
def run_agent(user_input: str, session_id: str) -> str:
    llm = ChatOpenAI(
        model="gpt-4o-mini",
        temperature=0,
        timeout=20,
        api_key=os.environ["OPENAI_API_KEY"],
        callbacks=[TokenLogger()],
    )
    tools = [DuckDuckGoSearchRun()]
    prompt = hub.pull("hwchase17/react")
    agent = create_react_agent(llm, tools, prompt)
    executor = AgentExecutor(
        agent=agent,
        tools=tools,
        max_iterations=8,
        max_execution_time=45,  # hard wall-clock timeout
        handle_parsing_errors=True,  # graceful degradation on malformed output
        verbose=False,
    )
    start = time.perf_counter()
    try:
        result = executor.invoke({"input": user_input})
        logger.info("session=%s latency=%.2fs status=ok", session_id, time.perf_counter() - start)
        return result["output"]
    except Exception as exc:
        logger.error("session=%s latency=%.2fs status=error error=%s", session_id, time.perf_counter() - start, exc)
        return "I was unable to complete that request. Please try again."

if __name__ == "__main__":
    answer = run_agent("What Python version is current and when was it released?", session_id="demo-001")
    print(answer)

How this code works

This code showcases the ReAct pattern, enabling an AI agent to intelligently reason and act by using tools to answer questions. Its primary job is to demonstrate a robust, production-grade agent that can search the web and respond to user queries, complete with logging and error handling. The core logic resides in the run_agent function, which orchestrates the entire process. It initializes a ChatOpenAI Large Language Model (LLM) using gpt-4o-mini, sets up a DuckDuckGoSearchRun tool for web access, and pulls a specialized ReAct prompt from langchain.hub to guide the agent's decision-making. These components are then combined by create_react_agent to define the agent's reasoning capabilities.

The AgentExecutor is responsible for bringing the agent to life, iteratively deciding whether to "think" or "act" by using its available tools until a final answer is reached. A custom TokenLogger provides valuable insights into LLM usage and costs by logging token_usage after each call. For resilience, the run_agent function is decorated with @retry from tenacity, automatically retrying calls that encounter openai.RateLimitError or openai.APITimeoutError. A subtle yet crucial detail for robust agents is handle_parsing_errors=True within AgentExecutor. This prevents the agent from crashing if the LLM generates slightly malformed output when trying to use a tool, allowing for graceful degradation rather than an abrupt failure, which is vital for real-world applications.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a ReAct agent that answers: 'What is the square root of the population of France?' It must use two tools: a web search tool (DuckDuckGo or a mock) and a calculator tool you write yourself. The agent should search for France's population, extract the number, then call the calculator to get the square root. Log every Thought and Observation to stdout.

python
# langchain >= 0.2, openai >= 1.0
import os
import math
from langchain.agents import create_react_agent, AgentExecutor
from langchain_openai import ChatOpenAI
from langchain_community.tools import DuckDuckGoSearchRun
from langchain.tools import tool
from langchain import hub

# TODO 1: Define a @tool called 'calculator'
# It should accept a math expression string and return the evaluated result.
# Hint: use Python's eval() with a restricted namespace, or math.sqrt directly.

# TODO 2: Create the LLM (gpt-4o-mini, temperature=0)

# TODO 3: Assemble the tools list (DuckDuckGoSearchRun + your calculator)

# TODO 4: Pull the react prompt from hub and create the agent + executor
# Set max_iterations=8 and verbose=True

# TODO 5: Invoke the agent with the question below
QUESTION = "What is the square root of the population of France?"
# result = executor.invoke({"input": QUESTION})
# print(result["output"])

Quick check

  1. In a ReAct loop, what signals the orchestrator to stop iterating and return a response to the user?

  2. A ReAct agent is hitting the max_iterations cap on a task that should only need 3 tool calls. What is the most likely root cause?

  3. Why does the Plan-and-Execute pattern use fewer LLM tokens than ReAct for the same task?

Self-check: Sketch the raw ReAct loop in pseudocode without referencing any framework. Then name one scenario where Plan-and-Execute would be a better choice than ReAct, and explain why the Thought trace is worth storing in production even after the task completes.