Phase 4: AI Agents & Autonomous Systems

Error handling for tool failures & unexpected results

Intermediate ~15 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you’re a super chef, ready to bake the most amazing birthday cake ever! You have a fantastic recipe, all your ingredients, and a kitchen full of tools: an electric mixer, a fancy oven, measuring cups, and more. When you’re baking, you don’t just follow the recipe blindly. You know that sometimes things go wrong, right? Maybe the mixer suddenly stops working, or you realize you’re out of eggs, or the oven is acting funny and making your cake too hot or not hot enough. If you just stopped every time something unexpected happened, you’d never get to eat that delicious cake!

That’s exactly what "error handling" is about when we build computer programs, especially programs that use other "tools" to get things done. Think of those tools as your kitchen appliances. Sometimes, when a program tries to use another tool (like asking a website for information), that tool might be busy, or broken, or give back a confusing answer. A basic program might just get stuck, waiting forever, or simply crash. But a smart program, like a smart chef, has a plan for when things don't go perfectly.

A good chef learns to tell the difference between a small hiccup (like the mixer cord coming loose – easy to fix!) and a bigger problem (like the oven completely breaking down). They might try plugging the mixer back in, or if that doesn't work, they might grab a whisk and mix by hand instead. If they're out of eggs, they might quickly check if they have an egg substitute, or adapt the recipe to make muffins instead. And they always check the cake to make sure it's actually baked, even if the recipe says 30 minutes, because ovens can be tricky! They don't want to serve a half-baked cake or a burnt one. They adjust their plan based on what actually happens.

So, when you learn about error handling, you're learning how to make your computer programs just as smart as that super chef. You’ll figure out how to teach your programs to notice when a "tool" isn't working right, understand why it failed, and then make a smart decision about what to do next. This means you can build powerful AI programs that don't just give up when something goes wrong, but can figure out how to keep going, find a different way, or even explain what happened, so they can still get the job done without spinning in circles or giving wrong answers.

Tool calls fail in more ways than most developers anticipate on first pass. The failure surface breaks into three layers: the transport layer (DNS failure, TCP timeout, TLS error), the protocol layer (HTTP 4xx/5xx, malformed JSON body, unexpected content-type), and the semantic layer (the request succeeded but the data is wrong, incomplete, or stale). Most beginners handle the middle layer and miss the other two entirely. Senior engineers treat all three as equally likely in production.

The most important architectural decision is error classification. Transient errors -- timeouts, 429 rate limits, 503 service unavailable -- are worth retrying with exponential backoff and jitter. Permanent errors -- 400 bad request, 404 not found, 401 unauthorized, schema validation failures -- should not be retried at all; they signal that the model or your code passed bad inputs, and retrying wastes time and money. When you return error context to the model, include this classification explicitly. A message like {"error": "rate_limited", "retry_after_seconds": 30, "is_transient": true} gives the model enough information to reason about whether to wait and retry, try a different tool, or ask the user what to do next. A generic {"error": "tool failed"} gives it nothing useful.

Semantic validation is where production systems diverge from demo systems. An API can return HTTP 200 with a JSON body that passes schema validation but is functionally wrong: a stock price API returning a ticker's price as 0.0, a geocoding API returning coordinates in the ocean, a database query returning an empty list when the user asked for their order history. These cases require application-level checks. After every tool call, ask: does this result make sense given the input? For numeric results, are values in a plausible range? For lists, does an empty result mean "no data exists" or "the query was wrong"? For timestamps, is the data recent enough to be actionable? Encode these checks as explicit validators, not ad-hoc conditionals scattered through your loop.

A real-world scenario: you're building a financial agent that calls a stock-price tool, a news-search tool, and a portfolio-database tool in sequence. The user asks "should I sell my NVDA position?" The stock price tool returns successfully. The news tool times out. The portfolio tool returns an empty list because the user ID was passed as a string but the API expected an integer. Without error handling, your agent either crashes, halts, or worse, answers the question with incomplete data and doesn't tell the user. With proper error handling: the portfolio tool failure is classified as a permanent input-validation error and the model is informed so it can ask the user to confirm their account ID; the news tool timeout triggers one retry, and if it fails again, the model is told news data is unavailable and it should caveat its answer accordingly. The agent stays useful instead of becoming a liability.

At scale, error handling strategy changes. At 10 users, you can afford synchronous retries inline. At 10,000 users, synchronous retries inside a request create latency spikes and can amplify thundering herd problems against rate-limited APIs. You need async retry logic, circuit breakers (stop calling a tool that has been failing for 60 seconds until it recovers), and timeout budgets that account for the entire agent turn, not just individual tool calls. At 10 million users, you're dealing with SLA management per tool: you need dashboards showing per-tool error rates, p99 latencies, and automatic fallback routing when a tool's error rate exceeds a threshold. The error handling code you write in phase one should be instrumented with metrics from day one, because retrofitting observability into an agentic loop later is painful.

Key Takeaways

  • Classify every tool error as transient or permanent before deciding how to respond.
  • Return structured error context to the model, not generic failure strings.
  • Validate tool output shape and semantic correctness, not just HTTP 200 status.
  • Design explicit fallback strategies so agents degrade gracefully when tools are unavailable.

Pro tips

  • Pass the error classification back to the model in the tool result, not just the error message. When the model sees "is_transient": true, it can reason about waiting; when it sees "is_transient": false, it knows to try a different approach. This is the difference between an agent that recovers intelligently and one that retries endlessly.
  • Set a budget timeout at the agent-turn level, not just per tool call. If your SLA is 10 seconds, a tool that retries three times with 2-4-8s backoff will blow your budget even if every retry eventually succeeds. Track cumulative elapsed time and short-circuit early with a degraded response.
  • Empty results are not the same as errors, but they're just as dangerous. A tool that returns an empty list on a bad query looks successful. Add a layer of validation that checks whether an empty or near-empty result is plausible given the input, and return a distinct "empty_result" error type so the model can ask the user for clarification rather than drawing false conclusions.
  • Circuit breakers belong at the tool level, not the request level. If a tool is returning errors 80% of the time, continuing to call it wastes tokens on the model processing useless tool results. Track rolling error rates per tool and open the circuit (skip the tool, return a cached or fallback result) until the error rate recovers.

Common pitfalls

  • Mistake: Retrying on 400 Bad Request errors because the code treats all errors as transient. Fix: Map HTTP 4xx to permanent errors immediately; only 429, 500, 502, 503, 504 justify retries.
  • Mistake: Returning raw exception stack traces to the model as the error message, bloating context. Fix: Log the full trace server-side; send the model a concise structured dict with error type, a short detail string, and a suggested action.
  • Mistake: Validating only that the HTTP response was 200 and the JSON parsed, ignoring field values. Fix: Add semantic validators for critical fields: type checks, range checks, and non-null assertions on every field your downstream logic depends on.
  • Mistake: Letting a failing tool silently produce a partial answer without informing the user. Fix: Always surface tool unavailability in the final response, even if the agent answers from other sources, so users know the answer may be incomplete.

How to respond to a tool failure

Option Use when Avoid when
Retry with exponential backoff Error is transient: timeout, 429, 503. You have remaining turn budget. Error is permanent (4xx, schema failure). Retry budget is exhausted. Turn timeout is close.
Return structured error to model and let it replan The model has an alternative tool or can ask the user for clarification to fix the input. The model has no alternatives and the task is impossible without this tool; just tell the user directly.
Use a fallback tool or cached result A lower-fidelity alternative exists (e.g., cached price, secondary data provider). Staleness or lower quality would produce a misleading answer without clear caveats to the user.
Abort the tool call and inform the user The tool is down, no fallback exists, and proceeding without it would produce an unreliable answer. The tool result is optional and the agent can still produce a useful partial answer.

Code Example

python
# openai>=1.30, requests>=2.31
import requests
import json

def search_weather(city: str) -> dict:
    """Tool wrapper with basic error handling."""
    try:
        resp = requests.get(
            "https://api.open-meteo.com/v1/forecast",
            params={"latitude": 0, "longitude": 0, "current_weather": True},
            timeout=5,
        )
        resp.raise_for_status()
        data = resp.json()
        # Validate expected fields exist
        if "current_weather" not in data:
            return {"error": "missing_field", "detail": "Response lacked 'current_weather' key"}
        return {"success": True, "weather": data["current_weather"]}
    except requests.exceptions.Timeout:
        return {"error": "timeout", "detail": "Weather API did not respond within 5 seconds"}
    except requests.exceptions.HTTPError as e:
        return {"error": "http_error", "status_code": e.response.status_code, "detail": str(e)}
    except Exception as e:
        return {"error": "unexpected", "detail": str(e)}

How this code works

This search_weather function serves as a robust tool to fetch weather data from an external API, designed to handle various issues gracefully. It starts with a try...except block, a fundamental pattern for writing reliable code by anticipating and responding to potential failures. Inside the try block, it uses requests.get to make an HTTP call to a weather service. A crucial detail is timeout=5, which ensures the function won't wait indefinitely if the API is slow. resp.raise_for_status() immediately converts any HTTP errors (like 404 Not Found or 500 Server Error) into a Python exception. After getting the data, the code checks if the expected "current_weather" key exists in the JSON response, returning a specific error if it's missing. A subtle point for beginners: despite taking a city argument, the latitude and longitude are hardcoded to 0 (the equator), meaning this function currently fetches weather for a fixed location, not the specified city.

The except blocks define how different problems are handled. If the timeout is hit, requests.exceptions.Timeout catches it, returning an error dictionary indicating the API didn't respond in time. Network or server-side issues that result in bad HTTP status codes are caught by requests.exceptions.HTTPError, providing the status_code for debugging. Finally, a broad except Exception acts as a catch-all for any other unforeseen problems, ensuring the program doesn't crash unexpectedly. Each error case returns a clear, structured dictionary, distinguishing between success and various types of failures for easier interpretation by the calling system.

Production-grade example

Typed error classification, tenacity retries only on transient errors, semantic validation, structured logging.

python
# openai>=1.30, tenacity>=8.2, structlog>=24.0
import os
import time
import structlog
import requests
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

log = structlog.get_logger()

TRANSIENT_STATUS_CODES = {429, 500, 502, 503, 504}

class TransientToolError(Exception):
    pass

class PermanentToolError(Exception):
    pass

def _classify_http_error(status_code: int, detail: str) -> None:
    if status_code in TRANSIENT_STATUS_CODES:
        raise TransientToolError(f"HTTP {status_code}: {detail}")
    raise PermanentToolError(f"HTTP {status_code}: {detail}")

@retry(
    retry=retry_if_exception_type(TransientToolError),
    stop=stop_after_attempt(3),
    wait=wait_exponential(multiplier=1, min=1, max=10),
    reraise=True,
)
def call_stock_price(ticker: str, timeout: float = 5.0) -> dict:
    api_key = os.environ["STOCK_API_KEY"]  # fail fast if missing
    start = time.monotonic()
    try:
        resp = requests.get(
            "https://api.example.com/v1/quote",
            params={"symbol": ticker},
            headers={"Authorization": f"Bearer {api_key}"},
            timeout=timeout,
        )
        latency_ms = (time.monotonic() - start) * 1000
        log.info("tool_call", tool="stock_price", ticker=ticker,
                 status=resp.status_code, latency_ms=round(latency_ms, 1))

        if not resp.ok:
            _classify_http_error(resp.status_code, resp.text[:200])

        data = resp.json()
        price = data.get("price")
        if price is None:
            raise PermanentToolError("Response missing 'price' field")
        if not isinstance(price, (int, float)) or price <= 0:
            raise PermanentToolError(f"Implausible price value: {price}")

        return {"success": True, "ticker": ticker, "price": price,
                "currency": data.get("currency", "USD")}

    except requests.exceptions.Timeout:
        log.warning("tool_timeout", tool="stock_price", ticker=ticker)
        raise TransientToolError(f"Timeout after {timeout}s")
    except TransientToolError:
        raise
    except PermanentToolError:
        raise
    except Exception as exc:
        log.error("tool_unexpected", tool="stock_price", ticker=ticker, error=str(exc))
        raise PermanentToolError(f"Unexpected error: {exc}") from exc

def safe_call_stock_price(ticker: str) -> dict:
    """Returns a structured result suitable for passing back to the model."""
    try:
        return call_stock_price(ticker)
    except TransientToolError as e:
        return {"error": "transient", "is_retryable": False,
                "detail": str(e), "suggestion": "Inform user the service is temporarily unavailable."}
    except PermanentToolError as e:
        return {"error": "permanent", "is_retryable": False,
                "detail": str(e), "suggestion": "Check ticker symbol or API credentials."}

How this code works

This code provides a robust system for calling external APIs, specifically fetching stock prices, making it ideal for AI tools needing reliable data. Its primary job is to handle various failure scenarios by classifying errors as either TransientToolError (temporary, worth retrying) or PermanentToolError (fatal, no retry needed). The safe_call_stock_price function wraps this complex logic, converting internal errors into structured, user-friendly results that an AI model can easily process and act upon, improving the overall resilience and user experience of AI-driven applications. It also uses structlog for detailed logging (log.info, log.warning) for observability.

The core call_stock_price function uses the tenacity library's @retry decorator. This automatically re-attempts the API call up to three times with wait_exponential backoff if a TransientToolError occurs. Inside, it calls the stock API using requests.get. If the HTTP response isn't successful, the _classify_http_error helper checks the status code to determine if it's transient (e.g., 429, 500-504) or permanent. It also validates the response content, raising PermanentToolError if data like price is missing or invalid. A subtle but crucial detail is the timeout parameter in requests.get; it prevents the tool from hanging indefinitely on slow networks, ensuring a TransientToolError is raised quickly for tenacity to retry. Finally, safe_call_stock_price catches these classified errors and formats them into a dictionary, including a suggestion for the AI.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a tool wrapper for a hypothetical currency-conversion API. It should: classify HTTP errors as transient or permanent, retry transient errors up to 3 times with backoff, validate that the returned rate is a positive float, and return a structured dict (success or error) suitable for passing back to a model in a tool-call loop.

python
# openai>=1.30, tenacity>=8.2, requests>=2.31
import requests
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

TRANSIENT_CODES = {429, 500, 502, 503, 504}

class TransientToolError(Exception):
    pass

class PermanentToolError(Exception):
    pass

# TODO: Add @retry decorator that only retries TransientToolError,
#       stops after 3 attempts, waits 1-8s exponentially
def call_currency_api(from_currency: str, to_currency: str) -> dict:
    try:
        resp = requests.get(
            "https://api.example.com/convert",
            params={"from": from_currency, "to": to_currency},
            timeout=4.0,
        )
        # TODO: classify non-2xx responses using TRANSIENT_CODES
        # TODO: parse JSON and validate 'rate' field is a positive float
        # TODO: return success dict
        pass
    except requests.exceptions.Timeout:
        # TODO: raise TransientToolError
        pass

def safe_convert(from_currency: str, to_currency: str) -> dict:
    # TODO: wrap call_currency_api, catch both error types,
    #       return structured error dicts for each
    pass

Quick check

  1. A tool call returns HTTP 429. What is the correct immediate response in a production agent?

  2. A weather tool returns HTTP 200 with JSON {"temp": null, "unit": "C"}. What should your tool wrapper do?

  3. Why should you send a structured error dict back to the model rather than raising an unhandled exception when a tool fails?

Self-check: Describe the three failure layers in a tool call (transport, protocol, semantic). For each layer, name one error type, whether it is transient or permanent, and what structured information you would return to the model to enable intelligent replanning.