Phase 4: AI Agents & Autonomous Systems

Tool schemas for APIs, databases, calculators & search engines

Intermediate ~15 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you have a super-smart robot chef who can make almost anything you ask for, from a simple sandwich to a fancy birthday cake. But here's the trick: this chef isn't magic! It needs to know exactly how to do things. If you just say, "make me a yummy cake," it might not know what kind of cake, what ingredients to use, or even where to find the oven. That's where something called a "tool schema" comes in.

Think of a tool schema like a super detailed recipe book for your robot chef. This book doesn't just list "cake"; it has an entry for "Chocolate Fudge Cake" and inside it says exactly: "To make Chocolate Fudge Cake, you need: Flour (how much?), Sugar (how much?), Eggs (how many?), and Cocoa Powder (how much?)." It even says which ingredients are super important (like flour for a cake) and which are optional. It's like a precise instruction manual that tells your robot chef exactly what 'cooking actions' it knows how to do, what 'ingredients' each action needs, and what each ingredient is for. Without this exact recipe, if you just said "make a yummy cake," the chef might get confused and try to use pickles instead of sugar, or try to bake something that doesn't exist in its kitchen!

In the world of computers, instead of a robot chef, we have really smart computer programs called "Large Language Models" (LLMs) – those are the fancy AIs that can chat with you and answer questions. And instead of recipes for food, these LLMs use tool schemas for things like talking to other computer programs (we call these "APIs" – Application Programming Interfaces), asking a database for information, or even using a calculator. The tool schema tells the LLM, "Hey, here's how you can ask for the weather: you need a city name and a date." Or, "Here's how you can search the internet: tell me what you want to look for."

So, when you learn to create these precise "recipe books" – these tool schemas – you're teaching your AI exactly how to use all the tools it has. This means your AI won't guess or make mistakes when it tries to use a calculator, look up facts, or even help you build new things online. Getting these "recipes" just right is the secret to making your AI helpful and accurate, so it always knows what to do and how to do it without getting mixed up!

A tool schema in the OpenAI format is a JSON object with three required fields: name, description, and parameters. The parameters field follows JSON Schema draft-07 conventions. The model receives these schemas concatenated into its context window before your user message, which means every schema costs tokens. A verbose schema with 12 parameters and long descriptions can easily consume 300-400 tokens per tool. Keep that in mind when you're assembling a toolkit of 20+ tools.

The mental model that matters here: the LLM is doing a classification and extraction task, not a lookup. It reads your description, the user's message, and the conversation history, then produces a structured prediction of which function to call and what to put in each slot. That means your description field is doing the same work as a retrieval query. If two tools have overlapping descriptions (say, search_documents and query_knowledge_base), the model will flip a coin. Name and describe tools so their decision boundaries are clear. A good test: if you printed only the name and description of each tool to a junior engineer, could they pick the right one for any given user request without seeing the full schema?

For real-world scenarios, consider a customer support agent that can look up orders, issue refunds, and search a help center. Your lookup_order tool should describe exactly what it returns (order status, items, shipping address) so the model knows whether to call it or to call search_help_center instead. The parameters object should use enum constraints wherever the valid values are finite: status_filter: { type: string, enum: [open, shipped, delivered, cancelled] }. Enums do two things: they prevent the model from generating arbitrary strings, and they communicate domain vocabulary that the model might not know from training data alone. For database tools specifically, never expose raw SQL as a parameter. Instead, model your schema as a typed function: query_orders(customer_id: string, status_filter: string, limit: integer). The SQL stays server-side; the model only fills typed slots.

The tradeoffs between different schema design choices are real. A single general-purpose query_database(sql: string) tool is easy to write but catastrophic to ship: the model will attempt arbitrary SQL including writes, and you'll spend weeks patching injection attempts. Narrow, purpose-built tools like get_customer_by_id and list_recent_orders are more tokens but radically safer and more reliable. On the other end, if you split too aggressively, you get tool sprawl and the model spends reasoning tokens deciding between search_orders_by_email vs find_orders_for_customer which do the same thing. Aim for one tool per distinct action, not one tool per possible input shape.

Scale changes the schema problem in two ways. At 10 users, you can afford 20 tools in every request because the extra context cost is noise. At 10 million users sending hundreds of requests per second, every extra tool schema is real money and real latency. The common production pattern is tool routing: a lightweight classifier (could be a small embedding similarity search) picks the 3-5 most relevant tools for a given message and injects only those into the context. Anthropic's approach with Claude recommends fewer than 10 tools per request for reliability. OpenAI's function calling also degrades when the tool list grows past 15-20 because the model's ability to discriminate between similar tools in a long context window weakens. Building a tool registry with metadata that supports semantic search is a standard pattern at scale: each tool has embeddings of its name and description, and at request time you retrieve the top-k tools based on the user message.

Cost and latency implications are straightforward to reason about. Schema tokens are prompt tokens, which are generally cheaper than completion tokens but still count. A toolkit of 10 tools averaging 150 tokens each adds 1500 tokens to every request. At GPT-4o pricing (illustrative: roughly $2.50 per million input tokens), that's fractions of a cent per call, but at 10 million calls per day it becomes meaningful. Latency-wise, the schema tokens add to time-to-first-token since the model must process them. Streaming tool calls (available in the OpenAI API) helps here: you can begin parsing the argument stream before the full response is complete, which lets you validate and prepare for execution in parallel.

Key Takeaways

  • A tool schema is the sole interface the LLM has to your function; precision prevents hallucinated arguments.
  • Mark only truly required parameters as required; optional parameters with defaults reduce failed calls.
  • Rich descriptions on parameters outperform type annotations alone for guiding argument selection.
  • Validate model-generated arguments against the schema before executing any tool call.

Pro tips

  • Put the decision logic in the description, not just the capability. Instead of 'Searches the web', write 'Searches the web for current events or facts that may have changed after the model's training cutoff. Do NOT use for math, reasoning, or historical facts the model already knows.' The model uses this to self-select correctly.
  • Never use additionalProperties: true in your parameter schemas. Locking the schema closed forces the model to use exactly the properties you defined, which prevents hallucinated extra keys that would silently fail or cause downstream errors.
  • When a parameter is an ID that the model can never invent (like a database row ID), mark it required but also document where it comes from: 'The order_id returned by list_orders. Do not fabricate.' This cues the model to chain tool calls rather than guess.
  • JSON Schema constraints like minimum, maximum, minLength, maxLength, and pattern are enforced by the model probabilistically, not guaranteed. Always re-validate model-generated arguments server-side with a library like jsonschema before passing them to real functions.

Common pitfalls

  • Mistake: Using a single generic run_query(sql: string) tool to keep schemas simple. Fix: Model each distinct database action as a separate typed function to prevent arbitrary writes and hallucinated SQL syntax.
  • Mistake: Omitting description from individual parameter fields and relying on the parameter name alone. Fix: Write a one-sentence description for every parameter, especially ones with non-obvious valid values or format constraints.
  • Mistake: Marking every parameter as required, so the model fails to call the tool when optional context is absent. Fix: Only mark parameters required if the function cannot execute without them; provide sensible defaults for everything else.
  • Mistake: Adding 20+ tools to every request regardless of relevance. Fix: Build a tool routing layer that selects 3-8 relevant tools per request using semantic similarity to the user message.

How to model different tool types as schemas

Option Use when Avoid when
REST API tool Calling external services (weather, payments, CRM). Map query params or request body fields directly to schema properties. The API has dynamic fields that change per endpoint; model a per-endpoint schema instead of one generic HTTP tool.
Database query tool (typed) Your agent needs to read structured data. Expose filters and column selectors as typed parameters, keep SQL server-side. You need the model to construct arbitrary cross-table joins; that complexity belongs in application logic, not tool arguments.
Calculator / math tool Precise arithmetic is needed and you cannot trust the model's floating-point reasoning. A simple evaluate(expression: string) wrapping a safe eval is enough. The calculation is simple enough that the model handles it reliably; unnecessary tool calls add latency.
Search engine tool User questions require information beyond the model's training cutoff or internal knowledge base. Your retrieval needs are over private documents; use a RAG tool with vector search instead of a public web search API.
Composite / multi-step tool A repeated sequence of low-level calls can be encapsulated as one higher-level tool to reduce round-trips. The steps have branching logic that depends on intermediate results; keep them separate so the model can adapt between calls.

Code Example

python
# openai>=1.0.0
import openai

client = openai.OpenAI()  # reads OPENAI_API_KEY from env

weather_tool = {
    "type": "function",
    "function": {
        "name": "get_current_weather",
        "description": "Return current temperature and conditions for a city.",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    "description": "City name, e.g. 'Austin, TX'"
                },
                "unit": {
                    "type": "string",
                    "enum": ["celsius", "fahrenheit"],
                    "description": "Temperature unit. Defaults to celsius."
                }
            },
            "required": ["city"]
        }
    }
}

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=[weather_tool],
)
print(response.choices[0].message.tool_calls)

How this code works

This code demonstrates how to teach an OpenAI language model about an external function, enabling it to intelligently decide when and how to "call" that function based on a user's request. It specifically shows how to define a tool for fetching current weather information. The process begins by initializing the openai client to connect with the AI service. The core of the example is the weather_tool dictionary, which outlines the function called get_current_weather. This definition includes a clear description that tells the AI what the tool does ("Return current temperature and conditions for a city.").

Inside weather_tool, the parameters section uses a JSON Schema-like structure to specify the inputs the get_current_weather function expects. It defines city as a string (e.g., 'Austin, TX') and unit as an enum allowing "celsius" or "fahrenheit," making city a required argument. When client.chat.completions.create is called, the model="gpt-4o" is given the user's messages and, crucially, the tools=[weather_tool]. The subtle point here is that the AI doesn't execute the get_current_weather function itself. Instead, the response.choices[0].message.tool_calls output contains a structured representation of the proposed function call, including the function's name and its arguments, which an external application would then use to perform the actual weather lookup.

Production-grade example

Adds schema validation, retry with backoff, timeout, token logging, and structured error paths.

python
# openai>=1.0.0, tenacity>=8.0.0
import os
import json
import time
import logging
import jsonschema
from openai import OpenAI, APITimeoutError, RateLimitError, APIError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

logger = logging.getLogger(__name__)
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)

SEARCH_TOOL = {
    "type": "function",
    "function": {
        "name": "web_search",
        "description": "Search the web for current information. Use when the user asks about recent events, prices, or facts that may have changed after your training cutoff.",
        "parameters": {
            "type": "object",
            "properties": {
                "query": {"type": "string", "description": "Search query, written as a concise phrase."},
                "num_results": {"type": "integer", "minimum": 1, "maximum": 10, "description": "Number of results to return. Default 3."},
            },
            "required": ["query"],
        },
    },
}

ARGUMENT_SCHEMA = {
    "type": "object",
    "properties": {
        "query": {"type": "string"},
        "num_results": {"type": "integer", "minimum": 1, "maximum": 10},
    },
    "required": ["query"],
}

@retry(
    retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
)
def call_with_tools(user_message: str) -> dict:
    start = time.monotonic()
    try:
        response = client.chat.completions.create(
            model="gpt-4o",
            messages=[{"role": "user", "content": user_message}],
            tools=[SEARCH_TOOL],
            tool_choice="auto",
        )
    except APIError as exc:
        logger.error("openai_api_error", extra={"status": exc.status_code, "msg": str(exc)})
        raise

    latency_ms = (time.monotonic() - start) * 1000
    usage = response.usage
    logger.info(
        "llm_call_complete",
        extra={
            "latency_ms": round(latency_ms, 1),
            "prompt_tokens": usage.prompt_tokens,
            "completion_tokens": usage.completion_tokens,
            "model": response.model,
        },
    )

    choice = response.choices[0].message
    if not choice.tool_calls:
        return {"type": "text", "content": choice.content}

    tool_call = choice.tool_calls[0]
    try:
        args = json.loads(tool_call.function.arguments)
        jsonschema.validate(instance=args, schema=ARGUMENT_SCHEMA)
    except (json.JSONDecodeError, jsonschema.ValidationError) as exc:
        logger.warning("invalid_tool_arguments", extra={"raw": tool_call.function.arguments, "error": str(exc)})
        return {"type": "error", "content": "Model produced invalid tool arguments.", "raw": tool_call.function.arguments}

    return {"type": "tool_call", "name": tool_call.function.name, "args": args, "call_id": tool_call.id}

How this code works

This code enables an AI assistant to intelligently decide when to perform a web search and then correctly format that search. It handles the definition of a web_search tool, how the AI uses it, and validates the AI's output to ensure reliability.

The SEARCH_TOOL dictionary describes the web_search tool to the AI, specifying its name, a helpful description for when to use it, and the parameters it expects, like query and an optional num_results. The call_with_tools function orchestrates the interaction. It uses @retry to automatically reattempt API calls if temporary issues like RateLimitError occur, making the system more robust. When client.chat.completions.create is called, tool_choice="auto" allows the AI to decide whether to use the defined SEARCH_TOOL or respond directly with text. If the AI suggests a tool call, the code parses the AI's generated arguments using json.loads and then critically validates them against ARGUMENT_SCHEMA using jsonschema.validate. This explicit validation is a subtle but important production-grade step: even though the AI received the SEARCH_TOOL schema as a guide, this second ARGUMENT_SCHEMA acts as a safety net, catching any malformed or unexpected arguments the AI might "hallucinate" before they cause errors in the actual search.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Define three tool schemas for a customer support agent: one to look up an order by ID, one to list a customer's recent orders by email, and one to search a help center by keyword. Then make a single chat completions request and print which tool the model selects for the message: 'I can't find my order from last Tuesday, my email is [email protected]'.

python
# openai>=1.0.0
import os, json
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

# TODO: Define lookup_order_by_id tool schema
lookup_order_tool = {
    "type": "function",
    "function": {
        "name": "lookup_order_by_id",
        "description": "TODO",
        "parameters": {
            "type": "object",
            "properties": {},  # TODO: add order_id parameter
            "required": [],
        },
    },
}

# TODO: Define list_orders_by_email tool schema
list_orders_tool = {}

# TODO: Define search_help_center tool schema
search_help_tool = {}

tools = [lookup_order_tool, list_orders_tool, search_help_tool]

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "I can't find my order from last Tuesday, my email is [email protected]"}],
    tools=tools,
)

# TODO: Print the name of the tool the model chose and the arguments it provided
tool_call = response.choices[0].message.tool_calls
print(tool_call)

Quick check

  1. A model repeatedly calls a tool with an extra invented argument that doesn't exist in your schema. What is the most likely root cause?

  2. You have 25 tools in your registry. What is the recommended production approach for passing tools to the model?

  3. Why should you validate model-generated tool arguments with jsonschema before executing the function, even though you already defined the schema for the model?

Self-check: Without looking at your notes, write the schema for a create_support_ticket tool that takes a required subject string, a required priority enum of low/medium/high, and an optional customer_id string. Explain why you would or wouldn't mark customer_id as required, and what you'd put in the description field to help the model decide when to call this tool versus a search_existing_tickets tool.