A tool schema in the OpenAI format is a JSON object with three required fields: name, description, and parameters. The parameters field follows JSON Schema draft-07 conventions. The model receives these schemas concatenated into its context window before your user message, which means every schema costs tokens. A verbose schema with 12 parameters and long descriptions can easily consume 300-400 tokens per tool. Keep that in mind when you're assembling a toolkit of 20+ tools.
The mental model that matters here: the LLM is doing a classification and extraction task, not a lookup. It reads your description, the user's message, and the conversation history, then produces a structured prediction of which function to call and what to put in each slot. That means your description field is doing the same work as a retrieval query. If two tools have overlapping descriptions (say, search_documents and query_knowledge_base), the model will flip a coin. Name and describe tools so their decision boundaries are clear. A good test: if you printed only the name and description of each tool to a junior engineer, could they pick the right one for any given user request without seeing the full schema?
For real-world scenarios, consider a customer support agent that can look up orders, issue refunds, and search a help center. Your lookup_order tool should describe exactly what it returns (order status, items, shipping address) so the model knows whether to call it or to call search_help_center instead. The parameters object should use enum constraints wherever the valid values are finite: status_filter: { type: string, enum: [open, shipped, delivered, cancelled] }. Enums do two things: they prevent the model from generating arbitrary strings, and they communicate domain vocabulary that the model might not know from training data alone. For database tools specifically, never expose raw SQL as a parameter. Instead, model your schema as a typed function: query_orders(customer_id: string, status_filter: string, limit: integer). The SQL stays server-side; the model only fills typed slots.
The tradeoffs between different schema design choices are real. A single general-purpose query_database(sql: string) tool is easy to write but catastrophic to ship: the model will attempt arbitrary SQL including writes, and you'll spend weeks patching injection attempts. Narrow, purpose-built tools like get_customer_by_id and list_recent_orders are more tokens but radically safer and more reliable. On the other end, if you split too aggressively, you get tool sprawl and the model spends reasoning tokens deciding between search_orders_by_email vs find_orders_for_customer which do the same thing. Aim for one tool per distinct action, not one tool per possible input shape.
Scale changes the schema problem in two ways. At 10 users, you can afford 20 tools in every request because the extra context cost is noise. At 10 million users sending hundreds of requests per second, every extra tool schema is real money and real latency. The common production pattern is tool routing: a lightweight classifier (could be a small embedding similarity search) picks the 3-5 most relevant tools for a given message and injects only those into the context. Anthropic's approach with Claude recommends fewer than 10 tools per request for reliability. OpenAI's function calling also degrades when the tool list grows past 15-20 because the model's ability to discriminate between similar tools in a long context window weakens. Building a tool registry with metadata that supports semantic search is a standard pattern at scale: each tool has embeddings of its name and description, and at request time you retrieve the top-k tools based on the user message.
Cost and latency implications are straightforward to reason about. Schema tokens are prompt tokens, which are generally cheaper than completion tokens but still count. A toolkit of 10 tools averaging 150 tokens each adds 1500 tokens to every request. At GPT-4o pricing (illustrative: roughly $2.50 per million input tokens), that's fractions of a cent per call, but at 10 million calls per day it becomes meaningful. Latency-wise, the schema tokens add to time-to-first-token since the model must process them. Streaming tool calls (available in the OpenAI API) helps here: you can begin parsing the argument stream before the full response is complete, which lets you validate and prepare for execution in parallel.
Key Takeaways
- A tool schema is the sole interface the LLM has to your function; precision prevents hallucinated arguments.
- Mark only truly required parameters as required; optional parameters with defaults reduce failed calls.
- Rich descriptions on parameters outperform type annotations alone for guiding argument selection.
- Validate model-generated arguments against the schema before executing any tool call.
Pro tips
- Put the decision logic in the description, not just the capability. Instead of 'Searches the web', write 'Searches the web for current events or facts that may have changed after the model's training cutoff. Do NOT use for math, reasoning, or historical facts the model already knows.' The model uses this to self-select correctly.
- Never use
additionalProperties: truein your parameter schemas. Locking the schema closed forces the model to use exactly the properties you defined, which prevents hallucinated extra keys that would silently fail or cause downstream errors. - When a parameter is an ID that the model can never invent (like a database row ID), mark it required but also document where it comes from: 'The order_id returned by list_orders. Do not fabricate.' This cues the model to chain tool calls rather than guess.
- JSON Schema constraints like
minimum,maximum,minLength,maxLength, andpatternare enforced by the model probabilistically, not guaranteed. Always re-validate model-generated arguments server-side with a library likejsonschemabefore passing them to real functions.
Common pitfalls
- Mistake: Using a single generic
run_query(sql: string)tool to keep schemas simple. Fix: Model each distinct database action as a separate typed function to prevent arbitrary writes and hallucinated SQL syntax. - Mistake: Omitting
descriptionfrom individual parameter fields and relying on the parameter name alone. Fix: Write a one-sentence description for every parameter, especially ones with non-obvious valid values or format constraints. - Mistake: Marking every parameter as required, so the model fails to call the tool when optional context is absent. Fix: Only mark parameters required if the function cannot execute without them; provide sensible defaults for everything else.
- Mistake: Adding 20+ tools to every request regardless of relevance. Fix: Build a tool routing layer that selects 3-8 relevant tools per request using semantic similarity to the user message.
How to model different tool types as schemas
| Option | Use when | Avoid when |
|---|---|---|
| REST API tool | Calling external services (weather, payments, CRM). Map query params or request body fields directly to schema properties. | The API has dynamic fields that change per endpoint; model a per-endpoint schema instead of one generic HTTP tool. |
| Database query tool (typed) | Your agent needs to read structured data. Expose filters and column selectors as typed parameters, keep SQL server-side. | You need the model to construct arbitrary cross-table joins; that complexity belongs in application logic, not tool arguments. |
| Calculator / math tool | Precise arithmetic is needed and you cannot trust the model's floating-point reasoning. A simple evaluate(expression: string) wrapping a safe eval is enough. | The calculation is simple enough that the model handles it reliably; unnecessary tool calls add latency. |
| Search engine tool | User questions require information beyond the model's training cutoff or internal knowledge base. | Your retrieval needs are over private documents; use a RAG tool with vector search instead of a public web search API. |
| Composite / multi-step tool | A repeated sequence of low-level calls can be encapsulated as one higher-level tool to reduce round-trips. | The steps have branching logic that depends on intermediate results; keep them separate so the model can adapt between calls. |
Code Example
# openai>=1.0.0
import openai
client = openai.OpenAI() # reads OPENAI_API_KEY from env
weather_tool = {
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Return current temperature and conditions for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. 'Austin, TX'"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit. Defaults to celsius."
}
},
"required": ["city"]
}
}
}
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
tools=[weather_tool],
)
print(response.choices[0].message.tool_calls)How this code works
This code demonstrates how to teach an OpenAI language model about an external function, enabling it to intelligently decide when and how to "call" that function based on a user's request. It specifically shows how to define a tool for fetching current weather information. The process begins by initializing the openai client to connect with the AI service. The core of the example is the weather_tool dictionary, which outlines the function called get_current_weather. This definition includes a clear description that tells the AI what the tool does ("Return current temperature and conditions for a city.").
Inside weather_tool, the parameters section uses a JSON Schema-like structure to specify the inputs the get_current_weather function expects. It defines city as a string (e.g., 'Austin, TX') and unit as an enum allowing "celsius" or "fahrenheit," making city a required argument. When client.chat.completions.create is called, the model="gpt-4o" is given the user's messages and, crucially, the tools=[weather_tool]. The subtle point here is that the AI doesn't execute the get_current_weather function itself. Instead, the response.choices[0].message.tool_calls output contains a structured representation of the proposed function call, including the function's name and its arguments, which an external application would then use to perform the actual weather lookup.
Production-grade example
Adds schema validation, retry with backoff, timeout, token logging, and structured error paths.
# openai>=1.0.0, tenacity>=8.0.0
import os
import json
import time
import logging
import jsonschema
from openai import OpenAI, APITimeoutError, RateLimitError, APIError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
logger = logging.getLogger(__name__)
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
SEARCH_TOOL = {
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web for current information. Use when the user asks about recent events, prices, or facts that may have changed after your training cutoff.",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "Search query, written as a concise phrase."},
"num_results": {"type": "integer", "minimum": 1, "maximum": 10, "description": "Number of results to return. Default 3."},
},
"required": ["query"],
},
},
}
ARGUMENT_SCHEMA = {
"type": "object",
"properties": {
"query": {"type": "string"},
"num_results": {"type": "integer", "minimum": 1, "maximum": 10},
},
"required": ["query"],
}
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
)
def call_with_tools(user_message: str) -> dict:
start = time.monotonic()
try:
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": user_message}],
tools=[SEARCH_TOOL],
tool_choice="auto",
)
except APIError as exc:
logger.error("openai_api_error", extra={"status": exc.status_code, "msg": str(exc)})
raise
latency_ms = (time.monotonic() - start) * 1000
usage = response.usage
logger.info(
"llm_call_complete",
extra={
"latency_ms": round(latency_ms, 1),
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
"model": response.model,
},
)
choice = response.choices[0].message
if not choice.tool_calls:
return {"type": "text", "content": choice.content}
tool_call = choice.tool_calls[0]
try:
args = json.loads(tool_call.function.arguments)
jsonschema.validate(instance=args, schema=ARGUMENT_SCHEMA)
except (json.JSONDecodeError, jsonschema.ValidationError) as exc:
logger.warning("invalid_tool_arguments", extra={"raw": tool_call.function.arguments, "error": str(exc)})
return {"type": "error", "content": "Model produced invalid tool arguments.", "raw": tool_call.function.arguments}
return {"type": "tool_call", "name": tool_call.function.name, "args": args, "call_id": tool_call.id}How this code works
This code enables an AI assistant to intelligently decide when to perform a web search and then correctly format that search. It handles the definition of a web_search tool, how the AI uses it, and validates the AI's output to ensure reliability.
The SEARCH_TOOL dictionary describes the web_search tool to the AI, specifying its name, a helpful description for when to use it, and the parameters it expects, like query and an optional num_results. The call_with_tools function orchestrates the interaction. It uses @retry to automatically reattempt API calls if temporary issues like RateLimitError occur, making the system more robust. When client.chat.completions.create is called, tool_choice="auto" allows the AI to decide whether to use the defined SEARCH_TOOL or respond directly with text. If the AI suggests a tool call, the code parses the AI's generated arguments using json.loads and then critically validates them against ARGUMENT_SCHEMA using jsonschema.validate. This explicit validation is a subtle but important production-grade step: even though the AI received the SEARCH_TOOL schema as a guide, this second ARGUMENT_SCHEMA acts as a safety net, catching any malformed or unexpected arguments the AI might "hallucinate" before they cause errors in the actual search.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Define three tool schemas for a customer support agent: one to look up an order by ID, one to list a customer's recent orders by email, and one to search a help center by keyword. Then make a single chat completions request and print which tool the model selects for the message: 'I can't find my order from last Tuesday, my email is [email protected]'.
# openai>=1.0.0
import os, json
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
# TODO: Define lookup_order_by_id tool schema
lookup_order_tool = {
"type": "function",
"function": {
"name": "lookup_order_by_id",
"description": "TODO",
"parameters": {
"type": "object",
"properties": {}, # TODO: add order_id parameter
"required": [],
},
},
}
# TODO: Define list_orders_by_email tool schema
list_orders_tool = {}
# TODO: Define search_help_center tool schema
search_help_tool = {}
tools = [lookup_order_tool, list_orders_tool, search_help_tool]
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "I can't find my order from last Tuesday, my email is [email protected]"}],
tools=tools,
)
# TODO: Print the name of the tool the model chose and the arguments it provided
tool_call = response.choices[0].message.tool_calls
print(tool_call)Quick check
A model repeatedly calls a tool with an extra invented argument that doesn't exist in your schema. What is the most likely root cause?
You have 25 tools in your registry. What is the recommended production approach for passing tools to the model?
Why should you validate model-generated tool arguments with jsonschema before executing the function, even though you already defined the schema for the model?
create_support_ticket tool that takes a required subject string, a required priority enum of low/medium/high, and an optional customer_id string. Explain why you would or wouldn't mark customer_id as required, and what you'd put in the description field to help the model decide when to call this tool versus a search_existing_tickets tool.