How these guardrails actually work
Budget limits track cumulative resource consumption across the agent's lifetime -- or across a single run -- and halt execution when a threshold is crossed. The most useful form is a dollar-equivalent budget computed from token usage, because raw token counts don't translate consistently across models or modalities. A GPT-4o call costs roughly 15x more per token than a gpt-4o-mini call (pricing changes frequently -- treat any number as illustrative). Your budget tracker needs to know which model was called for each turn and apply the right cost rate. Store this as a running total in a lightweight object that gets passed through every tool-call loop iteration. When the total crosses your threshold, the orchestrator returns a sentinel response and stops calling tools. Critically, this budget check happens before each LLM call, not after, so you don't overshoot by one expensive turn.
Rate limits are about time-windowed frequency, not total cost. You're solving two different problems simultaneously: protecting downstream services your agent calls (your own APIs, third-party APIs, databases) and respecting provider-side rate limits. These are separate controls. Provider rate limits (requests-per-minute, tokens-per-minute) are enforced externally and show up as 429 errors; your retry logic handles those. Your own rate limits are enforced internally and prevent the agent from making 50 database writes per second during a loop that's gone off the rails. Use a token bucket or a sliding window counter. The standard Python approach is a semaphore combined with asyncio.sleep, or a library like slowapi or limits for more sophisticated windowing. Set limits conservatively at first -- a well-behaved agent shouldn't need to call any single tool more than a handful of times per second.
A real production scenario
You've built a research agent that uses three tools: a web search tool (calls Tavily or Brave Search), a code execution tool (runs Python snippets), and a file-write tool (writes output to a report directory). Without controls, a prompt like "Research every company in the S&P 500 and write individual reports" could trigger 500 sequential search calls and 500 file writes before you notice. With budget limits, the agent halts after -- say -- $2.00 of total API spend. With rate limits, the search tool is capped at 5 calls per second so you don't burn through your Tavily daily quota in 90 seconds. And the code execution tool runs inside a Docker container so that a synthesized code snippet can't import subprocess and exfiltrate your API keys.
A senior engineer sets these limits by working backward from acceptable blast radius, not forward from "what the agent probably needs." If a single task should cost at most $0.50 and take at most 2 minutes, those become hard limits with no override path. If someone needs more, they file a request and you review it manually.
Tradeoffs vs. alternative approaches
You could handle budget enforcement at the API gateway layer (AWS API Gateway, Kong, or a custom proxy) rather than in application code. This is more operationally robust since it survives application bugs, but it's coarser -- the gateway doesn't know which agent run generated which token. Application-level budget tracking gives you per-run granularity and lets you surface budget-exceeded errors as structured responses rather than hard 429s. In production, use both: gateway-level as a circuit breaker, application-level for accounting. For sandboxing, the spectrum runs from process-level isolation (subprocess with limited permissions, cheap but weak) through container isolation (Docker with seccomp profiles and no-network flags, good default) to VM-level isolation (Firecracker or gVisor, strong but adds 100ms+ cold start). Most agents should default to Docker with dropped capabilities. Only go heavier if your threat model includes adversarial prompt injection that specifically targets sandbox escape.
What changes at scale
At 10 users, an in-memory token counter per request is fine. At 10,000 users, you need shared state -- Redis with atomic increment operations is the standard choice. Each agent turn does a INCRBY on a per-run key with a TTL equal to your run timeout. If the key doesn't exist yet, it gets created; if it exceeds budget, the agent halts before issuing the API call. This adds roughly 1ms of latency per turn, which is negligible compared to LLM round-trip times. At 10 million users, you're now thinking about multi-region Redis, per-tenant budget pools, and real-time spend dashboards with alerting. You also need to handle the race condition where two concurrent agent turns both check the budget, both see "under limit," and both proceed, pushing you over. Use Redis transactions (MULTI/EXEC) or Lua scripts for atomic check-and-increment. Sandboxing at scale means container orchestration: Kubernetes with namespace isolation, network policies that block inter-agent communication, and image scanning in CI so a compromised dependency doesn't become a vector.
Key Takeaways
- Implement budget limits in the orchestration layer, never in the system prompt.
- Use token-level and USD-level counters together; token counts alone miss embedding and image costs.
- Sandbox code execution with Docker or gVisor; chroot alone is not sufficient for adversarial input.
- Rate limit at the agent level independently of provider-side limits -- they serve different threat models.
Pro tips
- Track budget in USD-equivalent, not tokens. Token counts are model-specific and don't account for image inputs, embedding calls, or fine-tuned model surcharges that might run in the same agent loop.
- Your internal rate limit and the provider's rate limit solve different problems. The provider's limit is about their infrastructure; yours is about detecting runaway loops. An agent making 10 calls/second to a tool that should need 1 call/turn is a bug signal, not just a cost concern.
- When sandboxing code execution with Docker, use
--network noneand--read-onlyflags with an explicit tmpfs mount for /tmp. Most legitimate agent-generated code doesn't need network access or persistent writes, and this eliminates an entire class of exfiltration attacks. - Set your budget limit to trigger a structured halting response rather than raising an exception that kills the process. The calling code can then log the partial result, notify the user, and allow a resume-from-checkpoint if you've stored intermediate state.
Common pitfalls
- Mistake: Enforcing budget limits inside the system prompt ('don't spend more than $1'). Fix: The model cannot count its own API costs. Enforce limits in the orchestration layer with actual token/cost counters.
- Mistake: Relying solely on provider-side rate limits to prevent runaway loops. Fix: Add your own per-run call counter; provider limits are per API key and apply across all your users, not per agent run.
- Mistake: Using
subprocess.runwith shell=True for code execution sandboxing. Fix: Use a container with dropped Linux capabilities and no network. Shell=True with user-influenced input is a direct command injection vector. - Mistake: Setting a single global token budget shared across all concurrent agent runs. Fix: Use per-run budget keys (e.g., namespaced by run_id in Redis). A single expensive run should not starve all other users.
Sandboxing approach: when to use which
| Option | Use when | Avoid when |
|---|---|---|
| No sandbox (direct subprocess) | Agent only calls pre-defined, whitelisted functions you wrote; no user-influenced code paths. | Agent can execute model-generated code or process untrusted external content. |
| Docker container (--network none, dropped caps) | Agent executes model-generated Python or shell snippets in a controlled environment. | You need sub-50ms cold starts or cannot run Docker (some serverless platforms). |
| gVisor (runsc runtime) | Docker is insufficient because your threat model includes kernel exploit attempts from adversarial payloads. | Performance overhead of gVisor's syscall interception is unacceptable for your latency SLA. |
| Firecracker microVM | You need full VM-level isolation (e.g., multi-tenant code execution as a service) and can absorb 100-300ms cold starts. | You're sandboxing occasional tool calls inside a single-tenant agent; operational overhead is not justified. |
| AWS Lambda with VPC isolation | You're already on AWS, want managed infrastructure, and can tolerate Lambda cold starts and execution time limits. | Your sandboxed code needs more than 15 minutes of execution time or more than 10GB of memory. |
Code Example
# openai>=1.0.0, requires OPENAI_API_KEY env var
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
MAX_TOKENS_PER_RUN = 20_000 # roughly $0.10 at gpt-4o-mini input rates (illustrative)
def run_agent_turn(messages: list[dict], token_budget: dict) -> str:
"""Single agent turn with a rolling token budget."""
if token_budget["used"] >= MAX_TOKENS_PER_RUN:
return "[BUDGET_EXCEEDED] Agent halted: token budget exhausted."
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
max_tokens=min(1024, MAX_TOKENS_PER_RUN - token_budget["used"]),
)
usage = response.usage
token_budget["used"] += usage.total_tokens
return response.choices[0].message.content
# Usage
budget = {"used": 0}
for turn in range(10):
reply = run_agent_turn([{"role": "user", "content": "Next step?"}], budget)
print(f"Turn {turn}: {reply[:60]} | tokens used: {budget['used']}")
if "BUDGET_EXCEEDED" in reply:
breakHow this code works
This code demonstrates a fundamental safety guardrail for AI agents: implementing a token budget to control costs and prevent runaway agent behavior. It defines a maximum number of language model tokens an agent can use across multiple interaction turns, halting its operation once that limit is reached. This mechanism is vital for ensuring AI applications stay within predefined operational boundaries and cost constraints, especially in sandboxed or production environments.
The process begins by setting up an OpenAI client and defining MAX_TOKENS_PER_RUN, representing the total budget. The core logic resides in the run_agent_turn function. Before making an LLM call, it first checks if the token_budget["used"] has already exceeded MAX_TOKENS_PER_RUN. If so, it immediately returns a [BUDGET_EXCEEDED] message, stopping the agent. Crucially, the max_tokens argument in the client.chat.completions.create call is dynamically set using min(1024, MAX_TOKENS_PER_RUN - token_budget["used"]). This ensures the agent never requests a response longer than 1024 tokens or longer than its remaining budget, whichever is smaller. After each successful call, the usage.total_tokens (input + output) are added to the token_budget["used"], providing a rolling total to enforce the budget across turns.
Production-grade example
Redis atomic budgeting, sliding-window RPM, exponential backoff, structured JSON logging, graceful halting.
# openai>=1.0.0, redis>=5.0.0, tenacity>=8.0.0
import os
import time
import logging
import json
from openai import OpenAI, RateLimitError, APITimeoutError, APIError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
import redis
logger = logging.getLogger("agent.budget")
logging.basicConfig(level=logging.INFO, format='{"time":"%(asctime)s","level":"%(levelname)s","msg":%(message)s}')
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
redis_client = redis.Redis.from_url(os.environ["REDIS_URL"], decode_responses=True)
MODEL = "gpt-4o-mini"
BUDGET_USD = float(os.environ.get("AGENT_BUDGET_USD", "1.00"))
RATE_LIMIT_CALLS_PER_MIN = int(os.environ.get("AGENT_RPM", "20"))
COST_PER_1K_INPUT = 0.00015 # illustrative, verify at platform.openai.com/pricing
COST_PER_1K_OUTPUT = 0.00060 # illustrative
class BudgetExceededError(Exception):
pass
class RateLimitLocalError(Exception):
pass
def check_and_update_budget(run_id: str, input_tokens: int, output_tokens: int) -> float:
"""Atomically update spend; raise BudgetExceededError if over limit."""
cost = (input_tokens / 1000) * COST_PER_1K_INPUT + (output_tokens / 1000) * COST_PER_1K_OUTPUT
key = f"agent:budget:{run_id}"
# Lua script ensures atomic read-check-write
lua = """
local current = tonumber(redis.call('get', KEYS[1]) or 0)
local new = current + tonumber(ARGV[1])
if new > tonumber(ARGV[2]) then return -1 end
redis.call('set', KEYS[1], new, 'EX', 3600)
return new * 1000000
"""
result = redis_client.eval(lua, 1, key, cost, BUDGET_USD)
if result == -1:
raise BudgetExceededError(f"Run {run_id} exceeded ${BUDGET_USD:.2f} budget")
total_spent = result / 1_000_000
logger.info(json.dumps({"run_id": run_id, "turn_cost_usd": round(cost, 6), "total_usd": round(total_spent, 4)}))
return total_spent
def check_rate_limit(run_id: str) -> None:
"""Sliding window rate limit: max N calls per 60s per run."""
key = f"agent:rpm:{run_id}"
now = time.time()
pipe = redis_client.pipeline()
pipe.zremrangebyscore(key, 0, now - 60)
pipe.zadd(key, {str(now): now})
pipe.zcard(key)
pipe.expire(key, 120)
_, _, count, _ = pipe.execute()
if count > RATE_LIMIT_CALLS_PER_MIN:
raise RateLimitLocalError(f"Run {run_id} exceeded {RATE_LIMIT_CALLS_PER_MIN} RPM")
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
stop=stop_after_attempt(4),
wait=wait_exponential(multiplier=1, min=2, max=30),
reraise=True,
)
def call_llm(run_id: str, messages: list[dict]) -> str:
check_rate_limit(run_id)
response = client.chat.completions.create(
model=MODEL,
messages=messages,
max_tokens=1024,
)
usage = response.usage
check_and_update_budget(run_id, usage.prompt_tokens, usage.completion_tokens)
return response.choices[0].message.content
def run_agent(run_id: str, initial_messages: list[dict], max_turns: int = 10) -> str:
messages = list(initial_messages)
for turn in range(max_turns):
try:
reply = call_llm(run_id, messages)
logger.info(json.dumps({"run_id": run_id, "turn": turn, "reply_preview": reply[:80]}))
messages.append({"role": "assistant", "content": reply})
if "[DONE]" in reply:
return reply
except BudgetExceededError as e:
logger.warning(json.dumps({"run_id": run_id, "event": "budget_exceeded", "detail": str(e)}))
return f"[HALTED] {e}"
except RateLimitLocalError as e:
logger.warning(json.dumps({"run_id": run_id, "event": "rate_limited", "detail": str(e)}))
time.sleep(5) # graceful degradation: wait and let caller retry
return f"[THROTTLED] {e}"
except APIError as e:
logger.error(json.dumps({"run_id": run_id, "event": "api_error", "status": e.status_code}))
raise
return "[MAX_TURNS] Agent reached turn limit without completing."
How this code works
This code implements crucial safety guardrails for an AI agent, ensuring it operates within predefined financial and operational limits when interacting with an LLM. It actively prevents the agent from exceeding its allocated BUDGET_USD and from making too many call_llm requests per minute, which could trigger API rate limits or incur unnecessary costs. The run_agent function manages the agent's conversational turns, gracefully halting or throttling its operation if these limits are encountered.
Core to these guardrails are the check_and_update_budget and check_rate_limit functions. check_and_update_budget employs a Redis Lua script to atomically calculate and update the total cost for a specific run_id. This atomicity is a subtle but vital point: it prevents race conditions where multiple concurrent updates could lead to incorrect budget tracking. If an update would push the cost over BUDGET_USD, it raises a BudgetExceededError. Similarly, check_rate_limit utilizes Redis sorted sets to enforce a sliding window rate limit for RATE_LIMIT_CALLS_PER_MIN, raising RateLimitLocalError if the agent exceeds its allowed calls. Additionally, the call_llm function uses tenacity.retry to automatically handle transient OpenAI RateLimitError or APITimeoutError with exponential backoff for increased robustness.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a simple agent orchestrator with two guardrails: a per-run token budget of 5,000 tokens and a rate limit of 3 LLM calls per minute. The agent should stop gracefully when either limit is hit and print a structured summary showing how many tokens were used and how many turns completed.
import os
import time
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
TOKEN_BUDGET = 5_000
MAX_CALLS_PER_MINUTE = 3
def run_guarded_agent(task: str, max_turns: int = 10):
state = {
"tokens_used": 0,
"calls_this_minute": 0,
"minute_start": time.time(),
"turns_completed": 0,
}
messages = [{"role": "user", "content": task}]
for turn in range(max_turns):
# TODO: Check if token budget is exceeded; return summary if so
# TODO: Check and enforce rate limit (reset counter if 60s has elapsed)
# TODO: Call the LLM, update tokens_used and calls_this_minute
# TODO: Append assistant reply to messages
state["turns_completed"] += 1
# Simulate task completion check
if "done" in messages[-1]["content"].lower():
break
return {"status": "completed", **state}
result = run_guarded_agent("Count from 1 to 20, one number per message, then say done.")
print(result)
Quick check
An agent loop calls the LLM 200 times in 10 seconds, hitting your provider's rate limit. What does your internal per-run rate limit add that the provider's limit doesn't?
Why should budget limits be enforced in the orchestration layer rather than via a system prompt instruction?
You need to sandbox an agent that executes model-generated Python code. Which approach is the weakest security boundary?