Phase 4: AI Agents & Autonomous Systems

Budget limits, rate limits & sandboxing

Advanced ~17 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you have a super smart robot chef who can make any meal you ask for! It's really good at following recipes, but because it’s so powerful, it needs some special rules to make sure it doesn't accidentally cause a big mess or a big problem.

First, there are "budget limits." Think about ingredients. Your robot chef might decide it needs a hundred pounds of fancy cheese for one sandwich, or wants to keep ordering more and more expensive ingredients, even when it has enough. A budget limit is like telling the robot: "You only have $20 for this meal," or "You can only use two blocks of cheese." This stops it from spending all your money or using up all your supplies by accident. It’s a simple rule to keep things fair and make sure you don’t run out of everything.

Then, there are "rate limits." Let's say your robot chef needs to ask the pantry for an egg. That’s fine. But what if it asks for an egg, then immediately asks for another egg, then another, a thousand times a second? The poor pantry assistant would get totally overwhelmed! A rate limit is like telling the robot: "You can only ask the pantry for an ingredient once every few seconds," or "You can only send out five finished dishes every minute." This makes sure the robot doesn't overwhelm other parts of the kitchen or create a huge pile-up.

Finally, "sandboxing" is like giving your robot chef its own special, separate mini-kitchen. In this mini-kitchen, it can chop, mix, and experiment all it wants, but it can’t accidentally touch the real oven's power switch, or unlock the main restaurant door, or try to re-wire the lights. It’s a safe playpen where the robot can do its work without any chance of accidentally messing with the important, real kitchen systems.

These aren't just polite suggestions; they are really important rules that you set up before your robot even starts cooking its first meal. They make sure your super smart chef is helpful and safe, not a chaotic disaster. So when you build your own clever computer programs that can do lots of things, you'll use these ideas to make sure they always run smoothly and safely.

How these guardrails actually work

Budget limits track cumulative resource consumption across the agent's lifetime -- or across a single run -- and halt execution when a threshold is crossed. The most useful form is a dollar-equivalent budget computed from token usage, because raw token counts don't translate consistently across models or modalities. A GPT-4o call costs roughly 15x more per token than a gpt-4o-mini call (pricing changes frequently -- treat any number as illustrative). Your budget tracker needs to know which model was called for each turn and apply the right cost rate. Store this as a running total in a lightweight object that gets passed through every tool-call loop iteration. When the total crosses your threshold, the orchestrator returns a sentinel response and stops calling tools. Critically, this budget check happens before each LLM call, not after, so you don't overshoot by one expensive turn.

Rate limits are about time-windowed frequency, not total cost. You're solving two different problems simultaneously: protecting downstream services your agent calls (your own APIs, third-party APIs, databases) and respecting provider-side rate limits. These are separate controls. Provider rate limits (requests-per-minute, tokens-per-minute) are enforced externally and show up as 429 errors; your retry logic handles those. Your own rate limits are enforced internally and prevent the agent from making 50 database writes per second during a loop that's gone off the rails. Use a token bucket or a sliding window counter. The standard Python approach is a semaphore combined with asyncio.sleep, or a library like slowapi or limits for more sophisticated windowing. Set limits conservatively at first -- a well-behaved agent shouldn't need to call any single tool more than a handful of times per second.

A real production scenario

You've built a research agent that uses three tools: a web search tool (calls Tavily or Brave Search), a code execution tool (runs Python snippets), and a file-write tool (writes output to a report directory). Without controls, a prompt like "Research every company in the S&P 500 and write individual reports" could trigger 500 sequential search calls and 500 file writes before you notice. With budget limits, the agent halts after -- say -- $2.00 of total API spend. With rate limits, the search tool is capped at 5 calls per second so you don't burn through your Tavily daily quota in 90 seconds. And the code execution tool runs inside a Docker container so that a synthesized code snippet can't import subprocess and exfiltrate your API keys.

A senior engineer sets these limits by working backward from acceptable blast radius, not forward from "what the agent probably needs." If a single task should cost at most $0.50 and take at most 2 minutes, those become hard limits with no override path. If someone needs more, they file a request and you review it manually.

Tradeoffs vs. alternative approaches

You could handle budget enforcement at the API gateway layer (AWS API Gateway, Kong, or a custom proxy) rather than in application code. This is more operationally robust since it survives application bugs, but it's coarser -- the gateway doesn't know which agent run generated which token. Application-level budget tracking gives you per-run granularity and lets you surface budget-exceeded errors as structured responses rather than hard 429s. In production, use both: gateway-level as a circuit breaker, application-level for accounting. For sandboxing, the spectrum runs from process-level isolation (subprocess with limited permissions, cheap but weak) through container isolation (Docker with seccomp profiles and no-network flags, good default) to VM-level isolation (Firecracker or gVisor, strong but adds 100ms+ cold start). Most agents should default to Docker with dropped capabilities. Only go heavier if your threat model includes adversarial prompt injection that specifically targets sandbox escape.

What changes at scale

At 10 users, an in-memory token counter per request is fine. At 10,000 users, you need shared state -- Redis with atomic increment operations is the standard choice. Each agent turn does a INCRBY on a per-run key with a TTL equal to your run timeout. If the key doesn't exist yet, it gets created; if it exceeds budget, the agent halts before issuing the API call. This adds roughly 1ms of latency per turn, which is negligible compared to LLM round-trip times. At 10 million users, you're now thinking about multi-region Redis, per-tenant budget pools, and real-time spend dashboards with alerting. You also need to handle the race condition where two concurrent agent turns both check the budget, both see "under limit," and both proceed, pushing you over. Use Redis transactions (MULTI/EXEC) or Lua scripts for atomic check-and-increment. Sandboxing at scale means container orchestration: Kubernetes with namespace isolation, network policies that block inter-agent communication, and image scanning in CI so a compromised dependency doesn't become a vector.

Key Takeaways

  • Implement budget limits in the orchestration layer, never in the system prompt.
  • Use token-level and USD-level counters together; token counts alone miss embedding and image costs.
  • Sandbox code execution with Docker or gVisor; chroot alone is not sufficient for adversarial input.
  • Rate limit at the agent level independently of provider-side limits -- they serve different threat models.

Pro tips

  • Track budget in USD-equivalent, not tokens. Token counts are model-specific and don't account for image inputs, embedding calls, or fine-tuned model surcharges that might run in the same agent loop.
  • Your internal rate limit and the provider's rate limit solve different problems. The provider's limit is about their infrastructure; yours is about detecting runaway loops. An agent making 10 calls/second to a tool that should need 1 call/turn is a bug signal, not just a cost concern.
  • When sandboxing code execution with Docker, use --network none and --read-only flags with an explicit tmpfs mount for /tmp. Most legitimate agent-generated code doesn't need network access or persistent writes, and this eliminates an entire class of exfiltration attacks.
  • Set your budget limit to trigger a structured halting response rather than raising an exception that kills the process. The calling code can then log the partial result, notify the user, and allow a resume-from-checkpoint if you've stored intermediate state.

Common pitfalls

  • Mistake: Enforcing budget limits inside the system prompt ('don't spend more than $1'). Fix: The model cannot count its own API costs. Enforce limits in the orchestration layer with actual token/cost counters.
  • Mistake: Relying solely on provider-side rate limits to prevent runaway loops. Fix: Add your own per-run call counter; provider limits are per API key and apply across all your users, not per agent run.
  • Mistake: Using subprocess.run with shell=True for code execution sandboxing. Fix: Use a container with dropped Linux capabilities and no network. Shell=True with user-influenced input is a direct command injection vector.
  • Mistake: Setting a single global token budget shared across all concurrent agent runs. Fix: Use per-run budget keys (e.g., namespaced by run_id in Redis). A single expensive run should not starve all other users.

Sandboxing approach: when to use which

Option Use when Avoid when
No sandbox (direct subprocess) Agent only calls pre-defined, whitelisted functions you wrote; no user-influenced code paths. Agent can execute model-generated code or process untrusted external content.
Docker container (--network none, dropped caps) Agent executes model-generated Python or shell snippets in a controlled environment. You need sub-50ms cold starts or cannot run Docker (some serverless platforms).
gVisor (runsc runtime) Docker is insufficient because your threat model includes kernel exploit attempts from adversarial payloads. Performance overhead of gVisor's syscall interception is unacceptable for your latency SLA.
Firecracker microVM You need full VM-level isolation (e.g., multi-tenant code execution as a service) and can absorb 100-300ms cold starts. You're sandboxing occasional tool calls inside a single-tenant agent; operational overhead is not justified.
AWS Lambda with VPC isolation You're already on AWS, want managed infrastructure, and can tolerate Lambda cold starts and execution time limits. Your sandboxed code needs more than 15 minutes of execution time or more than 10GB of memory.

Code Example

python
# openai>=1.0.0, requires OPENAI_API_KEY env var
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

MAX_TOKENS_PER_RUN = 20_000  # roughly $0.10 at gpt-4o-mini input rates (illustrative)

def run_agent_turn(messages: list[dict], token_budget: dict) -> str:
    """Single agent turn with a rolling token budget."""
    if token_budget["used"] >= MAX_TOKENS_PER_RUN:
        return "[BUDGET_EXCEEDED] Agent halted: token budget exhausted."

    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=messages,
        max_tokens=min(1024, MAX_TOKENS_PER_RUN - token_budget["used"]),
    )
    usage = response.usage
    token_budget["used"] += usage.total_tokens
    return response.choices[0].message.content

# Usage
budget = {"used": 0}
for turn in range(10):
    reply = run_agent_turn([{"role": "user", "content": "Next step?"}], budget)
    print(f"Turn {turn}: {reply[:60]} | tokens used: {budget['used']}")
    if "BUDGET_EXCEEDED" in reply:
        break

How this code works

This code demonstrates a fundamental safety guardrail for AI agents: implementing a token budget to control costs and prevent runaway agent behavior. It defines a maximum number of language model tokens an agent can use across multiple interaction turns, halting its operation once that limit is reached. This mechanism is vital for ensuring AI applications stay within predefined operational boundaries and cost constraints, especially in sandboxed or production environments.

The process begins by setting up an OpenAI client and defining MAX_TOKENS_PER_RUN, representing the total budget. The core logic resides in the run_agent_turn function. Before making an LLM call, it first checks if the token_budget["used"] has already exceeded MAX_TOKENS_PER_RUN. If so, it immediately returns a [BUDGET_EXCEEDED] message, stopping the agent. Crucially, the max_tokens argument in the client.chat.completions.create call is dynamically set using min(1024, MAX_TOKENS_PER_RUN - token_budget["used"]). This ensures the agent never requests a response longer than 1024 tokens or longer than its remaining budget, whichever is smaller. After each successful call, the usage.total_tokens (input + output) are added to the token_budget["used"], providing a rolling total to enforce the budget across turns.

Production-grade example

Redis atomic budgeting, sliding-window RPM, exponential backoff, structured JSON logging, graceful halting.

python
# openai>=1.0.0, redis>=5.0.0, tenacity>=8.0.0
import os
import time
import logging
import json
from openai import OpenAI, RateLimitError, APITimeoutError, APIError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
import redis

logger = logging.getLogger("agent.budget")
logging.basicConfig(level=logging.INFO, format='{"time":"%(asctime)s","level":"%(levelname)s","msg":%(message)s}')

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
redis_client = redis.Redis.from_url(os.environ["REDIS_URL"], decode_responses=True)

MODEL = "gpt-4o-mini"
BUDGET_USD = float(os.environ.get("AGENT_BUDGET_USD", "1.00"))
RATE_LIMIT_CALLS_PER_MIN = int(os.environ.get("AGENT_RPM", "20"))
COST_PER_1K_INPUT = 0.00015   # illustrative, verify at platform.openai.com/pricing
COST_PER_1K_OUTPUT = 0.00060  # illustrative

class BudgetExceededError(Exception):
    pass

class RateLimitLocalError(Exception):
    pass

def check_and_update_budget(run_id: str, input_tokens: int, output_tokens: int) -> float:
    """Atomically update spend; raise BudgetExceededError if over limit."""
    cost = (input_tokens / 1000) * COST_PER_1K_INPUT + (output_tokens / 1000) * COST_PER_1K_OUTPUT
    key = f"agent:budget:{run_id}"
    # Lua script ensures atomic read-check-write
    lua = """
    local current = tonumber(redis.call('get', KEYS[1]) or 0)
    local new = current + tonumber(ARGV[1])
    if new > tonumber(ARGV[2]) then return -1 end
    redis.call('set', KEYS[1], new, 'EX', 3600)
    return new * 1000000
    """
    result = redis_client.eval(lua, 1, key, cost, BUDGET_USD)
    if result == -1:
        raise BudgetExceededError(f"Run {run_id} exceeded ${BUDGET_USD:.2f} budget")
    total_spent = result / 1_000_000
    logger.info(json.dumps({"run_id": run_id, "turn_cost_usd": round(cost, 6), "total_usd": round(total_spent, 4)}))
    return total_spent

def check_rate_limit(run_id: str) -> None:
    """Sliding window rate limit: max N calls per 60s per run."""
    key = f"agent:rpm:{run_id}"
    now = time.time()
    pipe = redis_client.pipeline()
    pipe.zremrangebyscore(key, 0, now - 60)
    pipe.zadd(key, {str(now): now})
    pipe.zcard(key)
    pipe.expire(key, 120)
    _, _, count, _ = pipe.execute()
    if count > RATE_LIMIT_CALLS_PER_MIN:
        raise RateLimitLocalError(f"Run {run_id} exceeded {RATE_LIMIT_CALLS_PER_MIN} RPM")

@retry(
    retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
    stop=stop_after_attempt(4),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    reraise=True,
)
def call_llm(run_id: str, messages: list[dict]) -> str:
    check_rate_limit(run_id)
    response = client.chat.completions.create(
        model=MODEL,
        messages=messages,
        max_tokens=1024,
    )
    usage = response.usage
    check_and_update_budget(run_id, usage.prompt_tokens, usage.completion_tokens)
    return response.choices[0].message.content

def run_agent(run_id: str, initial_messages: list[dict], max_turns: int = 10) -> str:
    messages = list(initial_messages)
    for turn in range(max_turns):
        try:
            reply = call_llm(run_id, messages)
            logger.info(json.dumps({"run_id": run_id, "turn": turn, "reply_preview": reply[:80]}))
            messages.append({"role": "assistant", "content": reply})
            if "[DONE]" in reply:
                return reply
        except BudgetExceededError as e:
            logger.warning(json.dumps({"run_id": run_id, "event": "budget_exceeded", "detail": str(e)}))
            return f"[HALTED] {e}"
        except RateLimitLocalError as e:
            logger.warning(json.dumps({"run_id": run_id, "event": "rate_limited", "detail": str(e)}))
            time.sleep(5)  # graceful degradation: wait and let caller retry
            return f"[THROTTLED] {e}"
        except APIError as e:
            logger.error(json.dumps({"run_id": run_id, "event": "api_error", "status": e.status_code}))
            raise
    return "[MAX_TURNS] Agent reached turn limit without completing."

How this code works

This code implements crucial safety guardrails for an AI agent, ensuring it operates within predefined financial and operational limits when interacting with an LLM. It actively prevents the agent from exceeding its allocated BUDGET_USD and from making too many call_llm requests per minute, which could trigger API rate limits or incur unnecessary costs. The run_agent function manages the agent's conversational turns, gracefully halting or throttling its operation if these limits are encountered.

Core to these guardrails are the check_and_update_budget and check_rate_limit functions. check_and_update_budget employs a Redis Lua script to atomically calculate and update the total cost for a specific run_id. This atomicity is a subtle but vital point: it prevents race conditions where multiple concurrent updates could lead to incorrect budget tracking. If an update would push the cost over BUDGET_USD, it raises a BudgetExceededError. Similarly, check_rate_limit utilizes Redis sorted sets to enforce a sliding window rate limit for RATE_LIMIT_CALLS_PER_MIN, raising RateLimitLocalError if the agent exceeds its allowed calls. Additionally, the call_llm function uses tenacity.retry to automatically handle transient OpenAI RateLimitError or APITimeoutError with exponential backoff for increased robustness.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a simple agent orchestrator with two guardrails: a per-run token budget of 5,000 tokens and a rate limit of 3 LLM calls per minute. The agent should stop gracefully when either limit is hit and print a structured summary showing how many tokens were used and how many turns completed.

python
import os
import time
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

TOKEN_BUDGET = 5_000
MAX_CALLS_PER_MINUTE = 3

def run_guarded_agent(task: str, max_turns: int = 10):
    state = {
        "tokens_used": 0,
        "calls_this_minute": 0,
        "minute_start": time.time(),
        "turns_completed": 0,
    }
    messages = [{"role": "user", "content": task}]

    for turn in range(max_turns):
        # TODO: Check if token budget is exceeded; return summary if so

        # TODO: Check and enforce rate limit (reset counter if 60s has elapsed)

        # TODO: Call the LLM, update tokens_used and calls_this_minute

        # TODO: Append assistant reply to messages
        state["turns_completed"] += 1

        # Simulate task completion check
        if "done" in messages[-1]["content"].lower():
            break

    return {"status": "completed", **state}

result = run_guarded_agent("Count from 1 to 20, one number per message, then say done.")
print(result)

Quick check

  1. An agent loop calls the LLM 200 times in 10 seconds, hitting your provider's rate limit. What does your internal per-run rate limit add that the provider's limit doesn't?

  2. Why should budget limits be enforced in the orchestration layer rather than via a system prompt instruction?

  3. You need to sandbox an agent that executes model-generated Python code. Which approach is the weakest security boundary?

Self-check: Explain why you would need both an application-level token budget counter AND a Redis-backed rate limiter in a multi-user agent deployment, and describe the specific failure mode that each one prevents but the other does not.