Phase 5: Production & Deployment

Token budgets & cost tracking per user, feature & endpoint

Intermediate ~14 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you have your very own amazing bakery, and you make delicious cookies, cakes, and pies. Every single treat you bake uses ingredients: flour, sugar, chocolate, eggs. These ingredients cost money, right? In the world of coding, when a powerful computer program (like an AI) does something for you, it uses tiny bits of information called "tokens." Think of tokens like your bakery ingredients – each one costs a little bit of money, and you don't want to run out or spend too much! If you don't keep track, you might suddenly find your flour cupboard empty and your piggy bank sad, wondering where all your money went.

To be a super-smart baker, you need to track your ingredients carefully. First, you'd want to know how much flour and sugar each customer uses. Maybe your friend Alex always orders giant chocolate cakes, and your friend Sam prefers small oatmeal cookies. You'd also want to know how much of each ingredient each recipe uses – is your new "super-fudgy brownie" recipe suddenly using twice as much expensive chocolate as you thought? And finally, if you have different workstations in your kitchen – like one oven for cookies and another for cakes – you’d want to know if one station is suddenly burning through ingredients faster because someone changed how they bake there.

This way of watching your ingredients is exactly what "token budgets and cost tracking" does for computer programs. It's like having a special clipboard where you write down, "Alex ordered a cake (recipe), made at the big oven (workstation), which used 50 tokens (ingredients)." By doing this, you can see exactly where all those tiny costs are going. You can even set a "budget," like saying, "I only have enough chocolate tokens for 10 cakes this week!"

So, when you build your own amazing computer programs, keeping track of these "ingredient" tokens means you'll never be surprised by a giant bill or an empty "pantry." You can quickly spot if a new feature you added is unexpectedly gobbling up tokens, or if one user is asking the AI to do so much work it's getting expensive. This helps you manage your virtual bakery efficiently and keep everyone happy, without breaking the bank.

At its core, token-based cost tracking is just request instrumentation with a pricing dimension. Every provider returns a usage object on each API response -- prompt_tokens, completion_tokens, and total_tokens. Those numbers, multiplied by the model's per-token rate, give you the dollar cost of that single call. The challenge isn't the math; it's attaching the right metadata so you can aggregate meaningfully later. A raw token count without knowing which user triggered it or which feature it served is nearly useless for optimization.

The mental model that works well in practice is treating each LLM call as a database query with an execution cost. You wouldn't ship a product without knowing which queries are slow. Same logic applies here. Instrument every call site to capture: timestamp, user_id (or a hashed tenant ID), feature slug (e.g., "doc_summarization", "chat_assistant"), endpoint label (e.g., "/api/v1/analyze"), model name, prompt tokens, completion tokens, and derived cost in USD. Persist that to a time-series-friendly store -- a Postgres table with an index on (user_id, created_at) works fine at moderate scale; ClickHouse or BigQuery at higher scale. Avoid logging raw prompt text by default: that's a data-privacy and storage cost problem.

Budget enforcement has two flavors: soft limits and hard limits. A soft limit warns the user or throttles their request rate when they approach their allocation. A hard limit rejects the request before it even reaches the LLM API. For a SaaS product with tiered plans, you might set a hard limit of 500k tokens per user per month on a free tier, with a soft warning at 80%. Implement hard limits as a pre-call guard: look up the user's token spend for the current window from your store, compare to their budget, and return an HTTP 429 with a clear message if exceeded. This prevents a single misbehaving user from generating a surprise invoice item. Implement the check efficiently -- a Redis counter with a 30-day TTL is faster than a Postgres aggregate for hot paths.

In the real world, this infrastructure pays off immediately in a multi-tenant SaaS setting. Imagine you've shipped a document summarization feature that calls GPT-4o. Three weeks after launch, your monthly OpenAI invoice jumps 3x. With no tracking, you're investigating cold. With tracking, you pull a GROUP BY feature query and see "doc_summarization" accounts for 71% of spend, and two enterprise trial accounts are summarizing hundreds of PDFs per day. You can immediately gate that feature behind a rate limit, contact those accounts about an enterprise plan, or switch the feature to a cheaper model (see the model-selection subtopic for that decision). None of those responses are available to you without the data.

At scale, the instrumentation approach changes shape. At 10 users, a SQLite table and a print statement is enough to learn the pattern. At 10k users, you want asynchronous logging so cost tracking doesn't add latency to the critical path -- fire-and-forget to a background queue (asyncio.create_task or a Celery worker). At 10M users, the ingestion volume of cost events is itself a real infrastructure problem. You batch-write to ClickHouse or BigQuery, run daily rollup aggregates, and query pre-aggregated tables rather than raw events. Budget checks at that scale hit Redis or Memcached for the current spend counter rather than querying raw event rows. The shape of the problem shifts from "how do I store this" to "how do I keep the check under 5ms." Start simple and let traffic drive the architecture.

Key Takeaways

  • Count tokens at the call site using provider SDKs or tiktoken; never estimate.
  • Tag every LLM call with user_id, feature, and endpoint before persisting cost data.
  • Enforce soft limits with warnings and hard limits with early exits before the API call.
  • Use time-windowed budgets (daily, monthly) so a burst doesn't silently exhaust a monthly quota.

Pro tips

  • Pre-count prompt tokens with tiktoken before the API call, not just after. This lets you enforce a per-request token cap and reject oversized inputs before you pay for them -- a single jailbreak attempt with a 100k-token injected document is otherwise a real line item.
  • Track completion tokens separately from prompt tokens. Completions are often 2-5x more expensive per token on some models, and a feature where users set max_tokens=4096 by default can silently dominate your bill even with modest prompt sizes.
  • Roll up cost data to a secondary table by (user_id, feature, date) at write time or via a nightly job. Querying raw event rows for a monthly user report at 10M events is slow; pre-aggregated rollups make dashboards instant and keep your analytics DB happy.
  • Redis counters for budget checks are fast but eventually consistent with your persistent store. Write the authoritative event log to Postgres or BigQuery asynchronously, and use Redis only as the hot-path gate. If Redis is unavailable, fail open with a log warning rather than blocking all LLM calls.

Common pitfalls

  • Mistake: Logging cost data synchronously on the critical path, adding 20-50ms to every response. Fix: Use asyncio.create_task or a background queue (Celery, ARQ) so the API response returns immediately.
  • Mistake: Using character count or word count as a proxy for token count. Fix: Always use tiktoken or the provider's token counting endpoint; character-to-token ratios vary by language and content type.
  • Mistake: Setting a single global monthly budget without user or feature dimension, so one heavy feature exhausts the budget for all users. Fix: Track budgets per user and per feature independently so enforcement is surgical.
  • Mistake: Forgetting to account for system prompt tokens in budget checks, which can be 300-800 tokens per call. Fix: Include the full message list -- system, few-shot examples, and conversation history -- in your pre-call token count.

Where to store token cost events

Option Use when Avoid when
Redis counters only You only need real-time budget enforcement and don't require historical cost queries. You need per-feature breakdowns, audit trails, or monthly invoice reconciliation.
Postgres event log Under ~50M events/month; team already operates Postgres; need SQL queries for ad-hoc analysis. Event volume exceeds what a single Postgres instance handles comfortably without heavy partitioning.
ClickHouse or BigQuery High event volume, complex aggregations, or BI dashboards querying raw events across millions of rows. You're early-stage; the operational overhead isn't justified until you have real scale.
Third-party (Helicone, LangSmith, OpenMeter) You want cost dashboards out of the box and don't want to build the pipeline from scratch. Your data cannot leave your infrastructure due to compliance requirements, or you need deep custom logic.

Code Example

python
# tiktoken 0.7.x, openai 1.x
import tiktoken
import openai

client = openai.OpenAI()  # reads OPENAI_API_KEY from env
enc = tiktoken.encoding_for_model("gpt-4o")

def count_tokens(messages: list[dict]) -> int:
    """Estimate prompt token count before sending."""
    total = 0
    for msg in messages:
        total += 4  # per-message overhead
        total += len(enc.encode(msg["content"]))
    return total + 2  # reply primer

messages = [{"role": "user", "content": "Summarize the history of Rome in 3 sentences."}]
print(f"Estimated prompt tokens: {count_tokens(messages)}")

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=messages,
)
usage = response.usage
print(f"Actual prompt={usage.prompt_tokens} completion={usage.completion_tokens} total={usage.total_tokens}")

How this code works

This code demonstrates how to estimate and accurately track token usage when interacting with OpenAI's chat models, a fundamental step in controlling costs for AI applications. It begins by importing the necessary tiktoken library for token counting and the openai library for API calls. An openai.OpenAI() client is initialized, automatically fetching the OPENAI_API_KEY from the environment. A tiktoken.encoding_for_model("gpt-4o") object is then set up to ensure token counting aligns with the chosen model's specific tokenization rules, which is crucial for accurate estimates.

The count_tokens function estimates the prompt's token usage before sending it to the API. It calculates tokens for each message's content using len(enc.encode(...)) and adds fixed overheads: 4 tokens per message and 2 tokens as a "reply primer" for the entire prompt. This per-message and primer overhead is a subtle detail often missed, but essential for aligning manual counts with the API's actual usage. After printing the Estimated prompt tokens, the code sends the messages to client.chat.completions.create, specifying model="gpt-4o-mini". The actual token usage, including prompt_tokens and completion_tokens, is then retrieved from response.usage, providing precise cost tracking data.

Production-grade example

Adds budget enforcement, async Redis spend tracking, structured cost logging, retries, and timeouts.

python
# openai>=1.0, redis>=5.0, structlog>=24.0
import asyncio
import os
import time
import structlog
import redis.asyncio as aioredis
from openai import AsyncOpenAI, APIStatusError, APITimeoutError

log = structlog.get_logger()
client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
redis_client = aioredis.from_url(os.environ["REDIS_URL"])

PRICE_PER_1K = {"gpt-4o": 0.005, "gpt-4o-mini": 0.000150}  # illustrative, verify with provider
HARD_LIMIT_TOKENS = int(os.environ.get("USER_MONTHLY_TOKEN_LIMIT", 500_000))

async def check_budget(user_id: str, model: str) -> None:
    key = f"token_spend:{user_id}:{time.strftime('%Y-%m')}"
    current = int(await redis_client.get(key) or 0)
    if current >= HARD_LIMIT_TOKENS:
        raise ValueError(f"User {user_id} exceeded monthly token budget ({current} tokens used)")

async def record_usage(user_id: str, feature: str, endpoint: str, model: str,
                       prompt_tokens: int, completion_tokens: int) -> None:
    total = prompt_tokens + completion_tokens
    cost_usd = (total / 1000) * PRICE_PER_1K.get(model, 0.005)
    key = f"token_spend:{user_id}:{time.strftime('%Y-%m')}"
    pipe = redis_client.pipeline()
    pipe.incrby(key, total)
    pipe.expire(key, 60 * 60 * 24 * 35)  # 35-day TTL
    await pipe.execute()
    log.info("llm_call_cost", user_id=user_id, feature=feature, endpoint=endpoint,
             model=model, prompt_tokens=prompt_tokens, completion_tokens=completion_tokens,
             total_tokens=total, cost_usd=round(cost_usd, 6))

async def tracked_chat(user_id: str, feature: str, endpoint: str,
                       messages: list[dict], model: str = "gpt-4o-mini") -> str:
    await check_budget(user_id, model)
    for attempt in range(3):
        try:
            t0 = time.monotonic()
            response = await client.chat.completions.create(model=model, messages=messages)
            latency_ms = round((time.monotonic() - t0) * 1000)
            usage = response.usage
            asyncio.create_task(record_usage(user_id, feature, endpoint, model,
                                             usage.prompt_tokens, usage.completion_tokens))
            log.info("llm_call_ok", user_id=user_id, model=model, latency_ms=latency_ms)
            return response.choices[0].message.content
        except APITimeoutError:
            wait = 2 ** attempt
            log.warning("llm_timeout_retry", attempt=attempt, wait_s=wait)
            await asyncio.sleep(wait)
        except APIStatusError as e:
            if e.status_code == 429:
                await asyncio.sleep(2 ** attempt)
            else:
                log.error("llm_api_error", status=e.status_code, body=e.message)
                raise
    raise RuntimeError("LLM call failed after 3 retries")

How this code works

This code's job is to manage and track the usage and cost of AI model calls (specifically OpenAI's chat completions) for individual users. It ensures users don't exceed a predefined HARD_LIMIT_TOKENS monthly budget and logs detailed cost information per user, feature, and endpoint, aiding in cost optimization. It initializes AsyncOpenAI for AI interactions and redis.asyncio for persistent data storage, loading necessary keys and URLs from environment variables. Illustrative PRICE_PER_1K values help calculate actual costs in USD.

The core logic first involves check_budget, which queries Redis using a token_spend:{user_id}:{year-month} key to verify a user's current token count against their limit, raising an error if exceeded. record_usage then calculates cost_usd and uses a Redis pipeline to incrby the user's monthly token total, setting a expire time of 35 days on the key for automatic monthly resets. It also logs detailed llm_call_cost data using structlog. The main tracked_chat function orchestrates this: it first calls check_budget, then attempts client.chat.completions.create with built-in retry logic for APITimeoutError or API errors (like a 429 rate limit). A subtle, but crucial, detail is using asyncio.create_task to dispatch record_usage to run in the background after a successful API call. This means the AI response is returned quickly to the user without waiting for database updates or detailed logging, improving perceived application performance.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a simple token budget enforcer. Write a function tracked_completion that wraps an OpenAI chat call, counts actual tokens used from the response, accumulates them per user in an in-memory dict, and raises a BudgetExceededError before calling the API if the user has already consumed more than 10,000 tokens. Print a cost summary after each call.

python
# openai>=1.0, tiktoken>=0.7
import os
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from env

class BudgetExceededError(Exception):
    pass

# In-memory token ledger: {user_id: total_tokens_used}
ledger: dict[str, int] = {}
USER_BUDGET = 10_000
PRICE_PER_1K = 0.000150  # gpt-4o-mini illustrative price

def tracked_completion(user_id: str, messages: list[dict], model: str = "gpt-4o-mini") -> str:
    # TODO 1: Check if user has exceeded USER_BUDGET; raise BudgetExceededError if so

    # TODO 2: Call the API

    # TODO 3: Extract token usage from response.usage

    # TODO 4: Update the ledger for this user

    # TODO 5: Print cost summary (tokens used this call, total for user, estimated USD)

    pass  # replace with return of message content

# Test it
for i in range(3):
    result = tracked_completion("user_42", [{"role": "user", "content": "Say hello in one sentence."}])
    print(result)

Quick check

  1. Why should you check a user's token budget BEFORE making the API call rather than after?

  2. Your async tracked_completion function logs cost data with await log_to_db(...) inline. What production problem does this cause?

  3. You notice your per-user token totals are 15% lower than the actual OpenAI invoice. What is the most likely cause?

Self-check: Explain how you would implement per-user monthly token budgets in a multi-tenant API with 10k users. Describe where budget state lives, how you enforce the limit without adding latency to 95% of requests, and what happens when your budget-check store is temporarily unreachable.