At its core, token-based cost tracking is just request instrumentation with a pricing dimension. Every provider returns a usage object on each API response -- prompt_tokens, completion_tokens, and total_tokens. Those numbers, multiplied by the model's per-token rate, give you the dollar cost of that single call. The challenge isn't the math; it's attaching the right metadata so you can aggregate meaningfully later. A raw token count without knowing which user triggered it or which feature it served is nearly useless for optimization.
The mental model that works well in practice is treating each LLM call as a database query with an execution cost. You wouldn't ship a product without knowing which queries are slow. Same logic applies here. Instrument every call site to capture: timestamp, user_id (or a hashed tenant ID), feature slug (e.g., "doc_summarization", "chat_assistant"), endpoint label (e.g., "/api/v1/analyze"), model name, prompt tokens, completion tokens, and derived cost in USD. Persist that to a time-series-friendly store -- a Postgres table with an index on (user_id, created_at) works fine at moderate scale; ClickHouse or BigQuery at higher scale. Avoid logging raw prompt text by default: that's a data-privacy and storage cost problem.
Budget enforcement has two flavors: soft limits and hard limits. A soft limit warns the user or throttles their request rate when they approach their allocation. A hard limit rejects the request before it even reaches the LLM API. For a SaaS product with tiered plans, you might set a hard limit of 500k tokens per user per month on a free tier, with a soft warning at 80%. Implement hard limits as a pre-call guard: look up the user's token spend for the current window from your store, compare to their budget, and return an HTTP 429 with a clear message if exceeded. This prevents a single misbehaving user from generating a surprise invoice item. Implement the check efficiently -- a Redis counter with a 30-day TTL is faster than a Postgres aggregate for hot paths.
In the real world, this infrastructure pays off immediately in a multi-tenant SaaS setting. Imagine you've shipped a document summarization feature that calls GPT-4o. Three weeks after launch, your monthly OpenAI invoice jumps 3x. With no tracking, you're investigating cold. With tracking, you pull a GROUP BY feature query and see "doc_summarization" accounts for 71% of spend, and two enterprise trial accounts are summarizing hundreds of PDFs per day. You can immediately gate that feature behind a rate limit, contact those accounts about an enterprise plan, or switch the feature to a cheaper model (see the model-selection subtopic for that decision). None of those responses are available to you without the data.
At scale, the instrumentation approach changes shape. At 10 users, a SQLite table and a print statement is enough to learn the pattern. At 10k users, you want asynchronous logging so cost tracking doesn't add latency to the critical path -- fire-and-forget to a background queue (asyncio.create_task or a Celery worker). At 10M users, the ingestion volume of cost events is itself a real infrastructure problem. You batch-write to ClickHouse or BigQuery, run daily rollup aggregates, and query pre-aggregated tables rather than raw events. Budget checks at that scale hit Redis or Memcached for the current spend counter rather than querying raw event rows. The shape of the problem shifts from "how do I store this" to "how do I keep the check under 5ms." Start simple and let traffic drive the architecture.
Key Takeaways
- Count tokens at the call site using provider SDKs or tiktoken; never estimate.
- Tag every LLM call with user_id, feature, and endpoint before persisting cost data.
- Enforce soft limits with warnings and hard limits with early exits before the API call.
- Use time-windowed budgets (daily, monthly) so a burst doesn't silently exhaust a monthly quota.
Pro tips
- Pre-count prompt tokens with tiktoken before the API call, not just after. This lets you enforce a per-request token cap and reject oversized inputs before you pay for them -- a single jailbreak attempt with a 100k-token injected document is otherwise a real line item.
- Track completion tokens separately from prompt tokens. Completions are often 2-5x more expensive per token on some models, and a feature where users set max_tokens=4096 by default can silently dominate your bill even with modest prompt sizes.
- Roll up cost data to a secondary table by (user_id, feature, date) at write time or via a nightly job. Querying raw event rows for a monthly user report at 10M events is slow; pre-aggregated rollups make dashboards instant and keep your analytics DB happy.
- Redis counters for budget checks are fast but eventually consistent with your persistent store. Write the authoritative event log to Postgres or BigQuery asynchronously, and use Redis only as the hot-path gate. If Redis is unavailable, fail open with a log warning rather than blocking all LLM calls.
Common pitfalls
- Mistake: Logging cost data synchronously on the critical path, adding 20-50ms to every response. Fix: Use asyncio.create_task or a background queue (Celery, ARQ) so the API response returns immediately.
- Mistake: Using character count or word count as a proxy for token count. Fix: Always use tiktoken or the provider's token counting endpoint; character-to-token ratios vary by language and content type.
- Mistake: Setting a single global monthly budget without user or feature dimension, so one heavy feature exhausts the budget for all users. Fix: Track budgets per user and per feature independently so enforcement is surgical.
- Mistake: Forgetting to account for system prompt tokens in budget checks, which can be 300-800 tokens per call. Fix: Include the full message list -- system, few-shot examples, and conversation history -- in your pre-call token count.
Where to store token cost events
| Option | Use when | Avoid when |
|---|---|---|
| Redis counters only | You only need real-time budget enforcement and don't require historical cost queries. | You need per-feature breakdowns, audit trails, or monthly invoice reconciliation. |
| Postgres event log | Under ~50M events/month; team already operates Postgres; need SQL queries for ad-hoc analysis. | Event volume exceeds what a single Postgres instance handles comfortably without heavy partitioning. |
| ClickHouse or BigQuery | High event volume, complex aggregations, or BI dashboards querying raw events across millions of rows. | You're early-stage; the operational overhead isn't justified until you have real scale. |
| Third-party (Helicone, LangSmith, OpenMeter) | You want cost dashboards out of the box and don't want to build the pipeline from scratch. | Your data cannot leave your infrastructure due to compliance requirements, or you need deep custom logic. |
Code Example
# tiktoken 0.7.x, openai 1.x
import tiktoken
import openai
client = openai.OpenAI() # reads OPENAI_API_KEY from env
enc = tiktoken.encoding_for_model("gpt-4o")
def count_tokens(messages: list[dict]) -> int:
"""Estimate prompt token count before sending."""
total = 0
for msg in messages:
total += 4 # per-message overhead
total += len(enc.encode(msg["content"]))
return total + 2 # reply primer
messages = [{"role": "user", "content": "Summarize the history of Rome in 3 sentences."}]
print(f"Estimated prompt tokens: {count_tokens(messages)}")
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
)
usage = response.usage
print(f"Actual prompt={usage.prompt_tokens} completion={usage.completion_tokens} total={usage.total_tokens}")How this code works
This code demonstrates how to estimate and accurately track token usage when interacting with OpenAI's chat models, a fundamental step in controlling costs for AI applications. It begins by importing the necessary tiktoken library for token counting and the openai library for API calls. An openai.OpenAI() client is initialized, automatically fetching the OPENAI_API_KEY from the environment. A tiktoken.encoding_for_model("gpt-4o") object is then set up to ensure token counting aligns with the chosen model's specific tokenization rules, which is crucial for accurate estimates.
The count_tokens function estimates the prompt's token usage before sending it to the API. It calculates tokens for each message's content using len(enc.encode(...)) and adds fixed overheads: 4 tokens per message and 2 tokens as a "reply primer" for the entire prompt. This per-message and primer overhead is a subtle detail often missed, but essential for aligning manual counts with the API's actual usage. After printing the Estimated prompt tokens, the code sends the messages to client.chat.completions.create, specifying model="gpt-4o-mini". The actual token usage, including prompt_tokens and completion_tokens, is then retrieved from response.usage, providing precise cost tracking data.
Production-grade example
Adds budget enforcement, async Redis spend tracking, structured cost logging, retries, and timeouts.
# openai>=1.0, redis>=5.0, structlog>=24.0
import asyncio
import os
import time
import structlog
import redis.asyncio as aioredis
from openai import AsyncOpenAI, APIStatusError, APITimeoutError
log = structlog.get_logger()
client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
redis_client = aioredis.from_url(os.environ["REDIS_URL"])
PRICE_PER_1K = {"gpt-4o": 0.005, "gpt-4o-mini": 0.000150} # illustrative, verify with provider
HARD_LIMIT_TOKENS = int(os.environ.get("USER_MONTHLY_TOKEN_LIMIT", 500_000))
async def check_budget(user_id: str, model: str) -> None:
key = f"token_spend:{user_id}:{time.strftime('%Y-%m')}"
current = int(await redis_client.get(key) or 0)
if current >= HARD_LIMIT_TOKENS:
raise ValueError(f"User {user_id} exceeded monthly token budget ({current} tokens used)")
async def record_usage(user_id: str, feature: str, endpoint: str, model: str,
prompt_tokens: int, completion_tokens: int) -> None:
total = prompt_tokens + completion_tokens
cost_usd = (total / 1000) * PRICE_PER_1K.get(model, 0.005)
key = f"token_spend:{user_id}:{time.strftime('%Y-%m')}"
pipe = redis_client.pipeline()
pipe.incrby(key, total)
pipe.expire(key, 60 * 60 * 24 * 35) # 35-day TTL
await pipe.execute()
log.info("llm_call_cost", user_id=user_id, feature=feature, endpoint=endpoint,
model=model, prompt_tokens=prompt_tokens, completion_tokens=completion_tokens,
total_tokens=total, cost_usd=round(cost_usd, 6))
async def tracked_chat(user_id: str, feature: str, endpoint: str,
messages: list[dict], model: str = "gpt-4o-mini") -> str:
await check_budget(user_id, model)
for attempt in range(3):
try:
t0 = time.monotonic()
response = await client.chat.completions.create(model=model, messages=messages)
latency_ms = round((time.monotonic() - t0) * 1000)
usage = response.usage
asyncio.create_task(record_usage(user_id, feature, endpoint, model,
usage.prompt_tokens, usage.completion_tokens))
log.info("llm_call_ok", user_id=user_id, model=model, latency_ms=latency_ms)
return response.choices[0].message.content
except APITimeoutError:
wait = 2 ** attempt
log.warning("llm_timeout_retry", attempt=attempt, wait_s=wait)
await asyncio.sleep(wait)
except APIStatusError as e:
if e.status_code == 429:
await asyncio.sleep(2 ** attempt)
else:
log.error("llm_api_error", status=e.status_code, body=e.message)
raise
raise RuntimeError("LLM call failed after 3 retries")How this code works
This code's job is to manage and track the usage and cost of AI model calls (specifically OpenAI's chat completions) for individual users. It ensures users don't exceed a predefined HARD_LIMIT_TOKENS monthly budget and logs detailed cost information per user, feature, and endpoint, aiding in cost optimization. It initializes AsyncOpenAI for AI interactions and redis.asyncio for persistent data storage, loading necessary keys and URLs from environment variables. Illustrative PRICE_PER_1K values help calculate actual costs in USD.
The core logic first involves check_budget, which queries Redis using a token_spend:{user_id}:{year-month} key to verify a user's current token count against their limit, raising an error if exceeded. record_usage then calculates cost_usd and uses a Redis pipeline to incrby the user's monthly token total, setting a expire time of 35 days on the key for automatic monthly resets. It also logs detailed llm_call_cost data using structlog. The main tracked_chat function orchestrates this: it first calls check_budget, then attempts client.chat.completions.create with built-in retry logic for APITimeoutError or API errors (like a 429 rate limit). A subtle, but crucial, detail is using asyncio.create_task to dispatch record_usage to run in the background after a successful API call. This means the AI response is returned quickly to the user without waiting for database updates or detailed logging, improving perceived application performance.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a simple token budget enforcer. Write a function tracked_completion that wraps an OpenAI chat call, counts actual tokens used from the response, accumulates them per user in an in-memory dict, and raises a BudgetExceededError before calling the API if the user has already consumed more than 10,000 tokens. Print a cost summary after each call.
# openai>=1.0, tiktoken>=0.7
import os
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from env
class BudgetExceededError(Exception):
pass
# In-memory token ledger: {user_id: total_tokens_used}
ledger: dict[str, int] = {}
USER_BUDGET = 10_000
PRICE_PER_1K = 0.000150 # gpt-4o-mini illustrative price
def tracked_completion(user_id: str, messages: list[dict], model: str = "gpt-4o-mini") -> str:
# TODO 1: Check if user has exceeded USER_BUDGET; raise BudgetExceededError if so
# TODO 2: Call the API
# TODO 3: Extract token usage from response.usage
# TODO 4: Update the ledger for this user
# TODO 5: Print cost summary (tokens used this call, total for user, estimated USD)
pass # replace with return of message content
# Test it
for i in range(3):
result = tracked_completion("user_42", [{"role": "user", "content": "Say hello in one sentence."}])
print(result)
Quick check
Why should you check a user's token budget BEFORE making the API call rather than after?
Your async tracked_completion function logs cost data with
await log_to_db(...)inline. What production problem does this cause?You notice your per-user token totals are 15% lower than the actual OpenAI invoice. What is the most likely cause?