A tokenizer's job is to map arbitrary text to a fixed vocabulary of integer IDs so the model can do math on it. The dominant approach today is Byte-Pair Encoding (BPE). BPE starts with individual bytes, then iteratively merges the most frequent adjacent pairs until it hits a target vocabulary size (GPT-4's cl100k_base vocabulary has roughly 100,000 tokens). The result is a vocabulary where common English words are single tokens, rarer words split into two or three tokens, and non-English or highly technical text can fragment heavily. "Unambiguous" is one token. "Xylophone" is two ("Xyl", "ophone" in some tokenizers). A UUID might be eight. This is not an implementation detail -- it directly affects how many of your expensive context window slots each piece of content consumes.
Here is the real-world scenario that trips up most developers: you are building a customer support bot. You store the last 20 messages in the conversation array and send them all with every request. At 10 users that works fine. At 1,000 users you notice latency creeping up and costs doubling weekly. The problem is that your average conversation grew to 3,000 tokens of history, your system prompt is 800 tokens, and you are using a model with a 128k context window priced at $X per million tokens. You are paying for 3,800 input tokens per turn without thinking about it. A senior engineer would instrument token counts from day one, set a rolling-window budget for history (e.g., keep only the last N turns that fit under 2,000 tokens), and summarize older context rather than drop it.
The tradeoffs versus alternative approaches are real. You could compress context by summarizing: cheaper per-turn but adds latency and a second LLM call. You could use retrieval-augmented generation (RAG) to only inject relevant context: great for large knowledge bases, but requires a retrieval layer. You could switch to a model with a larger context window: removes the truncation risk but does not always reduce cost -- a 1M-token window model may charge more per token than a 128k one. There is no free lunch; the right answer depends on your read-to-write ratio, latency budget, and how coherent the conversation needs to be across many turns.
At different scales, tokenization concerns shift. With 10 users you mostly care about correctness -- does your prompt fit and does the response complete? With 10,000 users you start caring about p95 latency (longer token counts mean longer time-to-first-token and longer streaming duration) and monthly costs. At 10 million users, token count directly determines your gross margin. A 200-token reduction in your average system prompt could save tens of thousands of dollars a month. Teams at that scale run automated prompt audits, A/B test prompt lengths for quality vs cost, and cache identical or near-identical prompt prefixes (a feature some providers offer natively) to avoid re-processing the same system prompt on every call.
Cost and latency are tightly coupled to token counts in two specific ways. First, time-to-first-token scales with input length because the model must process all input tokens before generating the first output token (this is the prefill phase). Second, output tokens are generated one at a time (autoregressive decoding), so total output length is the primary driver of end-to-end latency. If your use case allows it, constraining max_tokens in your API call both caps cost and guarantees a response time SLA. Always set max_tokens explicitly -- letting the model run until its natural stop means unbounded cost and latency in pathological cases.
Key Takeaways
- Count tokens before sending requests; never assume character count is a safe proxy.
- Context window = input tokens + output tokens; both consume the same fixed budget.
- Output tokens typically cost more than input tokens -- optimize response length deliberately.
- Use the model's own tokenizer (tiktoken, tokenizers library) for accurate counts in code.
Pro tips
- Different models use different tokenizers even at the same provider. GPT-3.5 and GPT-4 use cl100k_base but Claude uses its own BPE vocabulary. Run your actual prompt through the model's specific tokenizer before assuming a count is portable across providers.
- Non-English text tokenizes much less efficiently than English. A 1,000-character Chinese or Arabic string can consume 2-3x more tokens than an equivalent English string. If your app handles multilingual input, benchmark token consumption per language before setting context budgets.
- Caching the system prompt at the API level (Anthropic calls this 'prompt caching', some OpenAI endpoints support it natively) can cut costs by 80-90% on repeated calls that share a long prefix. It is one of the highest-leverage optimizations available once you are at scale.
- The token count you estimate client-side with tiktoken will sometimes differ by 1-3 tokens from what the server bills you. Always use the usage object in the API response for billing reconciliation, not your local estimate.
Common pitfalls
- Mistake: Using len(text.split()) as a token count estimate. Fix: Always use the model's tokenizer (tiktoken for OpenAI, tokenizers library for HuggingFace models); word count can be off by 30-50% for technical or mixed-language content.
- Mistake: Not setting max_tokens on the API call, letting the model respond at arbitrary length. Fix: Always set max_tokens; unbounded output drives unpredictable latency and cost spikes in production.
- Mistake: Counting only input tokens when monitoring costs. Fix: Log prompt_tokens and completion_tokens separately from the response usage object; output tokens are often 2-4x more expensive per token.
- Mistake: Assuming the full context window is safe to fill on every call. Fix: Leave headroom; at max context length, some models degrade in coherence (the 'lost-in-the-middle' problem) and latency spikes significantly.
When to use which context management strategy
| Option | Use when | Avoid when |
|---|---|---|
| Send full conversation history | Conversation is short (under 20 turns), coherence across entire history is critical, latency budget is relaxed. | Conversations grow long, cost per call is a constraint, or you hit the context window ceiling regularly. |
| Rolling window (keep last N turns) | You need a simple, low-latency solution and recent context is sufficient for the task. | The user refers back to facts established early in the conversation; dropped turns cause confusing responses. |
| Summarize older turns | You need coherence across long sessions but cannot afford full history; acceptable to add one extra LLM call. | Latency is tight (summarization adds a round-trip) or when exact wording from earlier messages matters. |
| RAG / inject only relevant context | You have a large knowledge base (docs, past tickets) and only a fraction is relevant per query. | The application requires strict turn-by-turn conversational memory; retrieval may miss implicit references. |
Code Example
# tiktoken 0.7.x -- pip install tiktoken
import tiktoken
enc = tiktoken.encoding_for_model("gpt-4o")
text = "Tokenization splits text into subword chunks before the model sees it."
tokens = enc.encode(text)
print(f"Text : {text}")
print(f"Token IDs : {tokens}")
print(f"Token count: {len(tokens)}")
print(f"Decoded : {[enc.decode([t]) for t in tokens]}")How this code works
This code demonstrates the crucial first step an LLM takes with any input text: tokenization. It uses the tiktoken library, which is the same tool OpenAI models like GPT-4o use, to break down a sentence into smaller, numerical units called tokens. Understanding this process is key to grasping how LLMs process information, manage input limits, and calculate costs. The setup involves ensuring tiktoken is installed with pip install tiktoken, then importing it. Next, tiktoken.encoding_for_model("gpt-4o") is used to load the specific tokenization rules that GPT-4o understands, ensuring the process mirrors what happens internally in the actual model.
Once the encoding rules are loaded into the enc object, the core of the process happens. enc.encode(text) takes the example sentence and converts it into a list of numbers, where each number is a unique token ID. The len(tokens) then reveals the total token count for that text, directly showing how much "space" the text would occupy in an LLM's context window. Finally, the list comprehension [enc.decode([t]) for t in tokens] reconstructs the original text piece by piece. A subtle point here is that enc.decode expects a list of token IDs, even when decoding a single token, which is why [t] is used. This output often shows that tokens aren't always whole words but can be parts of words or punctuation, reflecting the subword nature of LLM tokenization.
Production-grade example
Adds token budget guard, exponential-backoff retries, per-call cost logging, and explicit timeout.
# tiktoken 0.7.x, openai 1.x -- pip install tiktoken openai tenacity structlog
import os
import time
import logging
import tiktoken
from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
import structlog
log = structlog.get_logger()
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
MODEL = "gpt-4o-mini"
MAX_INPUT_TOKENS = 3_000
MAX_OUTPUT_TOKENS = 512
PRICE_PER_1K_INPUT = 0.00015 # illustrative -- verify current pricing
PRICE_PER_1K_OUTPUT = 0.00060 # illustrative
enc = tiktoken.encoding_for_model(MODEL)
def count_tokens(messages: list[dict]) -> int:
"""Approximate token count for a messages array (OpenAI chat format)."""
total = sum(4 + len(enc.encode(m.get("content", ""))) for m in messages)
return total + 2 # reply priming overhead
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
)
def chat(messages: list[dict]) -> str:
input_tokens = count_tokens(messages)
if input_tokens > MAX_INPUT_TOKENS:
raise ValueError(f"Prompt too large: {input_tokens} tokens (limit {MAX_INPUT_TOKENS})")
log.info("llm_request", model=MODEL, input_tokens=input_tokens)
t0 = time.monotonic()
response = client.chat.completions.create(
model=MODEL,
messages=messages,
max_tokens=MAX_OUTPUT_TOKENS,
stream=False,
)
latency_ms = int((time.monotonic() - t0) * 1000)
usage = response.usage
cost = (
usage.prompt_tokens / 1000 * PRICE_PER_1K_INPUT
+ usage.completion_tokens / 1000 * PRICE_PER_1K_OUTPUT
)
log.info(
"llm_response",
model=MODEL,
prompt_tokens=usage.prompt_tokens,
completion_tokens=usage.completion_tokens,
latency_ms=latency_ms,
estimated_cost_usd=round(cost, 6),
)
return response.choices[0].message.content
if __name__ == "__main__":
msgs = [{"role": "user", "content": "Explain BPE tokenization in two sentences."}]
print(chat(msgs))How this code works
This Python code provides a robust way to interact with OpenAI's LLMs, central to managing token limits and understanding cost implications. It configures an OpenAI client with an api_key and timeout, selecting a MODEL like "gpt-4o-mini" and defining MAX_INPUT_TOKENS, MAX_OUTPUT_TOKENS, and illustrative pricing for cost estimation. Crucially, it initializes tiktoken.encoding_for_model(MODEL) to accurately tokenize messages before sending them to the API.
The count_tokens function uses this encoder to approximate the token count for a message list, including a subtle adjustment of sum(4 + ...) and + 2 to account for OpenAI's internal overhead tokens per message and the reply. The chat function first calls count_tokens to check if the prompt exceeds MAX_INPUT_TOKENS, raising a ValueError if it’s too large. It then uses the @retry decorator from tenacity to automatically re-attempt API calls on RateLimitError or APITimeoutError. The client.chat.completions.create call sends the request, and upon receiving a response, the code calculates estimated_cost_usd based on usage.prompt_tokens and usage.completion_tokens, logging these details with structlog for monitoring.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a function that takes a list of chat messages (OpenAI format) and a token budget, then trims the oldest non-system messages until the total token count fits within the budget. Print a before/after token count. Use tiktoken with the gpt-4o-mini model. Do not drop the system message.
# pip install tiktoken
import tiktoken
MODEL = "gpt-4o-mini"
TOKEN_BUDGET = 200
enc = tiktoken.encoding_for_model(MODEL)
def count_tokens(messages: list[dict]) -> int:
# TODO: sum token counts for each message content + 4 overhead each + 2 priming
pass
def trim_to_budget(messages: list[dict], budget: int) -> list[dict]:
# TODO: keep the system message (role == 'system') always
# TODO: remove oldest non-system messages one at a time until under budget
pass
if __name__ == "__main__":
msgs = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Tell me about the French Revolution."},
{"role": "assistant", "content": "The French Revolution began in 1789 and fundamentally transformed France."},
{"role": "user", "content": "What caused it?"},
]
# TODO: print before count, trim, print after count and resulting messagesQuick check
A request uses 800 input tokens and the model returns 200 output tokens. If input costs $0.001/1k and output costs $0.004/1k, what is the total cost?
You set max_tokens=100 but the model naturally wants to produce 150 tokens. What happens?
Why does the same English sentence tokenize into more tokens when written in a less common programming language's identifier style (e.g., snake_case_long_variable_name)?