The mental model here is a capability-cost curve, not a single point. Every provider offers a family of models at different price points. OpenAI has gpt-4o-mini and gpt-4o. Anthropic has Haiku, Sonnet, and Opus. Google has Gemini Flash and Gemini Pro. Mistral has Mistral 7B, Mixtral 8x7B, and Mistral Large. Each family member exists because a meaningful fraction of real-world tasks do not require the full capability of the flagship. Your job is to figure out which fraction of your traffic that is.
The practical starting point is task taxonomy. Go through your application and classify every LLM call by what it actually does: extraction (pull structured fields from text), classification (bucket input into categories), summarization (shorten a document), generation (produce new content), reasoning (multi-step logic or code), and conversation (maintain context across turns). Extraction, classification, and simple summarization are almost always fine on smaller models. Reasoning, complex code generation, and long multi-turn conversations with nuanced context are where flagship models earn their price. The failure mode to watch for is treating your application as a single category when it contains multiple.
A real-world scenario: you're building a support ticket triage system that handles 50,000 tickets per day. Every ticket goes through three LLM calls: (1) classify the ticket into one of 12 categories, (2) extract structured fields like product name, severity, and affected version, (3) draft a suggested reply. If you run all three calls on Claude 3.5 Sonnet at roughly $3 per million output tokens (illustrative, check current pricing), your daily output token bill at an average of 200 output tokens per call is around $90/day just for output. Swapping steps 1 and 2 to Claude 3 Haiku at roughly $0.25 per million output tokens cuts that portion of your bill by 10x. Step 3 might still need Sonnet because the reply quality matters to your users. This selective swap often cuts total model cost by 50-70% with zero user-visible quality change.
The tradeoff against alternative approaches is worth spelling out. Prompt caching reduces cost when you have repeated prefixes but does nothing when every call is unique. Self-hosted open-source models can get you below the per-token costs of any API provider for high-volume tasks, but they introduce infrastructure complexity, GPU provisioning, and your own reliability SLA. Model selection via provider APIs sits in the middle: no infrastructure overhead, immediate access to new model versions, but you are still paying per token. For 10k monthly API calls, API-based tiering almost always beats self-hosting. For 10M calls on a narrow, stable task, self-hosting becomes worth evaluating. These approaches are not mutually exclusive: cache repeated calls, route by tier, and self-host only the highest-volume stable workloads.
Scale changes the math significantly. At 10 users, model choice is irrelevant to your bill. At 10k users, tiered routing pays for itself in the first week. At 10M users, you need routing logic that is itself fast and cheap, because adding a classification call to decide which model to use adds latency and its own cost. At that scale you tend to route on deterministic signals: if the input is under 200 tokens and matches a known template, it goes to the cheap model; only long or ambiguous inputs go to the flagship. You also need to instrument your routing layer carefully because a bug that routes everything to the flagship is a serious incident, not just a performance issue.
Latency is a second-order effect of model selection that often matters as much as cost. gpt-4o-mini produces first tokens roughly 2-3x faster than gpt-4o in typical conditions. For a real-time chat interface, that is the difference between a snappy and a sluggish experience. For a batch pipeline running overnight, latency is irrelevant. Factor in the usage context before choosing: interactive features should weight latency heavily, async workflows should weight throughput and cost instead.
Key Takeaways
- Profile your tasks first; most production workloads have 60-80% of calls that a cheaper model handles equally well.
- Build a routing layer that selects model tier based on task signals, not just a single global default.
- Always benchmark on your own data with your own eval rubric before committing to a model swap.
- Treat model selection as a continuous process: re-evaluate when providers release new models or change pricing.
Pro tips
- Run your eval suite on both the cheap and flagship model before shipping any routing logic. The cheap model often outperforms the flagship on narrow, well-defined tasks because it is less likely to overthink a simple classification prompt.
- Log model name, prompt token count, and completion token count on every single call from day one. Without that data, you cannot make a justified model swap decision later, and you will be flying blind when costs spike.
- The routing signal does not have to be a second LLM call. Heuristics like input length, presence of code blocks, number of reasoning steps requested, or which product feature triggered the call are cheap and fast routing signals you can compute locally.
- New model releases frequently obsolete your routing tiers. When Anthropic released Haiku, tasks that previously needed Sonnet dropped a tier. Subscribe to provider changelog emails and put a recurring calendar reminder to re-benchmark your tier boundaries quarterly.
Common pitfalls
- Mistake: Benchmarking models on public datasets or toy examples, not your production data. Fix: Always use a sample of real production inputs and a rubric tied to your actual quality bar before making a routing decision.
- Mistake: Routing by task type using only LLM-provided confidence scores. Fix: Confidence outputs from LLMs are poorly calibrated; use deterministic signals like input length or template matching as your primary routing gate.
- Mistake: Swapping the model and assuming quality held because aggregate metrics look fine. Fix: Segment your eval by input type and edge cases; aggregate accuracy can hide a 20% failure rate on a specific subcategory that matters most to users.
- Mistake: Never updating the model tier map after a provider releases new models. Fix: Treat model selection as a living config, not a one-time decision; new mid-tier models often outperform old flagships at a fraction of the cost.
When to use a smaller model vs a flagship model
| Option | Use when | Avoid when |
|---|---|---|
| Smaller/cheaper model (e.g., gpt-4o-mini, Claude Haiku) | Task is well-defined: classification, extraction, short summarization, template filling, or high-volume batch jobs where latency matters. | Task requires multi-step reasoning, nuanced judgment, long-context synthesis, or output quality directly drives revenue or safety outcomes. |
| Flagship model (e.g., gpt-4o, Claude 3.5 Sonnet) | Task involves complex reasoning, code generation, long multi-turn context, or you have no eval data yet and need reliable baseline quality. | Task is high volume and routine; the cost difference is an order of magnitude and your eval shows no quality gap on smaller models. |
| Fine-tuned smaller model | You have a narrow, stable, high-volume task and labeled training data; you want flagship-quality on that specific task at mini pricing. | Task distribution shifts frequently, you lack training data, or the fine-tuning iteration cycle is too slow for your release cadence. |
| Tiered routing (both model sizes) | Your application has mixed task complexity and you can derive a cheap routing signal (input length, template match, feature flag) without an extra LLM call. | All your requests genuinely require the same capability level, or routing logic complexity exceeds the savings it produces. |
Code Example
# openai>=1.0.0, anthropic>=0.25.0
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
TASK_MODEL_MAP = {
"classify": "gpt-4o-mini",
"extract": "gpt-4o-mini",
"summarize": "gpt-4o-mini",
"reason": "gpt-4o",
"draft_reply": "gpt-4o",
}
def call_llm(task_type: str, prompt: str) -> str:
model = TASK_MODEL_MAP.get(task_type, "gpt-4o-mini")
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_tokens=512,
)
return response.choices[0].message.content
# classify goes to mini, reasoning goes to flagship
print(call_llm("classify", "Classify this ticket: 'Login button not working on mobile'"))
print(call_llm("reason", "Debug this Python traceback and suggest a fix: ..."))How this code works
This code demonstrates a clever cost-optimization strategy: using different AI models based on the task's complexity. Instead of always defaulting to the most powerful and expensive models, it intelligently routes requests to a cheaper, smaller model or a more capable flagship model as needed. The TASK_MODEL_MAP dictionary is key to this, explicitly associating task types like "classify" and "extract" with the more economical gpt-4o-mini and more demanding tasks like "reason" or "draft_reply" with the powerful gpt-4o flagship.
The call_llm function orchestrates this selection. It takes a task_type and a prompt, then uses TASK_MODEL_MAP.get(task_type, "gpt-4o-mini") to choose the appropriate model. A subtle but important detail is the .get() method's second argument: if a task_type isn't explicitly listed in the map, the function silently defaults to gpt-4o-mini, ensuring cost efficiency even for unhandled cases. After selecting the model, it makes an API call via client.chat.completions.create to generate a response, limiting its length with max_tokens=512. This prevents overspending by matching model capability to task requirements.
Production-grade example
Adds per-call cost logging, typed retries, timeout config, and graceful model degradation on failure.
# openai>=1.0.0, tenacity>=8.0.0, structlog>=24.0.0
import os
import time
import structlog
from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
log = structlog.get_logger()
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
TASK_MODEL_MAP = {
"classify": ("gpt-4o-mini", 0.15, 0.60), # (model, $/1M in, $/1M out) -- illustrative
"extract": ("gpt-4o-mini", 0.15, 0.60),
"summarize": ("gpt-4o-mini", 0.15, 0.60),
"reason": ("gpt-4o", 2.50, 10.00),
"draft_reply":("gpt-4o", 2.50, 10.00),
}
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
)
def call_llm(
task_type: str,
prompt: str,
fallback_task: str = "classify",
) -> str:
model, cost_in, cost_out = TASK_MODEL_MAP.get(task_type, TASK_MODEL_MAP[fallback_task])
start = time.monotonic()
try:
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_tokens=512,
)
except APIStatusError as exc:
log.error("llm_api_error", task=task_type, model=model, status=exc.status_code)
raise
usage = response.usage
elapsed = time.monotonic() - start
estimated_cost = (usage.prompt_tokens * cost_in + usage.completion_tokens * cost_out) / 1_000_000
log.info(
"llm_call",
task=task_type,
model=model,
prompt_tokens=usage.prompt_tokens,
completion_tokens=usage.completion_tokens,
estimated_cost_usd=round(estimated_cost, 6),
latency_s=round(elapsed, 3),
)
return response.choices[0].message.content
# graceful degradation: complex task falls back to mini if flagship fails
def call_with_degradation(task_type: str, prompt: str) -> str:
try:
return call_llm(task_type, prompt)
except Exception:
log.warning("degrading_to_mini", original_task=task_type)
return call_llm("classify", prompt) # cheapest tier as last resortHow this code works
This code demonstrates how to optimize costs by selecting appropriate AI models for different tasks and building resilience into API calls. The TASK_MODEL_MAP dictionary is central, defining which specific model, like gpt-4o-mini for "classify" or gpt-4o for "reason," should be used for each task_type, along with their illustrative costs per million tokens. This direct mapping ensures that less complex tasks use cheaper, smaller models, while more complex ones leverage flagship models, directly supporting cost optimization.
The call_llm function handles the core interaction with the OpenAI API. It automatically retries failed calls for transient issues like RateLimitError or APITimeoutError thanks to the @retry decorator, improving reliability. A subtle but important detail is the fallback_task="classify" default argument within call_llm: if a task_type isn't found, it gracefully defaults to using the cheapest "classify" task model. Beyond this, call_with_degradation adds another layer of robustness; if an initial call_llm attempt fails completely (even after retries), it catches the error and retries the request with the "classify" task, ensuring some response is provided, albeit from the cheapest model, rather than a total failure.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a simple tiered router that classifies support tickets into 'billing', 'technical', or 'general' using gpt-4o-mini, then generates a suggested reply using gpt-4o only for 'technical' tickets (all others get a reply from gpt-4o-mini). Log the model used and estimated token cost for each call.
# openai>=1.0.0
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
TICKETS = [
"My invoice shows a charge I don't recognize.",
"The API returns a 500 error when I POST to /v2/process with a large payload.",
"How do I reset my password?",
]
def classify_ticket(text: str) -> str:
# TODO: call gpt-4o-mini to classify into billing/technical/general
# return the category string
pass
def draft_reply(category: str, text: str) -> str:
# TODO: use gpt-4o for 'technical', gpt-4o-mini for everything else
# log model name and estimated cost (input_tokens * 0.15/1M + output_tokens * 0.60/1M for mini)
pass
for ticket in TICKETS:
category = classify_ticket(ticket)
reply = draft_reply(category, ticket)
print(f"Category: {category}\nReply: {reply}\n")Quick check
You have a task classifying emails into 10 fixed categories with 200k calls per day. Which approach is most cost-effective without sacrificing quality?
Your routing layer uses a second LLM call to decide which model tier to use for the main call. What is the main problem with this approach at 10M calls per day?
After swapping from GPT-4o to gpt-4o-mini for summarization, your aggregate ROUGE score stays the same. Is that sufficient evidence the swap was safe?