Model routing is fundamentally a classification problem wrapped around a dispatch table. The orchestrator receives a request, produces a routing signal (simple/medium/complex, or a score, or a capability tag), and maps that signal to a model tier. The routing signal can come from heuristics, a dedicated classifier model, or even a cheap LLM call. Each approach has different latency and accuracy tradeoffs, which we will cover below.
The mental model to internalize: think of your model fleet as a staffing hierarchy. A call center routes quick account balance queries to an IVR, standard billing questions to a tier-1 agent, and disputed charges to a specialist. The routing logic itself is overhead, so it must be faster and cheaper than the savings it generates. If your classifier adds 100ms and a GPT-4o-mini call already takes 300ms, your simple-task path now takes 400ms. That is often fine, but you have to measure it.
A real-world scenario where routing pays off: imagine a customer support bot handling 50,000 messages per day. About 60% are simple deflections: order status, store hours, return policy. These are answerable with a 1B-parameter local model or a cheap hosted model in under 200ms. The remaining 40% involve account lookups, multi-step reasoning, or emotionally complex situations that benefit from a stronger model. Without routing, you pay premium rates for every message. With routing, your cost per message on the simple tier drops by roughly 10-20x (numbers are illustrative; benchmark your own use case). A senior engineer builds this incrementally: first instrument all requests to log prompt length, token counts, and model used. Then analyze the logs to find the natural complexity distribution. Only then design tiers around that distribution rather than guessing upfront.
Tradeoffs versus alternatives: the simplest alternative to routing is just using one model for everything. That is a valid starting point -- routing adds operational complexity and a new failure mode (misroutes). A second alternative is prompt compression: keep using the expensive model but shrink the input. Compression and routing are complementary, not competing. A third alternative is caching: store the output of identical or near-identical requests. Caching and routing stack well -- cache hits bypass routing entirely. The routing approach shines specifically when your request distribution has high variance in complexity, your cost-per-request budget is tight, and you can afford a small accuracy penalty on complex tasks occasionally landing on a weaker model.
Routing logic options, from simplest to most sophisticated: (1) Rule-based heuristics: token count, keyword presence, regex patterns. Zero latency overhead, no extra API call, but brittle. Works well for structured domains where complexity signals are reliable. (2) Embedding-based classifier: embed the prompt, run a trained logistic regression or small neural net against complexity labels from your evaluation dataset. Adds 20-50ms for the embedding call but generalizes better than pure rules. (3) Small LLM as router: send the prompt to a cheap, fast model (GPT-4o-mini, Claude Haiku, Gemini Flash) and ask it to return a structured routing decision. Adds one full LLM call but handles nuanced cases. Use structured output/JSON mode to make parsing reliable. (4) Learned router: train a small BERT-class classifier on labeled examples from your production logs where you know which model actually produced a better answer. This is the highest-accuracy approach but requires a labeled dataset and a training pipeline.
What changes at scale: at 10 users per day, heuristics are fine. At 10,000 users per day, you want to cache routing decisions alongside response caches so repeated prompts skip classification. At 10 million users per day, you need the routing layer to be horizontally scalable and stateless, emit structured logs to a data warehouse, and support A/B testing of routing policies without a full deploy. At that scale you also start routing for reasons beyond cost: some models are faster, some are offline (circuit breaker), some are rate-limited, and some are contractually required for certain data residency regions. The router becomes a policy engine. Libraries like LiteLLM and frameworks like LangGraph make it easier to express routing as configuration rather than code, which matters when your ML team and platform team both need to change routing logic independently.
Key Takeaways
- Route by complexity first; reserve expensive models only for tasks that genuinely need them.
- The router itself should be cheap and fast -- a classifier, heuristics, or a small LLM call.
- Measure routing accuracy in production; a misrouted hard task costs more than skipping routing entirely.
- Combine routing with caching: identical simple requests need neither a cheap nor an expensive model call.
Pro tips
- Log every routing decision with the actual model used AND the complexity signal that triggered it. After a week in production you will have a labeled dataset you can use to train a real classifier -- the logs pay for themselves.
- Set a confidence threshold on your classifier: if the signal is ambiguous (e.g., score between 0.4 and 0.6), default to the cheaper model and mark the request for human review. Ambiguous requests are rarely the ones where quality is critical.
- Route on output type, not just input complexity. A request for a 2000-token structured JSON report needs a different model than a 2000-token creative essay, even if both inputs look 'complex' by heuristics.
- When using a small LLM as your router, pin it to a specific model version and monitor its own failure rate separately. A router that fails open (always routes to complex) silently destroys your cost model; a router that fails closed (always routes to simple) silently degrades quality.
Common pitfalls
- Mistake: Using prompt length as your only complexity signal. Fix: Combine length with semantic cues (keywords, entities, presence of multi-step instructions) or use an embedding-based classifier trained on labeled examples.
- Mistake: Routing to a weaker model without telling the user or logging the decision. Fix: Always emit a structured log with tier, model, and routing reason so you can audit quality regressions post-hoc.
- Mistake: Building the router before you have production traffic data. Fix: Deploy with a single model first, instrument everything, then design routing tiers around the actual complexity distribution you observe.
- Mistake: Treating the router as a static config file that nobody owns. Fix: Assign the routing policy to a team, version it in source control, and A/B test changes with an evaluation harness before promoting to production.
Which routing strategy to use
| Option | Use when | Avoid when |
|---|---|---|
| Rule-based heuristics (token count, keywords) | Domain is narrow and complexity signals are reliable; you need zero latency overhead and no extra API calls. | Requests are linguistically diverse or complexity correlates poorly with surface features like length. |
| Embedding + trained classifier | You have labeled examples (even a few hundred) and need better generalization than rules, with modest latency budget (~50ms). | You have no labeled data yet or the embedding model adds unacceptable latency to your p99 budget. |
| Small LLM as router (structured output) | Task complexity is nuanced and hard to capture with rules; you can afford an extra fast LLM call (100-300ms, low cost). | Your router budget exceeds the savings from routing, or you cannot tolerate the added failure mode of a second LLM call. |
| No routing (single model for all requests) | Early in development, traffic is low, or the quality difference between tiers is unacceptable for your use case. | Daily request volume is high enough that the cost difference between tiers materially affects your unit economics. |
Code Example
# openai>=1.0.0, requires OPENAI_API_KEY in env
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
def classify_complexity(prompt: str) -> str:
"""Returns 'simple' or 'complex' based on heuristics."""
word_count = len(prompt.split())
has_reasoning_cue = any(kw in prompt.lower() for kw in ["explain", "compare", "analyze", "why", "how does"])
if word_count < 20 and not has_reasoning_cue:
return "simple"
return "complex"
def route_and_complete(user_message: str) -> str:
complexity = classify_complexity(user_message)
model = "gpt-4o-mini" if complexity == "simple" else "gpt-4o"
print(f"[router] complexity={complexity} -> model={model}")
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": user_message}],
)
return response.choices[0].message.content
print(route_and_complete("What is 2 + 2?"))
print(route_and_complete("Explain the tradeoffs between transformer and SSM architectures for long-context tasks."))How this code works
This Python code demonstrates how to dynamically choose the right AI model for a task based on its complexity, a common technique in AI Workflow Orchestration. Its job is to efficiently handle user requests by routing simpler questions to a faster, cheaper model and more demanding ones to a powerful, albeit more expensive, model. The code achieves this by first determining the complexity of a user's prompt and then using that information to select between OpenAI's gpt-4o-mini and gpt-4o models before generating a response.
The core logic resides in two functions. The classify_complexity function assesses a prompt by checking its word_count and the presence of reasoning_cue keywords like "explain" or "analyze". Prompts that are both short (under 20 words) and lack these cues are labeled simple; all others are complex. The route_and_complete function then uses this complexity to decide which model to use for the client.chat.completions.create call. A subtle but crucial detail for beginners is that the client initialization explicitly relies on os.environ["OPENAI_API_KEY"]. If this environment variable isn't properly set before running the code, it will fail with an error, as the program cannot authenticate with the OpenAI API. Finally, the code prints the chosen model and the AI's generated content.
Production-grade example
Adds retries with backoff, per-call latency/token logging, typed tiers, timeout, and graceful degradation to cheaper model on rate limit.
# openai>=1.0.0, structlog>=24.0.0, tenacity>=8.0.0
import os
import time
import structlog
from enum import Enum
from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
log = structlog.get_logger()
class Tier(str, Enum):
SIMPLE = "gpt-4o-mini"
COMPLEX = "gpt-4o"
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
def classify(prompt: str) -> Tier:
word_count = len(prompt.split())
reasoning_cues = {"explain", "compare", "analyze", "why", "how does", "tradeoff", "design"}
has_cue = any(kw in prompt.lower() for kw in reasoning_cues)
return Tier.COMPLEX if (word_count > 40 or has_cue) else Tier.SIMPLE
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=1, max=30),
stop=stop_after_attempt(4),
)
def _call_model(model: str, messages: list, max_tokens: int = 1024) -> dict:
return client.chat.completions.create(
model=model,
messages=messages,
max_tokens=max_tokens,
)
def route_and_complete(user_message: str, session_id: str = "anon") -> str:
tier = classify(user_message)
logger = log.bind(session_id=session_id, tier=tier.name, model=tier.value)
t0 = time.perf_counter()
try:
response = _call_model(
model=tier.value,
messages=[{"role": "user", "content": user_message}],
)
usage = response.usage
latency_ms = round((time.perf_counter() - t0) * 1000)
logger.info(
"routing.success",
latency_ms=latency_ms,
prompt_tokens=usage.prompt_tokens,
completion_tokens=usage.completion_tokens,
)
return response.choices[0].message.content
except RateLimitError:
logger.warning("routing.rate_limit", fallback_model=Tier.SIMPLE.value)
if tier == Tier.COMPLEX:
# Graceful degradation: fall back to cheaper model on rate limit
response = _call_model(model=Tier.SIMPLE.value,
messages=[{"role": "user", "content": user_message}])
return response.choices[0].message.content
raise
except APIStatusError as exc:
logger.error("routing.api_error", status_code=exc.status_code, message=str(exc))
raise
if __name__ == "__main__":
print(route_and_complete("What time does the store close?", session_id="u-001"))
print(route_and_complete("Compare attention mechanisms in transformers vs state space models.", session_id="u-002"))How this code works
This code intelligently routes user requests to different AI models based on their complexity, a core concept in AI Workflow Orchestration. Its job is to optimize resource usage and cost by using powerful, often more expensive, models like gpt-4o only when truly necessary, while defaulting to a cheaper, faster gpt-4o-mini for simpler tasks. This balances performance and efficiency.
The system works by first using the classify function to analyze the user's prompt. It checks the word_count and for specific "reasoning cues" (like "explain" or "analyze") to assign a Tier of either SIMPLE or COMPLEX. These tiers are mapped to specific model names using the Tier Enum. The main route_and_complete function then attempts to call the appropriate model via _call_model. A crucial aspect is the @retry decorator from tenacity, which makes _call_model resilient by automatically retrying requests if a RateLimitError or APITimeoutError occurs. A subtle but powerful design choice is how the code handles RateLimitError: if a COMPLEX task hits a rate limit, the system gracefully degrades by silently falling back to the Tier.SIMPLE model, ensuring a response even if the preferred model isn't immediately available. structlog provides detailed logging throughout this process.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a routing function that classifies a list of test prompts into 'simple', 'medium', or 'complex' tiers using at least two signals (e.g., token count + keyword presence). Print which model each prompt would be sent to, and calculate the estimated cost savings if the simple tier costs $0.15/1M tokens and the complex tier costs $5/1M tokens (illustrative numbers).
# TODO: fill in the classify_tier and estimate_savings functions
TEST_PROMPTS = [
"What is 5 times 8?",
"Summarize the French Revolution in one sentence.",
"Analyze the tradeoffs between microservices and monolithic architectures for a 10-person startup.",
"Hi",
"Explain how backpropagation works and why vanishing gradients are a problem in deep networks.",
"What is the capital of France?",
]
MODEL_MAP = {
"simple": "gpt-4o-mini", # ~$0.15/1M input tokens (illustrative)
"medium": "gpt-4o-mini", # same tier, different quota bucket
"complex": "gpt-4o", # ~$5/1M input tokens (illustrative)
}
def classify_tier(prompt: str) -> str:
# TODO: use at least two signals (e.g., word count + reasoning keywords)
# Return 'simple', 'medium', or 'complex'
pass
def estimate_savings(prompts: list[str]) -> dict:
# TODO: for each prompt, count tokens (approximate as word_count * 1.3),
# compute cost at complex rate vs routed rate, return totals
pass
if __name__ == "__main__":
for p in TEST_PROMPTS:
tier = classify_tier(p)
print(f"[{tier:7s}] -> {MODEL_MAP[tier]}: {p[:60]}")
savings = estimate_savings(TEST_PROMPTS)
print("\nCost summary:", savings)Quick check
Your router adds 200ms latency and saves $0.001 per simple request. At what daily volume does the latency cost likely outweigh the savings?
A router trained only on prompt length misclassifies a short but highly technical prompt as 'simple'. What is the most practical fix?
When does graceful degradation (falling back to a cheaper model on rate limit) cause more harm than good?