How authentication actually works
Every AI API call is an HTTPS request. The provider identifies you via an API key, typically sent as a bearer token in the Authorization header: Authorization: Bearer sk-.... The server validates the key, resolves your account, checks your quota, and only then forwards your request to inference infrastructure. The key is stateless -- there is no session, no login flow, and no token refresh. This means the key is the password. If it leaks via a git commit, a log file, or a client-side bundle, anyone who finds it can run requests billed to you. The correct approach is always to load keys from environment variables (os.environ["OPENAI_API_KEY"]) and use a secrets manager (AWS Secrets Manager, GCP Secret Manager, HashiCorp Vault) in production. The OpenAI SDK will automatically read OPENAI_API_KEY from the environment if you do not pass it explicitly, which is a sensible default.
How rate limiting works under the hood
Providers enforce multiple limit dimensions simultaneously. OpenAI, for example, limits requests per minute (RPM), tokens per minute (TPM), and sometimes tokens per day (TPD). Each dimension has its own counter, and hitting any one of them triggers a 429. Response headers tell you where you stand: x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests. Most developers ignore these headers until they have a production incident. A smarter approach is to read x-ratelimit-remaining-requests and proactively back off when it drops near zero, rather than waiting for the 429. Anthropic and Google follow similar patterns with slightly different header names. When you get a 429, the retry-after header (when present) tells you exactly how many seconds to wait. Trust it.
Real-world scenario: a document processing pipeline
Imagine you are processing 50,000 customer support tickets through gpt-4o-mini to extract sentiment and category. A naive loop with time.sleep(0) between calls will saturate your TPM limit within seconds and spend the rest of the run retrying. A production approach: use a token-aware rate limiter (track tokens in a sliding window, block when the window is full), process tickets in batches with asyncio, and cap concurrency with a semaphore sized to your RPM limit. Libraries like tenacity handle retry logic so you do not hand-roll it. For the pipeline specifically, consider OpenAI's Batch API, which processes requests asynchronously at a 50% cost reduction but with up to 24-hour turnaround -- a legitimate engineering tradeoff when latency does not matter.
Error classification: retryable vs non-retryable
Not all errors respond to retries. Grouping them clearly:
- 400 Bad Request: your payload is malformed (invalid JSON, unknown parameter, message exceeds context window). Retrying is pointless -- fix the code.
- 401 Unauthorized / 403 Forbidden: wrong key, revoked key, or insufficient permissions. Retrying is pointless -- fix the credential.
- 422 Unprocessable Entity: semantically invalid request (e.g., system message in a position the model does not support). Fix the structure.
- 429 Too Many Requests: transient rate limit hit. Retry with backoff after retry-after seconds.
- 500 / 502 / 503: server-side errors. Retry with backoff, but cap retries at 3-5.
- Timeout (no response): network or inference timeout. Retry with a longer timeout on the first retry.
A good error handler checks the status code first and branches on these classes rather than catching a generic exception and always retrying.
Exponential backoff with jitter
Exponential backoff means: wait 1s, then 2s, then 4s, then 8s, doubling each time up to a cap (usually 60s). The problem with pure exponential backoff is that if 100 concurrent workers all hit a 429 at the same time, they all wake up at the same moment and slam the API again -- this is the thundering herd. Jitter fixes it by adding randomness: wait = base * (2 ** attempt) + random.uniform(0, 1). Now workers wake up spread across a window and the second wave of requests is naturally staggered. The tenacity library implements this with wait_random_exponential. The backoff library is another option. Both are preferable to writing retry loops by hand because they handle edge cases like thread safety and max elapsed time.
What changes at scale
At 10 users, a simple retry loop works fine. At 10,000 concurrent users, per-key rate limits become a ceiling on throughput. You solve this by provisioning multiple API keys (one per deployment or service), using provider-specific features like OpenAI's "Tier" system (higher spend unlocks higher limits), and implementing a queue with a rate-limited worker pool rather than unbounded concurrency. At 10 million daily requests, you evaluate reserved capacity (Azure OpenAI's Provisioned Throughput Units, Google Vertex AI's dedicated endpoints) which trade variable costs for predictable latency and no rate limits. You also start caring deeply about p99 latency, not just average, because tail latency at scale directly affects user experience and retry amplification.
Key Takeaways
- Store API keys in environment variables, never in source code or version control.
- Retry only on transient errors (429, 500, 503) -- never on 400 or 401.
- Add random jitter to exponential backoff to prevent thundering-herd retries.
- Log response headers like
x-ratelimit-remainingto detect limit pressure early.
Pro tips
- Read the
x-ratelimit-remaining-requestsheader on every successful response and log it. You will see your limit pressure trend long before you start hitting 429s, giving you time to add a sleep or reduce concurrency proactively. - Differentiate between context-window errors (400 with a message about token limit) and other 400s. Context errors mean you need to truncate input -- retrying or escalating to a bigger model is a valid programmatic response, not just a bug.
- OpenAI's SDK exposes
response.headersvia the raw response object. Wrap your client with a thin logging shim that captures these headers on every call so you have an audit trail during incidents -- most teams realize they need this only after their first production outage. - Never share a single API key across environments. Use separate keys for dev, staging, and production so a leaked dev key cannot access production quota or billing, and so you can rotate one without touching the others.
Common pitfalls
- Mistake: Hardcoding API keys in source code or
.envfiles committed to git. Fix: Add.envto.gitignore, usepython-dotenvlocally, and inject secrets via environment variables in CI/CD and production. - Mistake: Catching all exceptions with a bare
except Exceptionand always retrying. Fix: Classify errors by status code first -- only retry 429, 500, 502, 503, and timeouts; raise immediately on 400, 401, 403. - Mistake: Using fixed sleep intervals between retries, causing thundering-herd when many workers fail simultaneously. Fix: Add random jitter to each wait interval using
wait_random_exponentialfrom tenacity or equivalent. - Mistake: Setting no timeout on API calls, allowing requests to hang indefinitely and exhaust thread pools. Fix: Always pass an explicit
timeout(10-30s is typical) and handleAPITimeoutErroras a retryable condition.
When to use which retry strategy
| Option | Use when | Avoid when |
|---|---|---|
| No retry | Interactive user-facing requests where you want immediate feedback and will surface the error to the UI. | Background jobs or pipelines where a single transient failure should not abort the entire run. |
| Fixed-interval retry (e.g., sleep 5s, retry 3x) | Simple scripts, low concurrency, or when the provider gives an explicit retry-after value to respect. | High-concurrency services where fixed intervals cause synchronized retry waves (thundering herd). |
Exponential backoff with jitter (tenacity wait_random_exponential) | Any production service with multiple concurrent callers. Default choice for most API retry scenarios. | When you have hard real-time latency requirements and cannot afford the added delay of multiple retry rounds. |
| Queue-based retry (Celery, SQS, BullMQ) | Long-running batch jobs, webhook processors, or any case where you need durability across process restarts. | Simple request-response flows where the added infrastructure complexity is not justified. |
Code Example
# openai>=1.0.0
import os
from openai import OpenAI, RateLimitError, APIStatusError
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
try:
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "What is 2 + 2?"}],
timeout=10,
)
print(response.choices[0].message.content)
except RateLimitError:
print("Hit rate limit -- slow down and retry later.")
except APIStatusError as e:
print(f"API error {e.status_code}: {e.message}")How this code works
This code demonstrates how to send a basic request to an OpenAI AI model and handle common issues like API errors or reaching usage limits. It begins by importing required modules: os to securely fetch the OPENAI_API_KEY from environment variables, and OpenAI, RateLimitError, APIStatusError from the openai library. An OpenAI client is then initialized using this key, authenticating all subsequent requests. The core task involves client.chat.completions.create, sending a simple user message to the gpt-4o-mini model, with a timeout=10 parameter ensuring the program won't wait indefinitely for a response.
To prevent the program from crashing, the API call is enclosed in a try...except block. This structure attempts the API interaction and then gracefully catches specific problems. If the application sends too many requests too quickly, a RateLimitError is caught, prompting a message to slow down. For other general HTTP-related issues, an APIStatusError is caught, allowing the program to report the status_code and message from the API. This robust error handling is crucial for building reliable applications that interact with external services.
Production-grade example
Adds retry logic with jitter, hard timeouts, non-retryable error short-circuit, and structured token/latency logging.
# openai>=1.0.0 tenacity>=8.2.0
import os
import time
import logging
import random
from openai import OpenAI, RateLimitError, APIStatusError, APITimeoutError
from tenacity import (
retry,
wait_random_exponential,
stop_after_attempt,
retry_if_exception_type,
before_sleep_log,
)
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
timeout=15.0, # hard timeout per request
)
RETRYABLE = (RateLimitError, APITimeoutError)
@retry(
retry=retry_if_exception_type(RETRYABLE),
wait=wait_random_exponential(multiplier=1, min=2, max=60),
stop=stop_after_attempt(5),
before_sleep=before_sleep_log(logger, logging.WARNING),
)
def chat_with_logging(messages: list[dict], model: str = "gpt-4o-mini") -> str:
start = time.monotonic()
try:
response = client.chat.completions.create(
model=model,
messages=messages,
)
except APIStatusError as e:
if e.status_code in (400, 401, 403, 422):
# Non-retryable: log and re-raise immediately
logger.error("Non-retryable API error", extra={"status": e.status_code, "detail": e.message})
raise
logger.warning("Retryable server error", extra={"status": e.status_code})
raise RateLimitError(message=e.message, response=e.response, body=e.body)
latency_ms = (time.monotonic() - start) * 1000
usage = response.usage
logger.info(
"llm_call_success",
extra={
"model": model,
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
"latency_ms": round(latency_ms, 1),
},
)
return response.choices[0].message.content
if __name__ == "__main__":
result = chat_with_logging([{"role": "user", "content": "Hello"}])
print(result)How this code works
This code provides a robust way to interact with the OpenAI API, ensuring reliable communication by automatically handling common issues like network timeouts, API rate limits, and various server errors. It leverages Python's logging module to track API calls and potential problems, while the OpenAI client is initialized with an api_key for authentication and a timeout to prevent requests from hanging indefinitely.
The core reliability comes from the @retry decorator from the tenacity library, which automatically re-attempts the chat_with_logging function if specific RETRYABLE exceptions like RateLimitError or APITimeoutError occur. It uses wait_random_exponential to pause for increasingly longer, random durations before retrying, preventing overwhelming the API, and stop_after_attempt to limit total retries. Inside the function, a try...except APIStatusError block catches other API responses. Crucially, errors like bad requests (400) or authentication issues (401) are considered non-retryable and are immediately re-raised. However, other server-side errors are *caught and re-raised as RateLimitError*, a subtle design choice that allows the external tenacity decorator to correctly interpret them as retryable, ensuring more resilient API calls. Successful calls are logged with model, tokens, and latency_ms for performance monitoring.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a function safe_complete(prompt, max_retries=4) that calls the OpenAI chat completions API, retries on 429 and 500-range errors using exponential backoff with jitter, raises immediately on non-retryable errors, and prints the attempt number and wait time before each retry.
import os
import time
import random
from openai import OpenAI, RateLimitError, APIStatusError, APITimeoutError
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
def safe_complete(prompt: str, max_retries: int = 4) -> str:
attempt = 0
while attempt <= max_retries:
try:
# TODO: call client.chat.completions.create with a 10s timeout
# TODO: return the message content on success
pass
except RateLimitError:
# TODO: compute wait = 2**attempt + random jitter, print attempt info, sleep
pass
except APIStatusError as e:
# TODO: if status is 500/502/503 treat like rate limit; else raise immediately
pass
except APITimeoutError:
# TODO: treat as retryable
pass
attempt += 1
raise RuntimeError("Max retries exceeded")
if __name__ == "__main__":
print(safe_complete("What is the capital of France?"))Quick check
Your API call returns HTTP 400 with 'context_length_exceeded'. What should your code do?
Why should you add random jitter to exponential backoff instead of using pure doubling waits?
Where should you store an OpenAI API key in a production Python service?