The mental model: cooperative multitasking on one thread
Asyncio is not threading and it is not multiprocessing. There is exactly one thread, one event loop, and it runs one chunk of code at a time. The magic happens at await: when your coroutine hits await client.chat.completions.create(...), it tells the event loop "I am going to wait for the network. Go run something else." The loop picks up another coroutine that is ready to advance, runs it until it hits its own await, and so on. When the first network response arrives, the OS notifies the event loop, which resumes your original coroutine right after the await. No real parallelism happens — just very efficient scheduling of wait time. This is why async is the right tool for I/O-bound work (LLM calls, database queries, HTTP requests) but does nothing for CPU-bound work (matrix math, tokenization over huge datasets). For CPU-bound tasks you want multiprocessing or concurrent.futures.ProcessPoolExecutor.
A real-world scenario: batch document summarization
You have 50 PDFs. Each one needs to be summarized by GPT-4o. Synchronous code processes them one at a time: if each call takes 3 seconds on average, you wait 150 seconds. With asyncio.gather, all 50 calls go out nearly simultaneously. You now wait roughly 3-5 seconds (dominated by the slowest call plus the time to send 50 requests). A senior engineer would not stop there though. They would add a asyncio.Semaphore(10) to cap at 10 concurrent requests, because OpenAI's tier-1 rate limits are typically around 500 RPM but also have a tokens-per-minute ceiling. Blasting 50 requests at once could trigger 429 errors. The semaphore acts as a concurrency valve: at most 10 calls run simultaneously, the rest queue up, and you stay comfortably under rate limits while still being far faster than sequential execution.
Tradeoffs vs alternative approaches
You have three realistic options for concurrent I/O in Python: asyncio, threading.Thread / concurrent.futures.ThreadPoolExecutor, and multiprocessing. Threading works and is simpler for one-off scripts — wrap your synchronous OpenAI client in a ThreadPoolExecutor and call it done. The cost is overhead: each thread uses ~8 MB of stack by default, and you still hold the GIL during Python bytecode execution (though not during the actual network wait, which is why threading still helps). Asyncio scales further because coroutines are cheap — you can have thousands in flight with negligible memory overhead. Multiprocessing is overkill here; it exists for CPU-bound problems. For production LLM pipelines, asyncio is the standard choice: every major SDK supports it natively, frameworks like FastAPI and LangChain assume it, and it composes cleanly with streaming responses.
What changes at scale
At 10 users: a simple asyncio.gather is fine. Run everything in one event loop, no rate limiting needed.
At 10,000 concurrent requests: you need explicit rate limiting (semaphores or a token bucket), retry logic with exponential backoff for 429 and 503 responses, and timeout handling so a slow call does not stall your whole batch. You probably also want a task queue (Celery, ARQ, or Redis Streams) so work survives process restarts.
At 10 million daily requests: the async pattern still applies at the per-process level, but you are now running many processes behind a load balancer. Observability becomes critical — you need per-request latency histograms, token usage metrics, and circuit breakers so a provider outage does not cascade. Cost tracking per request is non-negotiable at this scale; a single bug that sends 10x tokens costs real money.
Cost, latency, and reliability
Async does not change what you pay per token. It changes how fast you burn through your quota and how quickly users get responses. Be deliberate: batch processing benefits from high concurrency; user-facing features often need a timeout (e.g., 8 seconds) so a hung LLM call does not freeze a UI. For streaming responses — where the model sends tokens as they are generated — the async client gives you an async iterator you can forward to the browser over a Server-Sent Events connection. That pattern turns a 5-second wait into a progressive display that starts rendering in under a second, which dramatically improves perceived performance without changing actual latency.
Key Takeaways
- Use
asyncio.gatherto run multiple LLM API calls concurrently and cut wall-clock time drastically. - Async does not mean parallel CPU work — it means overlapping I/O wait times on one thread.
- Always use the SDK's async client (
AsyncOpenAI, notOpenAI) inside async functions. - Add a semaphore to cap concurrent requests and avoid hitting rate limits at scale.
Pro tips
- Never share a synchronous
OpenAI()client across async coroutines — the underlyinghttpxsession is not thread-safe in that context. Always instantiateAsyncOpenAI()once at module level and reuse it; it manages its own connection pool designed for concurrent async use. - Use
asyncio.TaskGroup(Python 3.11+) instead ofasyncio.gatherwhen you want the entire group cancelled the moment any single task raises. Withgather, one exception does not automatically stop the others, which can waste API quota and money. - Rate limits have two axes: requests per minute and tokens per minute. A semaphore caps concurrency but does not track tokens. For token-heavy workloads, implement a token bucket that estimates prompt tokens before each call using
tiktokenand blocks if the bucket is nearly full. - When profiling async LLM code, wall-clock time is the right metric, not CPU time. Use
time.perf_counter()around eachawaitto measure actual network latency, and log it per request. This is how you spot degraded providers before users do.
Common pitfalls
- Mistake: Calling
asyncio.run()inside an already-running event loop (e.g., inside a Jupyter cell or FastAPI handler). Fix: Useawaitdirectly in async contexts, ornest_asyncioin notebooks as a temporary shim. - Mistake: Using the synchronous
OpenAIclient inside anasync deffunction. Fix: Switch toAsyncOpenAI; the sync client blocks the event loop and cancels all concurrency gains. - Mistake: Firing hundreds of concurrent requests without a semaphore and getting flooded with 429 rate-limit errors. Fix: Wrap every call in
async with semaphore:and set the limit based on your tier's RPM. - Mistake: Ignoring exceptions from
asyncio.gatherby not inspectingreturn_exceptions=Trueresults. Fix: Check each result withisinstance(r, Exception)and handle or log failures explicitly.
When to use asyncio vs threading vs multiprocessing for LLM calls
| Option | Use when | Avoid when |
|---|---|---|
| asyncio + AsyncOpenAI | You have many concurrent LLM/HTTP calls in a long-running service or batch job | You need to run CPU-heavy preprocessing (tokenization of huge corpora) alongside calls |
| ThreadPoolExecutor + sync OpenAI client | You have existing sync code and want quick concurrency without a full async refactor | You need thousands of concurrent requests; thread overhead becomes significant past ~100 threads |
| multiprocessing / ProcessPoolExecutor | Your bottleneck is CPU-bound work like tokenization, PDF parsing, or local model inference | Your bottleneck is pure network I/O; process startup overhead is wasteful here |
| Task queue (Celery, ARQ) | Work must survive process restarts, needs prioritization, or spans multiple machines | You need low-latency results in a single request-response cycle |
Code Example
# openai>=1.0.0, Python 3.11+
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI() # reads OPENAI_API_KEY from env
async def ask(prompt: str) -> str:
response = await client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
)
return response.choices[0].message.content
async def main():
prompts = [
"Summarize async in one sentence.",
"What is a coroutine?",
"Give an example of an I/O-bound task.",
]
results = await asyncio.gather(*[ask(p) for p in prompts])
for prompt, result in zip(prompts, results):
print(f"Q: {prompt}\nA: {result}\n")
asyncio.run(main())How this code works
This code demonstrates how to make multiple requests to an LLM API concurrently, meaning it sends out several questions at once rather than waiting for each answer before asking the next. This significantly speeds up programs that need to interact with external services, like LLMs, by avoiding idle waiting time.
The program initializes an AsyncOpenAI client, which automatically picks up the OPENAI_API_KEY from the environment. The ask function is defined with async def, marking it as a coroutine. Inside ask, the await client.chat.completions.create line tells Python to pause this specific ask function's execution while the LLM processes the request, allowing other tasks to run. Once the LLM responds, ask resumes and returns the answer. The main function prepares a list of prompts, then uses await asyncio.gather(*[ask(p) for p in prompts]). This is where the magic happens: asyncio.gather takes all the ask calls and runs them in parallel. A subtle but important detail is the * before the list of ask calls; it unpacks the list, providing each ask task as a separate argument to gather, which is what it expects. Finally, asyncio.run(main()) kicks off the entire asynchronous process. The code then neatly pairs and prints each question with its corresponding LLM answer.
Production-grade example
Adds semaphore rate limiting, tenacity retries, per-call latency/token logging, and graceful partial-failure handling.
# openai>=1.0.0, Python 3.11+ | pip install openai tenacity structlog
import asyncio
import os
import time
import structlog
from openai import AsyncOpenAI, APIStatusError, APITimeoutError
from tenacity import (
retry, stop_after_attempt, wait_exponential,
retry_if_exception_type, before_sleep_log,
)
import logging
log = structlog.get_logger()
client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
SEMAPHORE = asyncio.Semaphore(8) # max 8 concurrent calls
@retry(
retry=retry_if_exception_type((APIStatusError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=1, max=30),
stop=stop_after_attempt(4),
before_sleep=before_sleep_log(logging.getLogger(), logging.WARNING),
)
async def call_llm(prompt: str, request_id: str) -> dict:
async with SEMAPHORE:
start = time.perf_counter()
try:
response = await client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
max_tokens=512,
)
except APIStatusError as e:
log.error("llm_api_error", request_id=request_id, status=e.status_code, msg=str(e))
raise
except APITimeoutError:
log.error("llm_timeout", request_id=request_id)
raise
elapsed = time.perf_counter() - start
usage = response.usage
log.info(
"llm_call_success",
request_id=request_id,
latency_s=round(elapsed, 3),
prompt_tokens=usage.prompt_tokens,
completion_tokens=usage.completion_tokens,
)
return {
"request_id": request_id,
"text": response.choices[0].message.content,
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
}
async def run_batch(prompts: list[str]) -> list[dict]:
tasks = [
call_llm(prompt, request_id=str(i))
for i, prompt in enumerate(prompts)
]
results = await asyncio.gather(*tasks, return_exceptions=True)
successes = [r for r in results if isinstance(r, dict)]
failures = [r for r in results if isinstance(r, Exception)]
if failures:
log.warning("batch_partial_failure", failed=len(failures), succeeded=len(successes))
return successes
if __name__ == "__main__":
sample_prompts = [f"Define AI term #{i}" for i in range(20)]
output = asyncio.run(run_batch(sample_prompts))
total_tokens = sum(r["prompt_tokens"] + r["completion_tokens"] for r in output)
print(f"Done. {len(output)} results, {total_tokens} total tokens used.")How this code works
This code provides a robust and efficient way to make many concurrent API calls to OpenAI's LLM, gpt-4o-mini. It's designed to process a batch of prompts quickly while adhering to API rate limits and gracefully handling transient network or API errors. The core logic uses Python's asyncio for concurrency, enabling the program to initiate multiple LLM requests without waiting for each one to complete sequentially. The AsyncOpenAI client handles the actual non-blocking API communication.
The call_llm function handles individual API requests. It’s decorated with @retry from tenacity, which automatically re-attempts calls that encounter APIStatusError or APITimeoutError with an exponential backoff, preventing immediate re-failures and giving the API time to recover. Crucially, async with SEMAPHORE: limits the number of simultaneous API calls to 8, preventing the application from overwhelming the OpenAI API. The run_batch function collects all call_llm tasks and runs them concurrently using asyncio.gather. A subtle but important detail is return_exceptions=True in asyncio.gather: this ensures that even if some LLM calls fail, the program collects results from all successful calls and reports on failures, rather than halting completely.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a function batch_summarize(texts: list[str]) -> list[str] that concurrently summarizes each text using the OpenAI async client, limits concurrency to 3 simultaneous calls via a semaphore, and returns results in the same order as the input list. Test it with at least 6 short texts.
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI() # needs OPENAI_API_KEY in env
# TODO: create an asyncio.Semaphore with limit 3
async def summarize_one(text: str, sem) -> str:
# TODO: use `async with sem:` to throttle concurrency
# TODO: call client.chat.completions.create with a summarization prompt
# TODO: return the response text
pass
async def batch_summarize(texts: list[str]) -> list[str]:
# TODO: create tasks for each text, preserving order
# TODO: gather all tasks and return results
pass
if __name__ == "__main__":
sample = [
"The quick brown fox jumps over the lazy dog.",
"Python was created by Guido van Rossum in 1991.",
"Asyncio enables concurrent I/O in Python.",
"GPT-4 is a large language model by OpenAI.",
"Semaphores limit concurrent access to a resource.",
"Rate limits protect APIs from overuse.",
]
results = asyncio.run(batch_summarize(sample))
for original, summary in zip(sample, results):
print(f"Original: {original}\nSummary: {summary}\n")Quick check
Why does replacing sequential LLM calls with
asyncio.gatherreduce total wall-clock time?You add
asyncio.Semaphore(5)to a batch of 50 LLM calls. What is the primary reason to do this?A FastAPI endpoint is already running inside an async event loop. What happens if you call
asyncio.run(some_coroutine())inside it?
asyncio instead of ThreadPoolExecutor for a service that makes 50 concurrent OpenAI calls, and describe the one situation where threading would still be the better choice.