How streaming actually works under the hood
LLMs generate tokens one at a time through autoregressive sampling. Without streaming, the server waits for the model to finish all tokens, serializes the complete response, and sends one HTTP response body. With streaming enabled, the provider's inference server flushes each token (or small batch of tokens) over a persistent TCP connection as soon as it is sampled. The wire format OpenAI and most providers use is Server-Sent Events: newline-delimited data: lines over a text/event-stream content type. Each chunk is a partial JSON object with a delta.content field. A final data: [DONE] sentinel closes the sequence. Your backend's job is to open that upstream connection, iterate the chunks, and re-emit them downstream -- either forwarding the raw SSE stream or re-framing into your own protocol.
A real-world scenario: customer support chat
Imagine you are building an internal support tool. The LLM generates a 400-token answer. At a typical inference speed of roughly 40-60 tokens/second (this varies widely by model, provider load, and hardware), that is 7-10 seconds to completion. Without streaming, the support agent stares at a blank box for 7-10 seconds, assumes the page hung, and refreshes. With streaming, they see the first word in under 300ms. The response feels instant even though total latency is identical. The pro approach here is to stream tokens directly from your FastAPI or Express route via SSE, keep the chunk size at the provider's natural token granularity (do not buffer on your side), and implement a heartbeat every 15 seconds so proxies and load balancers do not close idle connections.
Tradeoffs vs alternative approaches
SSE is the right default for LLM chat because communication is fundamentally one-directional during a turn: the server talks, the client listens. SSE runs over plain HTTP/1.1, works through most corporate proxies, requires no special upgrade handshake, and is natively supported by the browser EventSource API. WebSockets give you bidirectional channels, which matters if the user can interrupt a generation mid-stream, if you want the client to push tool-call results back, or if you are building a voice interface that sends audio chunks upstream. The cost of WebSockets is connection management complexity: you need a connection registry, heartbeat pings, and reconnect logic on both sides. HTTP/2 server push is a third option but has poor CDN and proxy support and is rarely worth the complexity. Long-polling (the naive fallback) adds 1-2 round trips of latency per chunk and should never be used for new LLM apps.
What changes at scale
At 10 concurrent users, a single FastAPI process with async generators handles everything fine. At 10,000 concurrent users, each open streaming connection holds a file descriptor and a small amount of memory for its async generator state. On a typical 2-vCPU container, you saturate file descriptor limits (default 1024 on many Linux distros) long before you saturate CPU. Raise ulimit -n to 65535, use an async framework (FastAPI with uvicorn, or Node.js), and put a horizontally scalable load balancer (AWS ALB, NGINX) in front. ALB supports HTTP/2 and keeps long-lived connections open; make sure its idle timeout is set longer than your longest expected LLM response (120 seconds is a safe floor). At 10 million users, you are almost certainly routing through a CDN or edge layer. Most CDNs do not support SSE out of the box -- Cloudflare Workers support TransformStream, Fastly supports streaming, but Vercel's free tier buffers responses. Test your CDN against a slow streaming endpoint before you commit.
Cost, latency, and reliability
Time-to-first-token (TTFT) is the metric that matters for perceived responsiveness; total generation time matters for throughput billing. Log both. TTFT includes network round-trip to the provider plus the time for the model to generate token 1 (which includes KV-cache setup). If your TTFT is high, the problem is usually provider cold starts or your own middleware adding latency before the upstream call. Token cost is identical whether you stream or not -- you pay for tokens generated, not HTTP packets. Where streaming does affect cost is in partial reads: if a client disconnects mid-stream (closed tab, mobile network loss), you should detect the disconnect and cancel the upstream call to the provider. OpenAI and Anthropic both support cancellation; failing to cancel means you pay for tokens the user never received. Use request.is_disconnected() in FastAPI or the AbortController pattern in Node.js to hook into connection close events.
Key Takeaways
- Use SSE for server-to-client LLM streaming; reserve WebSockets only when you need bidirectional real-time messages.
- Set a per-stream timeout, not just a connection timeout, to catch stalled generators.
- Always propagate provider errors inside the stream, not just at connection open.
- Log token counts and time-to-first-token separately; they diagnose different production problems.
Pro tips
- Set
X-Accel-Buffering: noin your SSE response headers. NGINX buffers streaming responses by default, which completely breaks SSE for clients behind it. This header disables that behavior. - Log time-to-first-token and total-stream-duration as separate metrics. High TTFT usually means provider cold start or middleware latency. High total duration with low TTFT means the model is running slowly, which is a different problem with a different fix.
- When proxying an OpenAI stream through your own backend, do not re-parse and re-serialize each chunk unless you need to inject data. Simply forward the raw
data:lines. Parsing adds 1-3ms per chunk and multiplies across thousands of connections. - Request OpenAI's
stream_options: {include_usage: true}to get a final usage chunk with exact token counts. Without it, you have to estimate costs from the response, which is error-prone for tool-call-heavy conversations.
Common pitfalls
- Mistake: Forgetting to cancel the upstream LLM call when the client disconnects. Fix: Check
request.is_disconnected()inside the generator loop and callstream.close()to stop paying for unread tokens. - Mistake: Using a synchronous generator inside an async FastAPI route, blocking the event loop. Fix: Use
async defgenerators withasync for chunk in streamso other requests are not starved during slow completions. - Mistake: Not setting a per-stream timeout, only a connection timeout. A stalled generator holds resources indefinitely. Fix: Wrap the streaming call in
asyncio.wait_for(...)with an explicit deadline of 60-120 seconds. - Mistake: Buffering all chunks on the backend before flushing to the client. Fix: Yield each
data:line immediately; never accumulate in a list. Buffering defeats the entire purpose of streaming.
When to use SSE vs WebSockets vs buffered HTTP for LLM responses
| Option | Use when | Avoid when |
|---|---|---|
| Server-Sent Events (SSE) | Server streams tokens to client; client sends one message per turn. Works through standard HTTP proxies and CDNs. | Client needs to interrupt or send data mid-stream, or you need sub-50ms bidirectional latency (e.g., voice). |
| WebSockets | Bidirectional real-time channel needed: user can cancel generation, send audio chunks, or push tool results back during a turn. | Your infrastructure (CDN, corporate proxy, some load balancers) does not support WebSocket upgrades, or you just need one-way streaming. |
| Buffered HTTP (no streaming) | Response must be post-processed entirely before display (e.g., structured JSON that is meaningless as partial output), or client is a batch pipeline. | Response takes more than 2-3 seconds to generate; users will perceive the wait as a hang. |
| HTTP/2 with chunked streaming | You control both client and server, your infra fully supports HTTP/2, and you want multiplexed streams over one connection. | Any CDN, proxy, or enterprise network is in the path that might downgrade or buffer HTTP/2 frames. |
Code Example
# openai>=1.0.0, fastapi>=0.110.0, uvicorn
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from openai import OpenAI
import os
app = FastAPI()
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
def token_stream(user_message: str):
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": user_message}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
# SSE format: each event is "data: <payload>\n\n"
yield f"data: {delta}\n\n"
yield "data: [DONE]\n\n"
@app.get("/chat")
def chat(message: str):
return StreamingResponse(token_stream(message), media_type="text/event-stream")How this code works
This Python code creates a simple web API using FastAPI that provides real-time chat responses from OpenAI's AI model. When a client requests the /chat endpoint with a message, the API streams the AI's reply back word-by-word, similar to how a chat interface displays text as it's generated. This is achieved through Server-Sent Events (SSE), enabling a continuous flow of data from the server to the client.
The core logic resides in the token_stream generator function. It calls client.chat.completions.create with stream=True to receive the AI's response in small pieces. As each delta (a small part of the generated text) arrives, the function yields it, formatted with a data: prefix and two newline characters (\n\n). This precise data: <payload>\n\n format is a subtle but critical detail; omitting the double newline would prevent the client from recognizing separate SSE events. Finally, the @app.get("/chat") endpoint uses StreamingResponse with media_type="text/event-stream" to send this formatted stream to the client.
Production-grade example
Adds retries, disconnect detection, TTFT logging, token cost tracking, and upstream cancellation on client drop.
# openai>=1.0.0, fastapi>=0.110.0, tenacity>=8.2.0, structlog>=24.0
import asyncio
import os
import time
import structlog
from fastapi import FastAPI, Request, HTTPException
from fastapi.responses import StreamingResponse
from openai import AsyncOpenAI, APIConnectionError, APIStatusError, RateLimitError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
log = structlog.get_logger()
app = FastAPI()
client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
RETRYABLE = (APIConnectionError, RateLimitError)
@retry(
retry=retry_if_exception_type(RETRYABLE),
wait=wait_exponential(multiplier=1, min=1, max=10),
stop=stop_after_attempt(3),
reraise=True,
)
async def _open_stream(messages: list[dict]):
return await client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
stream=True,
stream_options={"include_usage": True},
)
async def token_generator(request: Request, messages: list[dict], request_id: str):
t_start = time.monotonic()
first_token = True
total_tokens = 0
try:
stream = await _open_stream(messages)
async for chunk in stream:
if await request.is_disconnected():
log.info("client_disconnected", request_id=request_id)
await stream.close() # cancel upstream; stop paying for unused tokens
return
if chunk.usage: # final chunk with usage stats
total_tokens = chunk.usage.total_tokens
continue
delta = chunk.choices[0].delta.content if chunk.choices else None
if delta:
if first_token:
ttft = time.monotonic() - t_start
log.info("time_to_first_token", request_id=request_id, ttft_ms=round(ttft * 1000))
first_token = False
yield f"data: {delta}\n\n"
log.info("stream_complete", request_id=request_id,
total_tokens=total_tokens, duration_ms=round((time.monotonic() - t_start) * 1000))
yield "data: [DONE]\n\n"
except APIStatusError as exc:
log.error("provider_error", request_id=request_id, status=exc.status_code, body=exc.message)
yield f"data: {{\"error\": \"{exc.message}\"}}\n\n"
except asyncio.TimeoutError:
log.error("stream_timeout", request_id=request_id)
yield "data: {\"error\": \"upstream timeout\"}\n\n"
@app.post("/v1/chat/stream")
async def chat_stream(request: Request, body: dict):
messages = body.get("messages")
if not messages:
raise HTTPException(status_code=422, detail="messages required")
request_id = request.headers.get("x-request-id", "unknown")
headers = {"X-Accel-Buffering": "no", "Cache-Control": "no-cache"}
return StreamingResponse(
token_generator(request, messages, request_id),
media_type="text/event-stream",
headers=headers,
)How this code works
This code establishes a FastAPI endpoint, /v1/chat/stream, designed for real-time AI chat experiences. Its job is to accept user messages and stream back the AI's response one token at a time, providing a smooth, responsive interaction.
The process begins when chat_stream receives a request and passes messages to token_generator. This async generator first calls _open_stream, which robustly interacts with OpenAI using @retry from tenacity to automatically handle transient APIConnectionError or RateLimitError by retrying requests. As OpenAI sends back chunks, token_generator extracts the delta (individual tokens), measures time_to_first_token, and formats them as data: {delta}\n\n for a text/event-stream response. A subtle but crucial detail for production is handling request.is_disconnected(). If the client closes the connection mid-stream, await stream.close() is explicitly called to cancel the upstream OpenAI request, preventing unnecessary token usage and cost. The StreamingResponse then delivers these token chunks to the client in real time.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a FastAPI endpoint that streams tokens from gpt-4o-mini using SSE. The endpoint should accept a JSON body with a message field, stream each token as a separate data: event, emit data: [DONE] at the end, and print the total number of tokens yielded to stdout when the stream completes.
# openai>=1.0.0, fastapi>=0.110.0, uvicorn
import os
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
from openai import AsyncOpenAI
app = FastAPI()
client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"])
async def stream_tokens(message: str):
token_count = 0
# TODO: call client.chat.completions.create with stream=True
# TODO: iterate chunks, extract delta.content, yield SSE-formatted lines
# TODO: increment token_count for each non-empty delta
# TODO: after loop, print total token_count
# TODO: yield the [DONE] sentinel
pass
@app.post("/stream")
async def stream_endpoint(request: Request, body: dict):
message = body.get("message", "")
# TODO: return a StreamingResponse with media_type="text/event-stream"
passQuick check
A client disconnects mid-stream. What happens to the upstream LLM API call if you take no action?
Why is SSE preferred over WebSockets for a standard LLM chat turn?
You notice users see a blank chat bubble for 3 seconds before any text appears, even though total response time is only 5 seconds. Which metric is the real problem?