JSON is just a text format. It supports six value types: strings, numbers, booleans, null, arrays, and objects (key-value maps). Python's json.dumps() converts a dict or list to a JSON string; json.loads() reverses it. The requests library wraps both calls for you: pass json=your_dict to a request and it calls json.dumps under the hood and sets the Content-Type: application/json header. Call .json() on the response and it calls json.loads on the body. For most day-to-day API work, you rarely touch the json module directly.
Where the raw json module still matters is edge cases: reading a large JSON file from disk with json.load(file_handle), writing structured logs where you control the exact format, or handling non-standard float values that requests doesn't expose. Also important: json.dumps does not know how to serialize custom Python objects like datetime, NumPy arrays, or Pydantic models by default. You'll need a custom encoder or call .model_dump() on a Pydantic object first. This trips up nearly everyone building their first AI pipeline that includes timestamps or embeddings in a payload.
A real-world scenario: you're building a RAG pipeline that calls an embeddings API (say, POST /v1/embeddings on OpenAI), then POSTs the resulting vectors to a vector database REST API (Qdrant, Weaviate, Pinecone all offer HTTP APIs). The flow is: serialize your text chunks into the request body, receive a JSON response containing a list of float arrays, deserialize it, reshape it if needed, then serialize again for the second API call. If you skip validation between steps, a changed field name in the embeddings response silently propagates garbage into your vector store. Pydantic models as intermediate shapes catch that instantly.
Tradeoffs: requests is synchronous and battle-tested. It blocks the calling thread for the duration of the network round-trip, which is fine for scripts or background workers. Use httpx when you need async support (covered in aidev-python-async) or HTTP/2. aiohttp is another async option but httpx has a nearly identical API to requests, lowering the learning curve. For structured output from LLM APIs specifically, libraries like the OpenAI SDK or LangChain handle serialization internally, but understanding what they're doing under the hood makes debugging much faster when responses don't match the schema you expected.
At scale, the character of these problems changes. With 10 users, a misconfigured timeout is an annoyance. At 10k concurrent users, a 30-second default timeout on an LLM call means threads pile up, memory climbs, and your service falls over. At 10M users, you're thinking about connection pools (requests.Session reuses TCP connections), response streaming so you don't buffer megabytes of JSON in memory, and circuit breakers that fail fast when a downstream API is degraded rather than queuing requests indefinitely. Structured logging of latency, status codes, and token counts at each API call is not optional at that scale; it's how you diagnose cost spikes at 2am.
Cost and latency implications specific to AI work: LLM APIs charge per token, and the JSON envelope around your prompt (system message, conversation history, tool schemas) contributes to your input token count. Verbose Pydantic schemas passed as function definitions can meaningfully inflate cost. Measure your actual payload sizes. On latency, parsing a large JSON response is rarely the bottleneck compared to network time, but if you're deserializing thousands of embedding vectors on every request, consider orjson instead of the stdlib json module. orjson is 5-10x faster on large payloads and handles datetime and NumPy arrays natively.
Key Takeaways
- Use
requestsfor synchronous HTTP andhttpxfor async; don't mix them carelessly. - Always validate and type-check deserialized API responses with Pydantic before your app logic touches them.
- Set explicit timeouts on every outbound HTTP call, or one slow API will hang your entire service.
- Log request IDs and response status codes from LLM APIs to debug rate limits and billing issues.
Pro tips
- Use
requests.Session(orhttpx.Clientas a context manager) for any code path that makes more than one API call. Connection reuse cuts per-request overhead and is measurable at even modest call volumes. - When an LLM API response doesn't match the schema you expect, print
response.headersfirst. Thex-request-idheader from OpenAI lets you pull the exact request in their dashboard, which is faster than guessing from the response body alone. - Store raw JSON responses to disk or object storage during development before you parse them. When the upstream API changes its schema (and it will), you can replay real responses against your new parser without hitting the API again.
orjson.dumpsandorjson.loadsare drop-in replacements for the stdlibjsonmodule and handledatetime,UUID, and NumPy arrays without a custom encoder. Switch to it before you need it rather than after you're debugging a serialization crash in production.
Common pitfalls
- Mistake: Calling
requests.get(url)with no timeout. Fix: Always passtimeout=(connect_seconds, read_seconds), e.g.,timeout=(5, 30), or your thread blocks indefinitely on a slow API. - Mistake: Assuming
response.json()always succeeds. Fix: Checkresponse.status_codeandContent-Typefirst; a 500 error often returns HTML, causing aJSONDecodeErrorthat hides the real problem. - Mistake: Logging full request payloads containing user data or API keys. Fix: Log only non-sensitive fields like model name, token counts, and status codes; scrub or omit message content by default.
- Mistake: Passing
datetimeobjects directly tojson.dumps. Fix: Convert to ISO 8601 strings first, useorjson, or write a customdefaultencoder; stdlibjsonraisesTypeErroron unknown types.
When to use requests vs httpx vs aiohttp
| Option | Use when | Avoid when |
|---|---|---|
| requests | Synchronous scripts, CLI tools, or background workers where blocking I/O is acceptable. | You need async concurrency or HTTP/2 support. |
| httpx | You want async support with a nearly identical API to requests, or need HTTP/2. | Your team is already deep in aiohttp and migration cost isn't justified. |
| aiohttp | High-throughput async services where you need fine-grained control over the event loop and connection pools. | You're new to async Python; the API is more complex and error messages are harder to interpret. |
| LLM SDK (openai, anthropic) | Calling a specific provider's API; the SDK handles auth, retries, streaming, and schema changes for you. | You need to call a generic REST API or want provider-agnostic code. |
Code Example
# requests==2.31.0, json is stdlib
import json
import requests
# Serialize a Python dict to a JSON string and back
payload = {"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Hello"}]}
json_string = json.dumps(payload)
print("Serialized:", json_string)
parsed = json.loads(json_string)
print("Deserialized role:", parsed["messages"][0]["role"])
# Make a POST request to a public echo API
response = requests.post(
"https://httpbin.org/post",
json=payload, # requests serializes the dict and sets Content-Type automatically
timeout=10,
)
if response.status_code == 200:
data = response.json() # deserialization built in
print("Echo received:", data["json"]["model"])
else:
print("Request failed:", response.status_code)How this code works
This code demonstrates essential techniques for AI development: preparing data for models and interacting with web services. It first uses Python's built-in json library to manage data format. A Python dictionary, like the payload for an AI model's input, is converted into a standardized JSON string using json.dumps. This process, called serialization, makes data suitable for network transmission. Conversely, json.loads performs deserialization, converting a JSON string back into a Python dictionary, allowing easy access to specific data points, such as parsed["messages"][0]["role"].
Next, the code utilizes the requests library to send this data to a web service. A requests.post call targets an "echo" API with the payload dictionary. A crucial convenience here is that when json=payload is specified, requests automatically serializes the dictionary to a JSON string and correctly sets the Content-Type header, simplifying the request. If the response.status_code is 200, indicating success, response.json() then automatically deserializes the server's JSON response back into a Python dictionary, enabling straightforward extraction of values like data["json"]["model"], effectively completing the data exchange cycle.
Production-grade example
Adds Pydantic validation, typed timeouts, exponential backoff on 429s, and structured latency logging.
# httpx==0.27.0, pydantic==2.7.0, tenacity==8.3.0
import logging
import os
import time
from typing import Any
import httpx
from pydantic import BaseModel, ValidationError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
log = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
API_KEY = os.environ["OPENAI_API_KEY"] # never hardcode credentials
BASE_URL = "https://api.openai.com/v1"
TIMEOUT = httpx.Timeout(connect=5.0, read=60.0, write=10.0, pool=5.0)
class EmbeddingResponse(BaseModel):
model: str
data: list[dict[str, Any]]
usage: dict[str, int]
@retry(
retry=retry_if_exception_type((httpx.TimeoutException, httpx.HTTPStatusError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
reraise=True,
)
def get_embeddings(texts: list[str]) -> list[list[float]]:
payload = {"model": "text-embedding-3-small", "input": texts}
start = time.monotonic()
with httpx.Client(timeout=TIMEOUT) as client:
response = client.post(
f"{BASE_URL}/embeddings",
json=payload,
headers={"Authorization": f"Bearer {API_KEY}"},
)
latency_ms = (time.monotonic() - start) * 1000
if response.status_code == 429:
log.warning("rate_limited status=429 retry_after=%s", response.headers.get("Retry-After"))
response.raise_for_status() # triggers tenacity retry
response.raise_for_status()
try:
parsed = EmbeddingResponse.model_validate(response.json())
except ValidationError as exc:
log.error("schema_mismatch response=%s error=%s", response.text[:200], exc)
raise
log.info(
"embeddings_ok model=%s tokens=%s latency_ms=%.0f",
parsed.model, parsed.usage.get("total_tokens"), latency_ms,
)
return [item["embedding"] for item in parsed.data]How this code works
This code fetches AI model embeddings for given text inputs from OpenAI, demonstrating secure HTTP requests, robust error handling, and data validation. It sets up logging to track activity and securely loads the API_KEY from environment variables, preventing sensitive information from being hardcoded. The EmbeddingResponse pydantic BaseModel defines the expected structure of OpenAI's response, ensuring the data received conforms to a specific schema for reliable processing.
The get_embeddings function sends a POST request using httpx.Client. It’s decorated with @retry from tenacity, which automatically retries the request up to four times with an exponential backoff if httpx.TimeoutException or httpx.HTTPStatusError occurs. A subtle but important detail is that response.raise_for_status() is called twice: once to explicitly trigger a retry on certain error codes (like 429 rate limits), and again after the retry logic to ensure any remaining HTTP errors propagate. Finally, EmbeddingResponse.model_validate(response.json()) parses and strictly validates the JSON response against the defined pydantic model, catching ValidationError if the data structure doesn't match, before extracting and returning the embedding vectors.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a function called fetch_joke that calls the public API at https://official-joke-api.appspot.com/random_joke, deserializes the response into a Pydantic model with fields id, type, setup, and punchline, and returns the model. Add a timeout of 10 seconds and raise a clear error if the status code is not 200.
# requests==2.31.0, pydantic==2.7.0
import requests
from pydantic import BaseModel
JOKE_URL = "https://official-joke-api.appspot.com/random_joke"
# TODO: define a Pydantic model 'Joke' with fields: id (int), type (str), setup (str), punchline (str)
# TODO: implement fetch_joke() -> Joke
# 1. Make a GET request to JOKE_URL with a 10-second timeout
# 2. Raise an error if status_code != 200
# 3. Deserialize the JSON response into your Joke model and return it
def fetch_joke():
pass
if __name__ == "__main__":
joke = fetch_joke()
print(f"{joke.setup}\n-- {joke.punchline}")Quick check
What happens if you pass
json=my_dicttorequests.postinstead ofdata=json.dumps(my_dict)?Your script calls an LLM API and occasionally hangs for several minutes. No exception is raised. What is the most likely cause?
You receive a valid 200 response from an API but
response.json()raises a JSONDecodeError. What should you check first?
requests.post(url, json=payload) does differently from requests.post(url, data=json.dumps(payload)), and describe one scenario where that difference would break a real API call.