Phase 1: Programming & AI Foundations

JSON, HTTP requests & data serialization

Beginner ~12 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're at a restaurant, and you want to order a delicious burger. You can't just shout "BURGER!" across the room and hope the chef understands exactly what you want, right? You need a clear way to tell them your order, and they need a clear way to send it back to you. In the world of computers and building smart AI programs, it's very similar! Your AI program often needs to 'order' information from other computer programs (like asking for facts) or 'send' its own results somewhere else (like giving an answer).

When your computer program wants to 'order' something, it uses something called an HTTP request. Think of HTTP as the friendly waiter who takes your order from your table (your program) to the kitchen (another computer program somewhere on the internet). Now, how do you write your order so the chef understands it perfectly? That's where JSON comes in. JSON (which stands for JavaScript Object Notation, but don't worry about that fancy name for now) is like a special, super organized way of writing down your order. Instead of just saying 'burger', you'd write something like: 'food item: burger, cooked level: medium, add cheese: yes'. It's neat, tidy, and every computer program agrees to use this same special 'menu language' so there's no confusion. When your program gets ready to send this order, it takes what it knows (like 'I want a medium burger with cheese') and turns it into that specific JSON text format – that's called serializing the data. Then, when the kitchen sends back your delicious food (the information you asked for), it's also packaged in JSON. Your program then deserializes it, turning the JSON text back into something it can easily understand and use.

So, why is all this important for building AI? Well, imagine your AI program is trying to answer tricky questions, like a super smart assistant. It might need to ask another powerful AI, like a big brainy computer from a company like OpenAI, to help it figure out a complicated answer. Your AI program would use an HTTP request to send its question to OpenAI, carefully formatted in JSON. OpenAI's supercomputer would think really hard, generate an answer, and then send it back to your AI, also neatly packed in JSON. Your AI then unpacks that JSON, understands the answer, and can tell you!

Understanding how to 'order' information with HTTP and 'speak' the JSON language is super important. It means you can build AI programs that aren't stuck on your computer, but can talk to all sorts of other smart programs and databases across the internet. So, when you build your own AI, you'll be able to grab the latest weather, search for facts, or even chat with other AI assistants, all by using these exact same 'restaurant ordering' skills!

JSON is just a text format. It supports six value types: strings, numbers, booleans, null, arrays, and objects (key-value maps). Python's json.dumps() converts a dict or list to a JSON string; json.loads() reverses it. The requests library wraps both calls for you: pass json=your_dict to a request and it calls json.dumps under the hood and sets the Content-Type: application/json header. Call .json() on the response and it calls json.loads on the body. For most day-to-day API work, you rarely touch the json module directly.

Where the raw json module still matters is edge cases: reading a large JSON file from disk with json.load(file_handle), writing structured logs where you control the exact format, or handling non-standard float values that requests doesn't expose. Also important: json.dumps does not know how to serialize custom Python objects like datetime, NumPy arrays, or Pydantic models by default. You'll need a custom encoder or call .model_dump() on a Pydantic object first. This trips up nearly everyone building their first AI pipeline that includes timestamps or embeddings in a payload.

A real-world scenario: you're building a RAG pipeline that calls an embeddings API (say, POST /v1/embeddings on OpenAI), then POSTs the resulting vectors to a vector database REST API (Qdrant, Weaviate, Pinecone all offer HTTP APIs). The flow is: serialize your text chunks into the request body, receive a JSON response containing a list of float arrays, deserialize it, reshape it if needed, then serialize again for the second API call. If you skip validation between steps, a changed field name in the embeddings response silently propagates garbage into your vector store. Pydantic models as intermediate shapes catch that instantly.

Tradeoffs: requests is synchronous and battle-tested. It blocks the calling thread for the duration of the network round-trip, which is fine for scripts or background workers. Use httpx when you need async support (covered in aidev-python-async) or HTTP/2. aiohttp is another async option but httpx has a nearly identical API to requests, lowering the learning curve. For structured output from LLM APIs specifically, libraries like the OpenAI SDK or LangChain handle serialization internally, but understanding what they're doing under the hood makes debugging much faster when responses don't match the schema you expected.

At scale, the character of these problems changes. With 10 users, a misconfigured timeout is an annoyance. At 10k concurrent users, a 30-second default timeout on an LLM call means threads pile up, memory climbs, and your service falls over. At 10M users, you're thinking about connection pools (requests.Session reuses TCP connections), response streaming so you don't buffer megabytes of JSON in memory, and circuit breakers that fail fast when a downstream API is degraded rather than queuing requests indefinitely. Structured logging of latency, status codes, and token counts at each API call is not optional at that scale; it's how you diagnose cost spikes at 2am.

Cost and latency implications specific to AI work: LLM APIs charge per token, and the JSON envelope around your prompt (system message, conversation history, tool schemas) contributes to your input token count. Verbose Pydantic schemas passed as function definitions can meaningfully inflate cost. Measure your actual payload sizes. On latency, parsing a large JSON response is rarely the bottleneck compared to network time, but if you're deserializing thousands of embedding vectors on every request, consider orjson instead of the stdlib json module. orjson is 5-10x faster on large payloads and handles datetime and NumPy arrays natively.

Key Takeaways

  • Use requests for synchronous HTTP and httpx for async; don't mix them carelessly.
  • Always validate and type-check deserialized API responses with Pydantic before your app logic touches them.
  • Set explicit timeouts on every outbound HTTP call, or one slow API will hang your entire service.
  • Log request IDs and response status codes from LLM APIs to debug rate limits and billing issues.

Pro tips

  • Use requests.Session (or httpx.Client as a context manager) for any code path that makes more than one API call. Connection reuse cuts per-request overhead and is measurable at even modest call volumes.
  • When an LLM API response doesn't match the schema you expect, print response.headers first. The x-request-id header from OpenAI lets you pull the exact request in their dashboard, which is faster than guessing from the response body alone.
  • Store raw JSON responses to disk or object storage during development before you parse them. When the upstream API changes its schema (and it will), you can replay real responses against your new parser without hitting the API again.
  • orjson.dumps and orjson.loads are drop-in replacements for the stdlib json module and handle datetime, UUID, and NumPy arrays without a custom encoder. Switch to it before you need it rather than after you're debugging a serialization crash in production.

Common pitfalls

  • Mistake: Calling requests.get(url) with no timeout. Fix: Always pass timeout=(connect_seconds, read_seconds), e.g., timeout=(5, 30), or your thread blocks indefinitely on a slow API.
  • Mistake: Assuming response.json() always succeeds. Fix: Check response.status_code and Content-Type first; a 500 error often returns HTML, causing a JSONDecodeError that hides the real problem.
  • Mistake: Logging full request payloads containing user data or API keys. Fix: Log only non-sensitive fields like model name, token counts, and status codes; scrub or omit message content by default.
  • Mistake: Passing datetime objects directly to json.dumps. Fix: Convert to ISO 8601 strings first, use orjson, or write a custom default encoder; stdlib json raises TypeError on unknown types.

When to use requests vs httpx vs aiohttp

Option Use when Avoid when
requests Synchronous scripts, CLI tools, or background workers where blocking I/O is acceptable. You need async concurrency or HTTP/2 support.
httpx You want async support with a nearly identical API to requests, or need HTTP/2. Your team is already deep in aiohttp and migration cost isn't justified.
aiohttp High-throughput async services where you need fine-grained control over the event loop and connection pools. You're new to async Python; the API is more complex and error messages are harder to interpret.
LLM SDK (openai, anthropic) Calling a specific provider's API; the SDK handles auth, retries, streaming, and schema changes for you. You need to call a generic REST API or want provider-agnostic code.

Code Example

python
# requests==2.31.0, json is stdlib
import json
import requests

# Serialize a Python dict to a JSON string and back
payload = {"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Hello"}]}
json_string = json.dumps(payload)
print("Serialized:", json_string)

parsed = json.loads(json_string)
print("Deserialized role:", parsed["messages"][0]["role"])

# Make a POST request to a public echo API
response = requests.post(
    "https://httpbin.org/post",
    json=payload,      # requests serializes the dict and sets Content-Type automatically
    timeout=10,
)

if response.status_code == 200:
    data = response.json()   # deserialization built in
    print("Echo received:", data["json"]["model"])
else:
    print("Request failed:", response.status_code)

How this code works

This code demonstrates essential techniques for AI development: preparing data for models and interacting with web services. It first uses Python's built-in json library to manage data format. A Python dictionary, like the payload for an AI model's input, is converted into a standardized JSON string using json.dumps. This process, called serialization, makes data suitable for network transmission. Conversely, json.loads performs deserialization, converting a JSON string back into a Python dictionary, allowing easy access to specific data points, such as parsed["messages"][0]["role"].

Next, the code utilizes the requests library to send this data to a web service. A requests.post call targets an "echo" API with the payload dictionary. A crucial convenience here is that when json=payload is specified, requests automatically serializes the dictionary to a JSON string and correctly sets the Content-Type header, simplifying the request. If the response.status_code is 200, indicating success, response.json() then automatically deserializes the server's JSON response back into a Python dictionary, enabling straightforward extraction of values like data["json"]["model"], effectively completing the data exchange cycle.

Production-grade example

Adds Pydantic validation, typed timeouts, exponential backoff on 429s, and structured latency logging.

python
# httpx==0.27.0, pydantic==2.7.0, tenacity==8.3.0
import logging
import os
import time
from typing import Any

import httpx
from pydantic import BaseModel, ValidationError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

log = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

API_KEY = os.environ["OPENAI_API_KEY"]  # never hardcode credentials
BASE_URL = "https://api.openai.com/v1"
TIMEOUT = httpx.Timeout(connect=5.0, read=60.0, write=10.0, pool=5.0)

class EmbeddingResponse(BaseModel):
    model: str
    data: list[dict[str, Any]]
    usage: dict[str, int]

@retry(
    retry=retry_if_exception_type((httpx.TimeoutException, httpx.HTTPStatusError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
    reraise=True,
)
def get_embeddings(texts: list[str]) -> list[list[float]]:
    payload = {"model": "text-embedding-3-small", "input": texts}
    start = time.monotonic()
    with httpx.Client(timeout=TIMEOUT) as client:
        response = client.post(
            f"{BASE_URL}/embeddings",
            json=payload,
            headers={"Authorization": f"Bearer {API_KEY}"},
        )
    latency_ms = (time.monotonic() - start) * 1000
    if response.status_code == 429:
        log.warning("rate_limited status=429 retry_after=%s", response.headers.get("Retry-After"))
        response.raise_for_status()  # triggers tenacity retry
    response.raise_for_status()
    try:
        parsed = EmbeddingResponse.model_validate(response.json())
    except ValidationError as exc:
        log.error("schema_mismatch response=%s error=%s", response.text[:200], exc)
        raise
    log.info(
        "embeddings_ok model=%s tokens=%s latency_ms=%.0f",
        parsed.model, parsed.usage.get("total_tokens"), latency_ms,
    )
    return [item["embedding"] for item in parsed.data]

How this code works

This code fetches AI model embeddings for given text inputs from OpenAI, demonstrating secure HTTP requests, robust error handling, and data validation. It sets up logging to track activity and securely loads the API_KEY from environment variables, preventing sensitive information from being hardcoded. The EmbeddingResponse pydantic BaseModel defines the expected structure of OpenAI's response, ensuring the data received conforms to a specific schema for reliable processing.

The get_embeddings function sends a POST request using httpx.Client. It’s decorated with @retry from tenacity, which automatically retries the request up to four times with an exponential backoff if httpx.TimeoutException or httpx.HTTPStatusError occurs. A subtle but important detail is that response.raise_for_status() is called twice: once to explicitly trigger a retry on certain error codes (like 429 rate limits), and again after the retry logic to ensure any remaining HTTP errors propagate. Finally, EmbeddingResponse.model_validate(response.json()) parses and strictly validates the JSON response against the defined pydantic model, catching ValidationError if the data structure doesn't match, before extracting and returning the embedding vectors.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a function called fetch_joke that calls the public API at https://official-joke-api.appspot.com/random_joke, deserializes the response into a Pydantic model with fields id, type, setup, and punchline, and returns the model. Add a timeout of 10 seconds and raise a clear error if the status code is not 200.

python
# requests==2.31.0, pydantic==2.7.0
import requests
from pydantic import BaseModel

JOKE_URL = "https://official-joke-api.appspot.com/random_joke"

# TODO: define a Pydantic model 'Joke' with fields: id (int), type (str), setup (str), punchline (str)

# TODO: implement fetch_joke() -> Joke
#   1. Make a GET request to JOKE_URL with a 10-second timeout
#   2. Raise an error if status_code != 200
#   3. Deserialize the JSON response into your Joke model and return it

def fetch_joke():
    pass

if __name__ == "__main__":
    joke = fetch_joke()
    print(f"{joke.setup}\n-- {joke.punchline}")

Quick check

  1. What happens if you pass json=my_dict to requests.post instead of data=json.dumps(my_dict)?

  2. Your script calls an LLM API and occasionally hangs for several minutes. No exception is raised. What is the most likely cause?

  3. You receive a valid 200 response from an API but response.json() raises a JSONDecodeError. What should you check first?

Self-check: Without looking at your notes, explain what requests.post(url, json=payload) does differently from requests.post(url, data=json.dumps(payload)), and describe one scenario where that difference would break a real API call.