Phase 1: Programming & AI Foundations

Chat completions, embeddings & vision endpoints

Beginner ~15 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you’re running a super cool, super-smart restaurant, but instead of making food, it helps build amazing computer programs that can do all sorts of clever things. This restaurant has different specialized chefs, each incredibly good at one specific type of job. When you want your program to do something smart, you go to the right chef, or what computer programmers call an "endpoint," to get the job done. Knowing which chef to pick is key to making your programs brilliant!

First, there’s the Chat Chef. This chef is a master conversationalist! You give them a list of what people have said back and forth – like a conversation on a slip of paper. The Chat Chef reads it all, thinks for a moment, and then comes up with the perfect next thing to say, creating brand new words that fit right in. This is the chef you go to when you want your computer program to chat with someone, help write a story, or answer questions in a friendly way, just like a smart virtual assistant.

Next up is the Organizer Chef. This chef is less about talking and more about understanding the deep meaning of things. You give them a piece of text, like a recipe card from a giant cookbook. The Organizer Chef "tastes" the recipe and then gives it a special, secret "flavor code"—a string of numbers that describes its unique taste and meaning. Recipes with similar codes taste alike, even if their ingredients aren't exactly the same. This chef helps your program quickly find similar ideas in a huge collection of texts, like finding all the dessert recipes that are similar to your favorite chocolate cake, even if they don’t all say "chocolate."

And finally, meet the Eye Chef. This chef is amazing because you can show them a picture – like a photo of a delicious pizza – and even ask a question about it, such as "What toppings are on this pizza?" The Eye Chef "looks" at the picture with super vision and uses what it sees to describe what’s happening, identify objects, or answer your question based on the image itself. It's like having a food critic who can tell you everything about a dish just by looking at a photo, without tasting it!

So, when you’re building your own smart computer programs, remember to pick the right chef for the job. If you want your program to have a conversation, you send your request to the Chat Chef. If you need it to find similar ideas among tons of text, you use the Organizer Chef. And if you need your program to understand or describe pictures, you talk to the Eye Chef. This means you can create programs that talk, organize, and even see, just by sending your requests to the right specialist.

How chat completions actually work

The chat completions endpoint is a stateless HTTP call. You POST a JSON body containing a model name and a messages array; each message has a role (system, user, or assistant) and content. The model sees all messages as a single formatted prompt -- it has no memory between separate API calls. The "conversation history" you maintain is just a list you build and pass back on every request. The response includes a choices array (usually one item unless you set n > 1), usage counts in tokens, and a finish_reason telling you whether the model stopped naturally (stop) or hit the token limit (length). Streaming changes the transport, not the semantics: you get server-sent events with delta chunks instead of one big response, but the same token budget applies.

How embeddings actually work

An embedding model runs your input text through a transformer encoder and outputs a single fixed-length float vector -- 1536 dimensions for text-embedding-3-small, 3072 for text-embedding-3-large, 768 for Anthropic's upcoming embed models, 768 for Google's text-embedding-004. The vector encodes semantic position in a learned space: texts about similar topics end up geometrically close. You measure closeness with cosine similarity or dot product. The embedding model does not generate text; it only encodes. This makes it extremely fast and cheap compared to chat models. One practical implication: you can embed a 10,000-document corpus once, store the vectors in a database like pgvector or Pinecone, and then at query time embed the user's question and find the top-k nearest documents in milliseconds. That retrieved context then goes into a chat completion prompt -- this is the RAG pattern in one sentence.

How vision endpoints work

Vision support is not a separate endpoint; it is the same chat completions endpoint with a richer content format. Instead of a string, the content field becomes an array of typed blocks: text blocks and image blocks. Image blocks accept either a URL (the model fetches it server-side) or a base64-encoded data URI (you send the bytes directly). The model processes image tokens alongside text tokens. Image token cost is significant: a 512x512 image at high detail can cost several hundred tokens, while a full 1024x1024 image at high detail can cost over a thousand tokens before you write a single word of prompt. Always benchmark token counts with the API's usage field before assuming vision is affordable at scale. Anthropic's Claude and Google's Gemini models support the same multimodal pattern with slightly different payload shapes.

A real scenario: building a product catalog search

Suppose you're building semantic search over 50,000 product descriptions for an e-commerce site. The wrong approach is sending every product to a chat completion and asking "is this similar to X?" -- that costs dollars per search. The right approach: at index time, call the embeddings endpoint once per product, store the vectors. At search time, embed the user's query (one cheap call), compute cosine similarity against stored vectors, return the top 20. If you need a natural language answer about those products, pass them as context to a chat completion. Each tool does one job well.

Tradeoffs vs alternatives

For text generation, the main alternatives to chat completions are completions (legacy, single string, avoid for new projects) and fine-tuned model endpoints (same API shape, higher per-token cost, potentially better quality on narrow tasks). For embeddings, you can run open-source models like sentence-transformers/all-MiniLM-L6-v2 locally for zero marginal cost -- sensible if you're embedding millions of documents and have GPU capacity. For vision, local alternatives like LLaVA or Moondream exist but require infrastructure and deliver lower accuracy on complex images. For production at scale, the API services win on reliability; for high-volume predictable workloads, self-hosted wins on cost.

What changes at scale

At 10 users, call the API synchronously and return the result. At 10,000 users, you hit rate limits (requests per minute and tokens per minute are separate limits), need async calls with a semaphore to bound concurrency, and caching becomes valuable -- identical or near-identical prompts should hit a cache layer, not the API. At 10 million users, you're likely batching embedding jobs through a queue, using the Batch API for chat completions (OpenAI offers ~50% cost reduction for async batch jobs that tolerate 24-hour latency), caching embeddings aggressively, and potentially running a hybrid architecture where cheap open-source embedding models handle the bulk and the paid API handles edge cases. Latency budgets also diverge: embeddings are fast (under 200ms), streaming chat completions deliver first token in 300-800ms for small models, and vision can add 1-3 seconds on large images.

Key Takeaways

  • Use chat completions for any task that needs generated text -- chat, summarization, extraction, coding.
  • Use embeddings to compare meaning between texts; never use them to generate or answer anything.
  • Vision endpoints accept image URLs or base64 blobs alongside your text prompt in the same messages array.
  • Each endpoint type has a different cost structure -- embeddings are cheap, vision input tokens are expensive.

Pro tips

  • The embeddings endpoint accepts a list of strings in one call, not just one string. Batch your inputs to reduce round-trips and HTTP overhead -- up to 2048 strings per call for OpenAI's text-embedding-3 models.
  • finish_reason='length' in a chat completion means the model ran out of tokens mid-response, not that it finished cleanly. Always check this field in production; silently returning a truncated answer to a user is a correctness bug, not just a quality issue.
  • Vision tokens are priced differently from text tokens and vary by image resolution and detail setting. Set detail='low' for classification tasks where layout and fine details don't matter -- it dramatically cuts token cost and latency with minimal accuracy loss for many use cases.
  • Embedding models have their own token limits per input (8191 tokens for text-embedding-3-small). Inputs longer than the limit are silently truncated by some client versions. Chunk long documents before embedding them, and verify the model's specific limit in the docs.

Common pitfalls

  • Mistake: Calling a chat completion to check whether two texts are similar. Fix: Embed both texts and compute cosine similarity -- it's 100x cheaper and returns in under 200ms at scale.
  • Mistake: Rebuilding conversation history from scratch on every call instead of appending. Fix: Maintain a messages list in your session state and append each turn; pass the full list each time.
  • Mistake: Sending a raw base64 image string without the data URI prefix. Fix: Format it as data:image/jpeg;base64,<base64string> in the image_url field, or use the url key for publicly reachable URLs.
  • Mistake: Assuming the order of results from the embeddings endpoint matches input order. Fix: Sort by the index field in response.data before zipping with your input list -- the API does not guarantee order.

Which endpoint type to use for common AI tasks

Option Use when Avoid when
Chat Completions You need generated text: answers, summaries, code, extraction, classification with explanation. You need to compare or search across many documents -- too slow and expensive per comparison.
Embeddings You need to find similar items, cluster text, power semantic search, or build a retrieval index. You need the model to produce text or reason about content -- embeddings only encode, they don't generate.
Vision (multimodal chat completions) Your input contains images: OCR, layout understanding, image description, visual Q&A. You only have text input -- adding an image dramatically increases token cost and latency for no gain.
Local/open-source embedding model You're embedding millions of documents, have GPU capacity, and need to minimize marginal cost. You need the highest semantic quality, have no GPU infrastructure, or are indexing fewer than ~100k documents.

Code Example

python
# openai>=1.0.0
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from env

# 1. Chat completion
chat_response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is a transformer model in one sentence?"}
    ]
)
print(chat_response.choices[0].message.content)

# 2. Embedding
embed_response = client.embeddings.create(
    model="text-embedding-3-small",
    input="Transformer models use self-attention to process sequences."
)
vector = embed_response.data[0].embedding
print(f"Embedding dimension: {len(vector)}")  # 1536 for this model

# 3. Vision (image URL in the messages array)
vision_response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What object is in this image?"},
            {"type": "image_url", "image_url": {"url": "https://upload.wikimedia.org/wikipedia/commons/thumb/4/47/PNG_transparency_demonstration_1.png/280px-PNG_transparency_demonstration_1.png"}}
        ]
    }]
)
print(vision_response.choices[0].message.content)

How this code works

This code demonstrates interacting with OpenAI's API for three core functionalities: chat completions, text embeddings, and vision. It begins by setting up the API client using from openai import OpenAI and client = OpenAI(), which conveniently reads the API key from the environment. For chat completion, client.chat.completions.create sends a user query to gpt-4o-mini, a small, fast model. The messages array defines the conversation flow with role and content, where "system" sets the AI's persona and "user" provides the prompt. The response's generated text is then extracted and printed.

Next, the code performs text embedding using client.embeddings.create. It sends a sentence to text-embedding-3-small, a model designed to convert text into a numerical vector that captures its meaning. The resulting vector represents the input text in a high-dimensional space, and its len shows the embedding's dimension. Finally, the code uses client.chat.completions.create again, but this time for vision. The subtle difference here is how messages is structured: for image understanding, content becomes a list containing both a {"type": "text"} prompt and a {"type": "image_url"} dictionary with the image link. This allows the gpt-4o-mini model to process both text and visual information to answer questions about the image.

Production-grade example

Adds retries on rate limits, timeouts, structured token/latency logging, finish_reason guards, and batch-size validation.

python
# openai>=1.0.0, tenacity>=8.0.0
import os
import time
import logging
from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, wait_exponential, stop_after_attempt, retry_if_exception_type

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)

RETRYABLE = (RateLimitError, APITimeoutError)

@retry(
    retry=retry_if_exception_type(RETRYABLE),
    wait=wait_exponential(multiplier=1, min=2, max=60),
    stop=stop_after_attempt(4),
)
def chat_with_logging(messages: list[dict], model: str = "gpt-4o-mini") -> str:
    t0 = time.monotonic()
    try:
        response = client.chat.completions.create(
            model=model,
            messages=messages,
            stream=False,
        )
    except APIStatusError as e:
        # 4xx errors (bad request, context length) are not retryable
        log.error("Non-retryable API error", extra={"status": e.status_code, "body": e.body})
        raise
    latency_ms = (time.monotonic() - t0) * 1000
    usage = response.usage
    log.info(
        "chat_completion",
        extra={
            "model": model,
            "prompt_tokens": usage.prompt_tokens,
            "completion_tokens": usage.completion_tokens,
            "latency_ms": round(latency_ms, 1),
            "finish_reason": response.choices[0].finish_reason,
        },
    )
    if response.choices[0].finish_reason == "length":
        log.warning("Response truncated by token limit -- consider increasing max_tokens or reducing input")
    return response.choices[0].message.content


def embed_with_logging(texts: list[str], model: str = "text-embedding-3-small") -> list[list[float]]:
    # Batch up to 2048 inputs per call; split if needed
    if len(texts) > 2048:
        raise ValueError(f"Batch size {len(texts)} exceeds API limit of 2048")
    t0 = time.monotonic()
    response = client.embeddings.create(model=model, input=texts)
    latency_ms = (time.monotonic() - t0) * 1000
    log.info("embeddings", extra={"model": model, "count": len(texts), "tokens": response.usage.total_tokens, "latency_ms": round(latency_ms, 1)})
    return [item.embedding for item in sorted(response.data, key=lambda x: x.index)]


if __name__ == "__main__":
    answer = chat_with_logging([{"role": "user", "content": "Name three vector databases."}])
    print(answer)
    vecs = embed_with_logging(["pgvector", "Pinecone", "Weaviate"])
    print(f"Got {len(vecs)} vectors of dim {len(vecs[0])}") 

How this code works

This code provides production-ready functions for interacting with OpenAI's Chat Completions and Embeddings APIs, designed to be robust and informative. It begins by configuring logging to capture important details about API calls, then initializes an OpenAI client using an API key from environment variables and a 30-second timeout.

The chat_with_logging function is wrapped with a @retry decorator from the tenacity library. This automatically re-attempts API calls if transient issues like RateLimitError or APITimeoutError occur, increasing the wait time between retries up to a set maximum. This ensures resilience. However, for non-retryable APIStatusError (like 4xx client errors), it logs the error and immediately raises it. The embed_with_logging function generates text embeddings and handles a subtle API constraint: it proactively checks if len(texts) > 2048 to prevent exceeding OpenAI's maximum batch size for embeddings in a single call, raising a ValueError if this limit is hit. Both functions log crucial metrics like latency_ms and token usage for monitoring.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a tiny semantic similarity checker. Embed three short sentences about programming languages, then embed a query sentence. Compute cosine similarity between the query vector and each of the three sentence vectors. Print the sentences ranked by similarity to the query. Use OpenAI's text-embedding-3-small model.

python
# openai>=1.0.0
import os
import math
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

sentences = [
    "Python is widely used for data science and machine learning.",
    "JavaScript runs in the browser and powers interactive websites.",
    "Rust guarantees memory safety without a garbage collector.",
]
query = "Which language is best for building neural networks?"

def get_embeddings(texts: list[str]) -> list[list[float]]:
    # TODO: call client.embeddings.create with model="text-embedding-3-small"
    # return a list of vectors in input order
    pass

def cosine_similarity(a: list[float], b: list[float]) -> float:
    # TODO: implement dot(a,b) / (norm(a) * norm(b))
    pass

# TODO: embed sentences and query, compute similarities, print ranked results

Quick check

  1. You need to find the 10 most relevant documents from a corpus of 100,000 for a given user query. What is the correct approach?

  2. A chat completion returns with finish_reason='length'. What does this mean for your application?

  3. How do you send an image to a vision-capable model using the OpenAI API?

Self-check: Describe in your own words when you would use an embedding versus a chat completion for a search feature, and explain what information finish_reason gives you that the response text alone does not.