How chat completions actually work
The chat completions endpoint is a stateless HTTP call. You POST a JSON body containing a model name and a messages array; each message has a role (system, user, or assistant) and content. The model sees all messages as a single formatted prompt -- it has no memory between separate API calls. The "conversation history" you maintain is just a list you build and pass back on every request. The response includes a choices array (usually one item unless you set n > 1), usage counts in tokens, and a finish_reason telling you whether the model stopped naturally (stop) or hit the token limit (length). Streaming changes the transport, not the semantics: you get server-sent events with delta chunks instead of one big response, but the same token budget applies.
How embeddings actually work
An embedding model runs your input text through a transformer encoder and outputs a single fixed-length float vector -- 1536 dimensions for text-embedding-3-small, 3072 for text-embedding-3-large, 768 for Anthropic's upcoming embed models, 768 for Google's text-embedding-004. The vector encodes semantic position in a learned space: texts about similar topics end up geometrically close. You measure closeness with cosine similarity or dot product. The embedding model does not generate text; it only encodes. This makes it extremely fast and cheap compared to chat models. One practical implication: you can embed a 10,000-document corpus once, store the vectors in a database like pgvector or Pinecone, and then at query time embed the user's question and find the top-k nearest documents in milliseconds. That retrieved context then goes into a chat completion prompt -- this is the RAG pattern in one sentence.
How vision endpoints work
Vision support is not a separate endpoint; it is the same chat completions endpoint with a richer content format. Instead of a string, the content field becomes an array of typed blocks: text blocks and image blocks. Image blocks accept either a URL (the model fetches it server-side) or a base64-encoded data URI (you send the bytes directly). The model processes image tokens alongside text tokens. Image token cost is significant: a 512x512 image at high detail can cost several hundred tokens, while a full 1024x1024 image at high detail can cost over a thousand tokens before you write a single word of prompt. Always benchmark token counts with the API's usage field before assuming vision is affordable at scale. Anthropic's Claude and Google's Gemini models support the same multimodal pattern with slightly different payload shapes.
A real scenario: building a product catalog search
Suppose you're building semantic search over 50,000 product descriptions for an e-commerce site. The wrong approach is sending every product to a chat completion and asking "is this similar to X?" -- that costs dollars per search. The right approach: at index time, call the embeddings endpoint once per product, store the vectors. At search time, embed the user's query (one cheap call), compute cosine similarity against stored vectors, return the top 20. If you need a natural language answer about those products, pass them as context to a chat completion. Each tool does one job well.
Tradeoffs vs alternatives
For text generation, the main alternatives to chat completions are completions (legacy, single string, avoid for new projects) and fine-tuned model endpoints (same API shape, higher per-token cost, potentially better quality on narrow tasks). For embeddings, you can run open-source models like sentence-transformers/all-MiniLM-L6-v2 locally for zero marginal cost -- sensible if you're embedding millions of documents and have GPU capacity. For vision, local alternatives like LLaVA or Moondream exist but require infrastructure and deliver lower accuracy on complex images. For production at scale, the API services win on reliability; for high-volume predictable workloads, self-hosted wins on cost.
What changes at scale
At 10 users, call the API synchronously and return the result. At 10,000 users, you hit rate limits (requests per minute and tokens per minute are separate limits), need async calls with a semaphore to bound concurrency, and caching becomes valuable -- identical or near-identical prompts should hit a cache layer, not the API. At 10 million users, you're likely batching embedding jobs through a queue, using the Batch API for chat completions (OpenAI offers ~50% cost reduction for async batch jobs that tolerate 24-hour latency), caching embeddings aggressively, and potentially running a hybrid architecture where cheap open-source embedding models handle the bulk and the paid API handles edge cases. Latency budgets also diverge: embeddings are fast (under 200ms), streaming chat completions deliver first token in 300-800ms for small models, and vision can add 1-3 seconds on large images.
Key Takeaways
- Use chat completions for any task that needs generated text -- chat, summarization, extraction, coding.
- Use embeddings to compare meaning between texts; never use them to generate or answer anything.
- Vision endpoints accept image URLs or base64 blobs alongside your text prompt in the same messages array.
- Each endpoint type has a different cost structure -- embeddings are cheap, vision input tokens are expensive.
Pro tips
- The embeddings endpoint accepts a list of strings in one call, not just one string. Batch your inputs to reduce round-trips and HTTP overhead -- up to 2048 strings per call for OpenAI's text-embedding-3 models.
- finish_reason='length' in a chat completion means the model ran out of tokens mid-response, not that it finished cleanly. Always check this field in production; silently returning a truncated answer to a user is a correctness bug, not just a quality issue.
- Vision tokens are priced differently from text tokens and vary by image resolution and detail setting. Set detail='low' for classification tasks where layout and fine details don't matter -- it dramatically cuts token cost and latency with minimal accuracy loss for many use cases.
- Embedding models have their own token limits per input (8191 tokens for text-embedding-3-small). Inputs longer than the limit are silently truncated by some client versions. Chunk long documents before embedding them, and verify the model's specific limit in the docs.
Common pitfalls
- Mistake: Calling a chat completion to check whether two texts are similar. Fix: Embed both texts and compute cosine similarity -- it's 100x cheaper and returns in under 200ms at scale.
- Mistake: Rebuilding conversation history from scratch on every call instead of appending. Fix: Maintain a messages list in your session state and append each turn; pass the full list each time.
- Mistake: Sending a raw base64 image string without the data URI prefix. Fix: Format it as
data:image/jpeg;base64,<base64string>in the image_url field, or use the url key for publicly reachable URLs. - Mistake: Assuming the order of results from the embeddings endpoint matches input order. Fix: Sort by the index field in response.data before zipping with your input list -- the API does not guarantee order.
Which endpoint type to use for common AI tasks
| Option | Use when | Avoid when |
|---|---|---|
| Chat Completions | You need generated text: answers, summaries, code, extraction, classification with explanation. | You need to compare or search across many documents -- too slow and expensive per comparison. |
| Embeddings | You need to find similar items, cluster text, power semantic search, or build a retrieval index. | You need the model to produce text or reason about content -- embeddings only encode, they don't generate. |
| Vision (multimodal chat completions) | Your input contains images: OCR, layout understanding, image description, visual Q&A. | You only have text input -- adding an image dramatically increases token cost and latency for no gain. |
| Local/open-source embedding model | You're embedding millions of documents, have GPU capacity, and need to minimize marginal cost. | You need the highest semantic quality, have no GPU infrastructure, or are indexing fewer than ~100k documents. |
Code Example
# openai>=1.0.0
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from env
# 1. Chat completion
chat_response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is a transformer model in one sentence?"}
]
)
print(chat_response.choices[0].message.content)
# 2. Embedding
embed_response = client.embeddings.create(
model="text-embedding-3-small",
input="Transformer models use self-attention to process sequences."
)
vector = embed_response.data[0].embedding
print(f"Embedding dimension: {len(vector)}") # 1536 for this model
# 3. Vision (image URL in the messages array)
vision_response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What object is in this image?"},
{"type": "image_url", "image_url": {"url": "https://upload.wikimedia.org/wikipedia/commons/thumb/4/47/PNG_transparency_demonstration_1.png/280px-PNG_transparency_demonstration_1.png"}}
]
}]
)
print(vision_response.choices[0].message.content)How this code works
This code demonstrates interacting with OpenAI's API for three core functionalities: chat completions, text embeddings, and vision. It begins by setting up the API client using from openai import OpenAI and client = OpenAI(), which conveniently reads the API key from the environment. For chat completion, client.chat.completions.create sends a user query to gpt-4o-mini, a small, fast model. The messages array defines the conversation flow with role and content, where "system" sets the AI's persona and "user" provides the prompt. The response's generated text is then extracted and printed.
Next, the code performs text embedding using client.embeddings.create. It sends a sentence to text-embedding-3-small, a model designed to convert text into a numerical vector that captures its meaning. The resulting vector represents the input text in a high-dimensional space, and its len shows the embedding's dimension. Finally, the code uses client.chat.completions.create again, but this time for vision. The subtle difference here is how messages is structured: for image understanding, content becomes a list containing both a {"type": "text"} prompt and a {"type": "image_url"} dictionary with the image link. This allows the gpt-4o-mini model to process both text and visual information to answer questions about the image.
Production-grade example
Adds retries on rate limits, timeouts, structured token/latency logging, finish_reason guards, and batch-size validation.
# openai>=1.0.0, tenacity>=8.0.0
import os
import time
import logging
from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, wait_exponential, stop_after_attempt, retry_if_exception_type
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
RETRYABLE = (RateLimitError, APITimeoutError)
@retry(
retry=retry_if_exception_type(RETRYABLE),
wait=wait_exponential(multiplier=1, min=2, max=60),
stop=stop_after_attempt(4),
)
def chat_with_logging(messages: list[dict], model: str = "gpt-4o-mini") -> str:
t0 = time.monotonic()
try:
response = client.chat.completions.create(
model=model,
messages=messages,
stream=False,
)
except APIStatusError as e:
# 4xx errors (bad request, context length) are not retryable
log.error("Non-retryable API error", extra={"status": e.status_code, "body": e.body})
raise
latency_ms = (time.monotonic() - t0) * 1000
usage = response.usage
log.info(
"chat_completion",
extra={
"model": model,
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
"latency_ms": round(latency_ms, 1),
"finish_reason": response.choices[0].finish_reason,
},
)
if response.choices[0].finish_reason == "length":
log.warning("Response truncated by token limit -- consider increasing max_tokens or reducing input")
return response.choices[0].message.content
def embed_with_logging(texts: list[str], model: str = "text-embedding-3-small") -> list[list[float]]:
# Batch up to 2048 inputs per call; split if needed
if len(texts) > 2048:
raise ValueError(f"Batch size {len(texts)} exceeds API limit of 2048")
t0 = time.monotonic()
response = client.embeddings.create(model=model, input=texts)
latency_ms = (time.monotonic() - t0) * 1000
log.info("embeddings", extra={"model": model, "count": len(texts), "tokens": response.usage.total_tokens, "latency_ms": round(latency_ms, 1)})
return [item.embedding for item in sorted(response.data, key=lambda x: x.index)]
if __name__ == "__main__":
answer = chat_with_logging([{"role": "user", "content": "Name three vector databases."}])
print(answer)
vecs = embed_with_logging(["pgvector", "Pinecone", "Weaviate"])
print(f"Got {len(vecs)} vectors of dim {len(vecs[0])}") How this code works
This code provides production-ready functions for interacting with OpenAI's Chat Completions and Embeddings APIs, designed to be robust and informative. It begins by configuring logging to capture important details about API calls, then initializes an OpenAI client using an API key from environment variables and a 30-second timeout.
The chat_with_logging function is wrapped with a @retry decorator from the tenacity library. This automatically re-attempts API calls if transient issues like RateLimitError or APITimeoutError occur, increasing the wait time between retries up to a set maximum. This ensures resilience. However, for non-retryable APIStatusError (like 4xx client errors), it logs the error and immediately raises it. The embed_with_logging function generates text embeddings and handles a subtle API constraint: it proactively checks if len(texts) > 2048 to prevent exceeding OpenAI's maximum batch size for embeddings in a single call, raising a ValueError if this limit is hit. Both functions log crucial metrics like latency_ms and token usage for monitoring.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a tiny semantic similarity checker. Embed three short sentences about programming languages, then embed a query sentence. Compute cosine similarity between the query vector and each of the three sentence vectors. Print the sentences ranked by similarity to the query. Use OpenAI's text-embedding-3-small model.
# openai>=1.0.0
import os
import math
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
sentences = [
"Python is widely used for data science and machine learning.",
"JavaScript runs in the browser and powers interactive websites.",
"Rust guarantees memory safety without a garbage collector.",
]
query = "Which language is best for building neural networks?"
def get_embeddings(texts: list[str]) -> list[list[float]]:
# TODO: call client.embeddings.create with model="text-embedding-3-small"
# return a list of vectors in input order
pass
def cosine_similarity(a: list[float], b: list[float]) -> float:
# TODO: implement dot(a,b) / (norm(a) * norm(b))
pass
# TODO: embed sentences and query, compute similarities, print ranked resultsQuick check
You need to find the 10 most relevant documents from a corpus of 100,000 for a given user query. What is the correct approach?
A chat completion returns with finish_reason='length'. What does this mean for your application?
How do you send an image to a vision-capable model using the OpenAI API?