Phase 1: Programming & AI Foundations

Temperature, top-p sampling & output quality

Beginner ~12 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

When an AI (like the kind that writes stories or answers questions) creates text, it's a bit like a chef deciding the next ingredient for a new recipe. For every spot in a sentence, it thinks of many different words that could go there – some super likely, others less so. But if it always picked only the most obvious word, its writing would be super boring and repetitive! To make its writing more interesting, creative, or accurate, we use two special controls: "Temperature" and "top-p sampling."

Let's stick with our baking analogy for cookies. * If you set the Temperature very low, it's like sticking exactly to the original recipe. The recipe says "sugar," so you use sugar. You'll get very predictable, reliable cookies every time – good for consistency, but maybe a bit plain. The AI will pick very common, expected words. * If you crank the Temperature up high, it's like your recipe says, "Feel free to experiment!" When it's time for sugar, you might think, "Hmm, maybe I'll try honey instead, or even maple syrup!" The AI starts to consider less common words more often. Your cookies could turn out surprisingly unique, or a bit weird if you go too wild!

Now, top-p sampling is a different kind of control. Imagine you're looking at all the possible ingredients you could put into your cookies. You know some are super popular (sugar, flour), some are less popular but still good (oatmeal), and some are just silly for cookies (muddy socks!). Top-p helps you draw a line. It says, "Okay, I'm only going to consider a certain percentage of the most likely and sensible ingredients." * If you set top-p very high (like 95%), you're saying, "Let me look at almost all the plausible ingredients, even slightly unusual ones." The AI will have a wider but still sensible choice of words. * If you set top-p very low (like 60%), you're saying, "Nope, I'm only going to look at the absolute most common and guaranteed ingredients, like sugar and plain chocolate chips." This makes the AI stick to a very narrow, very safe list of words.

As someone who helps build AI, you get to be the "master chef" for its language. You'll adjust these two controls for almost every AI job! * If you need the AI to write a factual report, like explaining how volcanoes work, use a low Temperature and a low top-p. This ensures it sticks to accurate, expected words and doesn't "hallucinate" (make up) facts. * But if you're asking the AI to brainstorm new superhero ideas or write a funny story, turn the Temperature up and maybe increase top-p a bit. This gives the AI more freedom to explore unusual words and unexpected ideas, making its writing much more imaginative. By understanding these controls, you can make the AI produce exactly the kind of text you need!

Under the hood, an LLM produces a vector of raw scores called logits -- one per token in its vocabulary (often 50,000 to 100,000 tokens). These get passed through a softmax to become probabilities. Temperature is applied before that softmax: every logit is divided by the temperature value T. When T is 0.1, logits get blown up by 10x relative to each other, so the highest-scoring token's probability approaches 1.0 and everything else approaches 0. When T is 2.0, logits are squashed toward each other, making even unlikely tokens competitive. Temperature = 1.0 is the neutral baseline -- the model samples from its raw learned distribution.

Top-p (nucleus sampling) works differently. After the softmax produces probabilities, sort all tokens in descending probability order. Walk down the list accumulating probability mass. Stop when the cumulative total first exceeds p (e.g., 0.9). Discard every token below that cutoff and renormalize the remaining ones. The model then samples from that smaller nucleus. The key insight is that the nucleus size adapts dynamically: for a highly confident prediction the nucleus might be just 2 tokens; for an uncertain prediction it might be 400. This adaptive quality is why top-p often beats a fixed top-k cutoff.

A real-world scenario: you are building a customer support bot for a SaaS product. For intent classification and structured JSON extraction ("what is the user's account ID?"), you want temperature=0.0 and top_p=1.0. Deterministic output means your downstream parser never crashes on unexpected formatting. For the free-text reply to the user explaining the resolution, temperature=0.5 and top_p=0.9 gives natural-sounding variation without the model going off-script. A senior engineer would not use a single global setting; they would configure sampling parameters per call type, potentially exposing them as per-feature config values in a settings file.

The tradeoff versus greedy decoding (always picking the argmax token) is that greedy decoding is fast and fully reproducible but suffers from repetition loops and gets stuck in local probability maxima. Beam search explores multiple candidate sequences simultaneously and selects the highest-probability complete sequence, which can improve coherence for translation tasks but is expensive and still deterministic. Sampling with temperature and top-p introduces useful stochasticity that avoids repetition while keeping costs linear in sequence length. Most production LLM APIs (OpenAI, Anthropic, Google Vertex) expose temperature and top-p because they are cheap to compute and effective in practice.

At scale, sampling parameters affect more than output quality -- they affect latency and cost indirectly. A high-temperature model tends to generate more tokens before reaching a natural stopping point (EOS token probability is also part of the distribution). If you are paying per output token and your temperature is unnecessarily high for a factual task, you are burning money on verbose, wandering completions. At 10 users this is noise; at 10 million requests per day the difference between a mean output length of 80 tokens versus 120 tokens at a hypothetical $0.60 per million tokens is material. Monitor your p95 output token counts by call site. Another scale concern: if your application needs reproducibility (audit trails, regression testing), set temperature=0.0. Note that even at temperature=0, some providers do not guarantee bit-for-bit identical outputs across infrastructure changes, so never rely on exact string matching for logic -- use semantic checks or structured output parsing instead.

Key Takeaways

  • Temperature rescales logits before softmax; lower values concentrate probability mass on top tokens.
  • Top-p sampling restricts candidates to the smallest token set covering cumulative probability p.
  • Use temperature near 0 for deterministic tasks like code or structured extraction, higher for creative tasks.
  • Temperature and top-p interact; changing both simultaneously makes debugging output quality harder.

Pro tips

  • Never set both temperature and top_p to non-default values simultaneously during debugging. Change one at a time so you know which parameter caused the output shift.
  • Temperature=0 does not mean the output is always identical across API calls. Floating-point non-determinism and load balancing across GPU replicas can produce rare token-level differences. Use structured outputs or JSON mode for true reliability.
  • For few-shot prompting evaluations, run each test case at temperature=0 first to get a stable baseline, then increase temperature only after you know the baseline quality is acceptable.
  • Top-p at 1.0 with a low temperature is often more stable than top-p at 0.8 with a moderate temperature because you preserve the full distribution shape that the model was trained on; nucleus sampling at low p can create discontinuous jumps in the renormalized distribution.

Common pitfalls

  • Mistake: Using high temperature for structured output tasks like JSON extraction. Fix: Set temperature=0.0 and use the provider's JSON mode or structured outputs feature to eliminate parse errors.
  • Mistake: Assuming temperature=0 guarantees reproducibility across deployments. Fix: Use hash-based output validation or semantic assertions in tests, never exact string equality.
  • Mistake: Setting top_p very low (e.g., 0.5) and expecting concise output. Fix: Top_p controls candidate breadth, not length; use max_tokens to constrain length, and a stop sequence to end at a natural boundary.
  • Mistake: Using the same sampling settings for every call type in an application. Fix: Define named profiles per task category and store them in config, so tuning one task does not break another.

Which sampling settings to use by task type

Option Use when Avoid when
temperature=0.0, top_p=1.0 Structured extraction, SQL generation, classification, any output that feeds a parser. Creative writing or tasks where response variation is a feature; outputs become repetitive over many calls.
temperature=0.3–0.5, top_p=0.9 Customer support replies, summarization, factual Q&A where natural phrasing variation is acceptable. Strict determinism is required or you need to maximize creative diversity.
temperature=0.8–1.0, top_p=0.95 Brainstorming, marketing copy, generating varied training examples, creative storytelling. Factual accuracy is critical; higher temperatures increase hallucination risk on knowledge-dependent tasks.
temperature=1.0, top_p=0.5–0.7 (low nucleus) You want high-temperature diversity but need to suppress rare nonsensical tokens from a large vocabulary. The nucleus is so narrow that renormalization artifacts cause unnatural word repetition; test carefully.

Code Example

python
# openai-python >= 1.0.0
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from env

def complete(prompt: str, temperature: float, top_p: float) -> str:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        temperature=temperature,
        top_p=top_p,
        max_tokens=256,
    )
    return response.choices[0].message.content

# Deterministic -- good for structured extraction
print(complete("List the capital cities of France, Germany, Italy.", temperature=0.0, top_p=1.0))

# Creative -- good for brainstorming
print(complete("Give me 5 unusual names for a coffee shop.", temperature=0.9, top_p=0.95))

How this code works

This Python code demonstrates how to interact with an LLM using the openai library, specifically showcasing how temperature and top_p parameters influence the generated text's creativity and determinism. It starts by importing OpenAI and initializing a client instance, which automatically retrieves the necessary OPENAI_API_KEY from the environment. A reusable function, complete, is defined to streamline calls to the LLM. This function takes a prompt string along with temperature and top_p floating-point values. Inside, it constructs a chat completion request to the model="gpt-4o-mini", passing the prompt as a user message and directly applying the specified temperature and top_p values. A max_tokens=256 limit is also set for the response length, and the function returns the actual text content from response.choices[0].message.content.

The code then makes two distinct calls to this complete function to illustrate the parameters in action. The first example, labeled "Deterministic," requests capital cities with temperature=0.0 and top_p=1.0. A temperature of 0.0 makes the model highly predictable, always selecting the most probable next word, ideal for factual or structured tasks. A subtle point for beginners is that when temperature is 0.0, top_p has little practical effect; the model’s choices are already constrained to the highest probability. The second example, "Creative," asks for unusual coffee shop names using temperature=0.9 and top_p=0.95. A higher temperature introduces more randomness, encouraging the model to explore less obvious word choices, while top_p further defines the set of tokens the model can choose from, leading to more diverse and creative outputs suitable for brainstorming.

Production-grade example

Adds retry with backoff, per-profile sampling config, token logging, latency tracking, and typed error handling.

python
# openai-python >= 1.0.0, tenacity >= 8.0
import logging
import os
import time
from typing import Literal

from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)

SAMPLING_PROFILES = {
    "deterministic": {"temperature": 0.0, "top_p": 1.0},
    "balanced":      {"temperature": 0.4, "top_p": 0.9},
    "creative":      {"temperature": 0.9, "top_p": 0.95},
}

@retry(
    retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
)
def chat(
    prompt: str,
    profile: Literal["deterministic", "balanced", "creative"] = "balanced",
    model: str = "gpt-4o-mini",
    max_tokens: int = 512,
) -> str:
    params = SAMPLING_PROFILES[profile]
    t0 = time.monotonic()
    try:
        response = client.chat.completions.create(
            model=model,
            messages=[{"role": "user", "content": prompt}],
            max_tokens=max_tokens,
            **params,
        )
    except APIStatusError as exc:
        logger.error("api_error status=%s body=%s", exc.status_code, exc.body)
        raise

    latency_ms = (time.monotonic() - t0) * 1000
    usage = response.usage
    logger.info(
        "llm_call model=%s profile=%s prompt_tokens=%d completion_tokens=%d latency_ms=%.1f",
        model, profile, usage.prompt_tokens, usage.completion_tokens, latency_ms,
    )
    return response.choices[0].message.content


if __name__ == "__main__":
    print(chat("Extract the ISO date from: 'Meeting on the third of March 2025'", profile="deterministic"))
    print(chat("Suggest five taglines for a hiking gear brand.", profile="creative"))

How this code works

This code's job is to demonstrate how to interact with a Large Language Model (LLM) and control its output style using temperature and top_p sampling parameters. It defines SAMPLING_PROFILES like "deterministic", "balanced", and "creative," each mapping to specific temperature and top_p values. The core chat function sends a prompt to the gpt-4o-mini model using client.chat.completions.create, applying these chosen parameters. The if __name__ == "__main__": section illustrates this by using a "deterministic" profile for precise information extraction and a "creative" profile for brainstorming.

To make LLM interactions reliable, the chat function uses the @retry decorator from the tenacity library. This automatically retries API calls if temporary issues like RateLimitError or APITimeoutError occur, ensuring the program doesn't crash from transient problems. The stop_after_attempt(4) parameter limits these retries. A subtle but important detail is the profile: Literal[...] = "balanced" parameter in the chat function's definition. This sets "balanced" as the default sampling profile, meaning if a call to chat doesn't explicitly specify a profile, it will produce a moderately varied response rather than a strictly precise or highly imaginative one. The code also logs llm_call details, including latency and token usage, for monitoring.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Write a Python script that sends the same prompt ('Write a one-sentence product description for a smart water bottle.') to gpt-4o-mini three times each at temperature=0.0 and temperature=1.0. Print all six responses and observe how consistent or varied they are. Then find the temperature value where outputs start to noticeably diverge.

python
# openai-python >= 1.0.0
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
PROMPT = "Write a one-sentence product description for a smart water bottle."

def sample(temperature: float) -> str:
    # TODO: call client.chat.completions.create with gpt-4o-mini
    # use temperature param, set max_tokens=100
    # return the message content
    pass

for temp in [0.0, 1.0]:
    print(f"\n--- temperature={temp} ---")
    for i in range(3):
        # TODO: call sample() and print the result
        pass

# TODO: experiment with temperatures between 0.0 and 1.0
# and note at what value outputs start to vary across runs

Quick check

  1. You set temperature=0.1 for a creative writing task and the model keeps generating very similar openings across 20 runs. What is the most direct cause?

  2. Top-p=0.9 means the model samples from tokens that together account for 90% of the probability mass. How does nucleus size change when the model is highly confident about the next token?

  3. You need to extract a JSON object from user input reliably. Which parameter combination is most appropriate?

Self-check: Without looking at your notes, explain what happens to the token probability distribution when you set temperature=0.2 versus temperature=1.5, and describe a concrete task where each setting would be the right choice and why.