Under the hood, an LLM produces a vector of raw scores called logits -- one per token in its vocabulary (often 50,000 to 100,000 tokens). These get passed through a softmax to become probabilities. Temperature is applied before that softmax: every logit is divided by the temperature value T. When T is 0.1, logits get blown up by 10x relative to each other, so the highest-scoring token's probability approaches 1.0 and everything else approaches 0. When T is 2.0, logits are squashed toward each other, making even unlikely tokens competitive. Temperature = 1.0 is the neutral baseline -- the model samples from its raw learned distribution.
Top-p (nucleus sampling) works differently. After the softmax produces probabilities, sort all tokens in descending probability order. Walk down the list accumulating probability mass. Stop when the cumulative total first exceeds p (e.g., 0.9). Discard every token below that cutoff and renormalize the remaining ones. The model then samples from that smaller nucleus. The key insight is that the nucleus size adapts dynamically: for a highly confident prediction the nucleus might be just 2 tokens; for an uncertain prediction it might be 400. This adaptive quality is why top-p often beats a fixed top-k cutoff.
A real-world scenario: you are building a customer support bot for a SaaS product. For intent classification and structured JSON extraction ("what is the user's account ID?"), you want temperature=0.0 and top_p=1.0. Deterministic output means your downstream parser never crashes on unexpected formatting. For the free-text reply to the user explaining the resolution, temperature=0.5 and top_p=0.9 gives natural-sounding variation without the model going off-script. A senior engineer would not use a single global setting; they would configure sampling parameters per call type, potentially exposing them as per-feature config values in a settings file.
The tradeoff versus greedy decoding (always picking the argmax token) is that greedy decoding is fast and fully reproducible but suffers from repetition loops and gets stuck in local probability maxima. Beam search explores multiple candidate sequences simultaneously and selects the highest-probability complete sequence, which can improve coherence for translation tasks but is expensive and still deterministic. Sampling with temperature and top-p introduces useful stochasticity that avoids repetition while keeping costs linear in sequence length. Most production LLM APIs (OpenAI, Anthropic, Google Vertex) expose temperature and top-p because they are cheap to compute and effective in practice.
At scale, sampling parameters affect more than output quality -- they affect latency and cost indirectly. A high-temperature model tends to generate more tokens before reaching a natural stopping point (EOS token probability is also part of the distribution). If you are paying per output token and your temperature is unnecessarily high for a factual task, you are burning money on verbose, wandering completions. At 10 users this is noise; at 10 million requests per day the difference between a mean output length of 80 tokens versus 120 tokens at a hypothetical $0.60 per million tokens is material. Monitor your p95 output token counts by call site. Another scale concern: if your application needs reproducibility (audit trails, regression testing), set temperature=0.0. Note that even at temperature=0, some providers do not guarantee bit-for-bit identical outputs across infrastructure changes, so never rely on exact string matching for logic -- use semantic checks or structured output parsing instead.
Key Takeaways
- Temperature rescales logits before softmax; lower values concentrate probability mass on top tokens.
- Top-p sampling restricts candidates to the smallest token set covering cumulative probability p.
- Use temperature near 0 for deterministic tasks like code or structured extraction, higher for creative tasks.
- Temperature and top-p interact; changing both simultaneously makes debugging output quality harder.
Pro tips
- Never set both temperature and top_p to non-default values simultaneously during debugging. Change one at a time so you know which parameter caused the output shift.
- Temperature=0 does not mean the output is always identical across API calls. Floating-point non-determinism and load balancing across GPU replicas can produce rare token-level differences. Use structured outputs or JSON mode for true reliability.
- For few-shot prompting evaluations, run each test case at temperature=0 first to get a stable baseline, then increase temperature only after you know the baseline quality is acceptable.
- Top-p at 1.0 with a low temperature is often more stable than top-p at 0.8 with a moderate temperature because you preserve the full distribution shape that the model was trained on; nucleus sampling at low p can create discontinuous jumps in the renormalized distribution.
Common pitfalls
- Mistake: Using high temperature for structured output tasks like JSON extraction. Fix: Set temperature=0.0 and use the provider's JSON mode or structured outputs feature to eliminate parse errors.
- Mistake: Assuming temperature=0 guarantees reproducibility across deployments. Fix: Use hash-based output validation or semantic assertions in tests, never exact string equality.
- Mistake: Setting top_p very low (e.g., 0.5) and expecting concise output. Fix: Top_p controls candidate breadth, not length; use max_tokens to constrain length, and a stop sequence to end at a natural boundary.
- Mistake: Using the same sampling settings for every call type in an application. Fix: Define named profiles per task category and store them in config, so tuning one task does not break another.
Which sampling settings to use by task type
| Option | Use when | Avoid when |
|---|---|---|
| temperature=0.0, top_p=1.0 | Structured extraction, SQL generation, classification, any output that feeds a parser. | Creative writing or tasks where response variation is a feature; outputs become repetitive over many calls. |
| temperature=0.3–0.5, top_p=0.9 | Customer support replies, summarization, factual Q&A where natural phrasing variation is acceptable. | Strict determinism is required or you need to maximize creative diversity. |
| temperature=0.8–1.0, top_p=0.95 | Brainstorming, marketing copy, generating varied training examples, creative storytelling. | Factual accuracy is critical; higher temperatures increase hallucination risk on knowledge-dependent tasks. |
| temperature=1.0, top_p=0.5–0.7 (low nucleus) | You want high-temperature diversity but need to suppress rare nonsensical tokens from a large vocabulary. | The nucleus is so narrow that renormalization artifacts cause unnatural word repetition; test carefully. |
Code Example
# openai-python >= 1.0.0
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from env
def complete(prompt: str, temperature: float, top_p: float) -> str:
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=temperature,
top_p=top_p,
max_tokens=256,
)
return response.choices[0].message.content
# Deterministic -- good for structured extraction
print(complete("List the capital cities of France, Germany, Italy.", temperature=0.0, top_p=1.0))
# Creative -- good for brainstorming
print(complete("Give me 5 unusual names for a coffee shop.", temperature=0.9, top_p=0.95))How this code works
This Python code demonstrates how to interact with an LLM using the openai library, specifically showcasing how temperature and top_p parameters influence the generated text's creativity and determinism. It starts by importing OpenAI and initializing a client instance, which automatically retrieves the necessary OPENAI_API_KEY from the environment. A reusable function, complete, is defined to streamline calls to the LLM. This function takes a prompt string along with temperature and top_p floating-point values. Inside, it constructs a chat completion request to the model="gpt-4o-mini", passing the prompt as a user message and directly applying the specified temperature and top_p values. A max_tokens=256 limit is also set for the response length, and the function returns the actual text content from response.choices[0].message.content.
The code then makes two distinct calls to this complete function to illustrate the parameters in action. The first example, labeled "Deterministic," requests capital cities with temperature=0.0 and top_p=1.0. A temperature of 0.0 makes the model highly predictable, always selecting the most probable next word, ideal for factual or structured tasks. A subtle point for beginners is that when temperature is 0.0, top_p has little practical effect; the model’s choices are already constrained to the highest probability. The second example, "Creative," asks for unusual coffee shop names using temperature=0.9 and top_p=0.95. A higher temperature introduces more randomness, encouraging the model to explore less obvious word choices, while top_p further defines the set of tokens the model can choose from, leading to more diverse and creative outputs suitable for brainstorming.
Production-grade example
Adds retry with backoff, per-profile sampling config, token logging, latency tracking, and typed error handling.
# openai-python >= 1.0.0, tenacity >= 8.0
import logging
import os
import time
from typing import Literal
from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
SAMPLING_PROFILES = {
"deterministic": {"temperature": 0.0, "top_p": 1.0},
"balanced": {"temperature": 0.4, "top_p": 0.9},
"creative": {"temperature": 0.9, "top_p": 0.95},
}
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
)
def chat(
prompt: str,
profile: Literal["deterministic", "balanced", "creative"] = "balanced",
model: str = "gpt-4o-mini",
max_tokens: int = 512,
) -> str:
params = SAMPLING_PROFILES[profile]
t0 = time.monotonic()
try:
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_tokens=max_tokens,
**params,
)
except APIStatusError as exc:
logger.error("api_error status=%s body=%s", exc.status_code, exc.body)
raise
latency_ms = (time.monotonic() - t0) * 1000
usage = response.usage
logger.info(
"llm_call model=%s profile=%s prompt_tokens=%d completion_tokens=%d latency_ms=%.1f",
model, profile, usage.prompt_tokens, usage.completion_tokens, latency_ms,
)
return response.choices[0].message.content
if __name__ == "__main__":
print(chat("Extract the ISO date from: 'Meeting on the third of March 2025'", profile="deterministic"))
print(chat("Suggest five taglines for a hiking gear brand.", profile="creative"))How this code works
This code's job is to demonstrate how to interact with a Large Language Model (LLM) and control its output style using temperature and top_p sampling parameters. It defines SAMPLING_PROFILES like "deterministic", "balanced", and "creative," each mapping to specific temperature and top_p values. The core chat function sends a prompt to the gpt-4o-mini model using client.chat.completions.create, applying these chosen parameters. The if __name__ == "__main__": section illustrates this by using a "deterministic" profile for precise information extraction and a "creative" profile for brainstorming.
To make LLM interactions reliable, the chat function uses the @retry decorator from the tenacity library. This automatically retries API calls if temporary issues like RateLimitError or APITimeoutError occur, ensuring the program doesn't crash from transient problems. The stop_after_attempt(4) parameter limits these retries. A subtle but important detail is the profile: Literal[...] = "balanced" parameter in the chat function's definition. This sets "balanced" as the default sampling profile, meaning if a call to chat doesn't explicitly specify a profile, it will produce a moderately varied response rather than a strictly precise or highly imaginative one. The code also logs llm_call details, including latency and token usage, for monitoring.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Write a Python script that sends the same prompt ('Write a one-sentence product description for a smart water bottle.') to gpt-4o-mini three times each at temperature=0.0 and temperature=1.0. Print all six responses and observe how consistent or varied they are. Then find the temperature value where outputs start to noticeably diverge.
# openai-python >= 1.0.0
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
PROMPT = "Write a one-sentence product description for a smart water bottle."
def sample(temperature: float) -> str:
# TODO: call client.chat.completions.create with gpt-4o-mini
# use temperature param, set max_tokens=100
# return the message content
pass
for temp in [0.0, 1.0]:
print(f"\n--- temperature={temp} ---")
for i in range(3):
# TODO: call sample() and print the result
pass
# TODO: experiment with temperatures between 0.0 and 1.0
# and note at what value outputs start to vary across runsQuick check
You set temperature=0.1 for a creative writing task and the model keeps generating very similar openings across 20 runs. What is the most direct cause?
Top-p=0.9 means the model samples from tokens that together account for 90% of the probability mass. How does nucleus size change when the model is highly confident about the next token?
You need to extract a JSON object from user input reliably. Which parameter combination is most appropriate?