Few-shot prompting works because transformer-based language models perform a form of in-context learning. When the model processes your prompt, the attention mechanism lets later tokens attend to the example pairs you provided earlier. The model effectively uses those examples as a soft constraint on its probability distribution for the next token. It is not updating weights. It is reading a short specification at inference time and biasing generation toward outputs that are consistent with the demonstrated pattern. This is why example formatting matters so much: inconsistent label placement, mixed delimiters, or varied structure create conflicting signals the model has to resolve, and it sometimes resolves them wrong.
Consider a real scenario: you are building a customer support triage system that classifies incoming tickets into four queues (billing, technical, shipping, other) and extracts the urgency level (low/medium/high). Zero-shot gets maybe 80% accuracy on clean text, but it breaks on short ambiguous messages like "still waiting" or on non-native English. A fine-tuned classifier would be ideal, but you have 400 examples, not 40,000, and your team does not want to own a training pipeline. Few-shot prompting with 3-4 labeled examples per class (12-16 total) typically gets you to 90-93% without touching model weights. The right approach is to pick examples that cover edge cases -- tickets that could belong to two queues, tickets with implicit urgency -- not just the clean textbook cases.
The tradeoffs versus alternatives are real and worth internalizing. Against zero-shot: few-shot costs 200-600 extra tokens per call and adds 10-30ms of latency at typical rates, but buys you 5-15 percentage points of accuracy on ambiguous tasks. Against fine-tuning: few-shot needs no training data pipeline, no GPU time, no model hosting, and adapts instantly when your task definition changes -- but it consumes context window on every call and will never match a properly fine-tuned model on a well-defined, high-volume task. Against retrieval-augmented few-shot (dynamically selecting examples from a database): static few-shot is simpler and cheaper, while dynamic selection adds latency and an embedding cost but lets you handle a wider input distribution. Choose static few-shot first; only add retrieval when you can measure it helping.
At 10 users, none of this matters operationally. At 10,000 users, the token overhead becomes visible in your billing dashboard. If your few-shot block is 800 tokens and your average user message is 50 tokens, you are spending 94% of your input tokens on examples. At that scale, audit whether fine-tuning or a smaller model with better prompting would be cheaper. At 10 million users, the math almost always points toward fine-tuning for the stable, high-volume classification path and reserving few-shot for the long tail of edge cases. Cost per call compounds fast: 800 extra tokens times 10M calls per month at $0.15 per million input tokens is $1,200 per month in prompt overhead alone -- before you touch output tokens or latency SLAs.
A few construction rules that actually move the needle: first, keep your examples in the same format the model will see at inference time, including realistic noise like typos or short messages if those exist in production. Second, put your strongest, clearest examples first. Research suggests models weight earlier examples more heavily in long few-shot blocks, though this varies by model. Third, use a consistent and unambiguous delimiter between input and output -- a newline with a labeled prefix ("Sentiment:") is more robust than custom XML tags, which some models treat as markup to be continued rather than as a structure to be followed. Fourth, set temperature to 0 when you want deterministic, format-constrained output from few-shot prompts. You are doing classification, not creative generation.
Key Takeaways
- Embed 2-5 labeled input-output pairs in your prompt to demonstrate the exact behavior you want.
- Example quality beats example quantity -- diverse, representative, and consistently formatted examples outperform more examples.
- Few-shot costs tokens, so measure whether accuracy gains justify the added latency and spend.
- Order and label formatting both affect output quality -- experiment deliberately, not randomly.
Pro tips
- Test your few-shot examples against a held-out sample of real production inputs before shipping. Examples built from imagined scenarios often miss the long tail of real user input, and you won't know until you measure.
- When the model keeps producing almost-right output (correct class, wrong casing or punctuation), the fix is almost always in your example formatting, not in adding more examples. Make every example letter-perfect for the format you expect.
- For multi-label or multi-field extraction tasks, show at least one example where a field is absent or null. If you only show examples where every field is populated, the model will hallucinate values rather than return null.
- Dynamic few-shot selection -- embedding your example bank and retrieving the k-nearest examples for each input at runtime -- consistently outperforms a static set on high-variance inputs. It adds one embedding call latency, but the accuracy lift is often worth it for production systems.
Common pitfalls
- Mistake: Using imagined, idealized examples that don't match real input distribution. Fix: Pull examples from actual production or staging data, including messy and edge-case inputs.
- Mistake: Inconsistent formatting across examples (mixed delimiters, label positions, whitespace). Fix: Template your examples programmatically so every one follows the exact same structure.
- Mistake: Adding more examples to fix a flawed pattern instead of fixing the examples themselves. Fix: Diagnose which input type is failing, then fix or replace the relevant example rather than appending more.
- Mistake: Ignoring token cost of few-shot blocks at scale. Fix: Log prompt_tokens per call and model the monthly cost. Revisit fine-tuning if few-shot overhead exceeds a meaningful share of total spend.
When to use few-shot vs alternatives
| Option | Use when | Avoid when |
|---|---|---|
| Zero-shot | The task is well-defined, the model handles it reliably, and token budget is tight. | Output format is strict or the task involves domain-specific labeling the model has not seen. |
| Static few-shot | Input distribution is fairly uniform, you have 3-8 strong examples, and simplicity matters. | Inputs vary widely enough that no single static example set covers the real distribution. |
| Dynamic few-shot (retrieval) | You have a large example bank and input variance is high; accuracy is more important than latency. | You cannot afford an extra embedding call or do not have enough labeled examples to build a bank. |
| Fine-tuning | Task is stable, volume is high (millions of calls), and you have hundreds of labeled examples. | Task definition changes frequently or you need to ship this week without a training pipeline. |
Code Example
# openai>=1.0.0
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from environment
FEW_SHOT_PROMPT = """Classify the sentiment of each review as POSITIVE, NEGATIVE, or NEUTRAL.
Review: "The noise-cancelling is incredible, battery lasts forever."
Sentiment: POSITIVE
Review: "Broke after two weeks. Customer service never replied."
Sentiment: NEGATIVE
Review: "It's fine. Does what it says on the box."
Sentiment: NEUTRAL
Review: "{review}"
Sentiment:"""
def classify_review(review: str) -> str:
prompt = FEW_SHOT_PROMPT.format(review=review)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
max_tokens=10,
temperature=0,
)
return response.choices[0].message.content.strip()
print(classify_review("Sounds great but the fit is terrible."))How this code works
This code classifies product review sentiment (POSITIVE, NEGATIVE, NEUTRAL) using a "few-shot prompting" technique with an AI model. It starts by setting up the OpenAI client with client = OpenAI(), which conveniently reads the necessary API key from the environment without needing explicit configuration in the code itself. The core of this approach lies in the FEW_SHOT_PROMPT string. This multiline text provides the AI with several examples of reviews and their corresponding sentiments, effectively "teaching" the model the desired output format and classification style for new, unseen reviews.
The classify_review function takes a new review, inserts it into the FEW_SHOT_PROMPT template using .format(review=review), and sends this complete prompt to the AI. The client.chat.completions.create method handles the interaction with the specified model="gpt-4o-mini". A crucial detail is max_tokens=10, which limits the AI's response length. This ensures the model provides only the single sentiment word (like "POSITIVE") and doesn't generate additional text. Setting temperature=0 also makes the AI's classification very consistent and deterministic. Finally, the code extracts the sentiment from the AI's response using response.choices[0].message.content and prints the result.
Production-grade example
Adds retries with backoff, timeouts, token logging, unexpected-label validation, and graceful error return.
# openai>=1.0.0, tenacity>=8.0
import logging
import os
import time
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
from openai import OpenAI, RateLimitError, APITimeoutError, APIConnectionError
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=10.0)
EXAMPLES = [
("Broke after two weeks. No refund offered.", "NEGATIVE"),
("Fast shipping, exactly as described.", "POSITIVE"),
("Average. Nothing special.", "NEUTRAL"),
]
def build_prompt(examples: list[tuple[str, str]], review: str) -> str:
lines = ["Classify the sentiment as POSITIVE, NEGATIVE, or NEUTRAL.\n"]
for text, label in examples:
lines.append(f'Review: "{text}"')
lines.append(f"Sentiment: {label}\n")
lines.append(f'Review: "{review}"')
lines.append("Sentiment:")
return "\n".join(lines)
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError, APIConnectionError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
)
def classify_review(review: str, examples: list[tuple[str, str]] = EXAMPLES) -> dict:
prompt = build_prompt(examples, review)
start = time.monotonic()
try:
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
max_tokens=10,
temperature=0,
)
except Exception as exc:
log.error("LLM call failed", extra={"error": str(exc), "review_snippet": review[:60]})
return {"sentiment": "UNKNOWN", "error": str(exc), "latency_ms": None}
latency_ms = round((time.monotonic() - start) * 1000)
usage = response.usage
label = response.choices[0].message.content.strip().upper()
valid = {"POSITIVE", "NEGATIVE", "NEUTRAL"}
if label not in valid:
log.warning("Unexpected label", extra={"raw": label, "review_snippet": review[:60]})
label = "UNKNOWN"
log.info(
"classify_review",
extra={
"label": label,
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
"latency_ms": latency_ms,
},
)
return {"sentiment": label, "prompt_tokens": usage.prompt_tokens, "latency_ms": latency_ms}
if __name__ == "__main__":
result = classify_review("Sounds great but the fit is terrible.")
print(result)How this code works
This Python code classifies the sentiment of product reviews as "POSITIVE", "NEGATIVE", or "NEUTRAL" using a few-shot prompting technique with a large language model. It begins by defining EXAMPLES – a list of existing reviews and their correct sentiments. The build_prompt function takes these examples and a new review, formatting them into a clear, instructional message for the AI model. The core classify_review function then sends this prepared prompt to an OpenAI model like gpt-4o-mini, which learns the desired classification pattern from the provided demonstrations before processing the new review.
For robust communication with the AI, the classify_review function is decorated with @retry. This subtle but crucial feature automatically re-attempts the AI call multiple times if temporary issues like RateLimitError or network problems occur, making the system more resilient. Inside, client.chat.completions.create sends the structured prompt, specifying max_tokens=10 to encourage a concise sentiment label and temperature=0 for consistent responses. After receiving the AI's output, the code validates that the extracted sentiment label is one of the expected values. If the AI provides anything unexpected, the code gracefully defaults the sentiment to "UNKNOWN" to handle edge cases and logs a warning.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a few-shot prompt that extracts structured data from a raw product review: the product name, a sentiment label (POSITIVE/NEGATIVE/NEUTRAL), and a one-sentence summary. Test it on at least three different reviews and verify the output is consistently formatted as JSON. Pay attention to what happens when information is missing.
# openai>=1.0.0
from openai import OpenAI
import json
client = OpenAI() # set OPENAI_API_KEY in your environment
# TODO: write 2-3 input-output example pairs that show the model
# how to produce JSON with keys: product_name, sentiment, summary.
# Use triple-quoted strings for readability.
FEW_SHOT_EXAMPLES = """
# your examples here
"""
def extract_review_data(review_text: str) -> dict:
# TODO: build the full prompt by combining FEW_SHOT_EXAMPLES
# with the new review_text, then call the API.
prompt = ""
# TODO: call client.chat.completions.create with temperature=0
# and parse the response as JSON
raw_output = ""
# TODO: return a Python dict (use json.loads)
return {}
# Test inputs
test_reviews = [
"The AirPods Pro fit perfectly and the sound quality is outstanding.",
"Keyboard stopped working after a month. Very disappointed.",
"It arrived on time.",
]
for review in test_reviews:
result = extract_review_data(review)
print(json.dumps(result, indent=2))Quick check
You add a 5th example to fix a recurring misclassification but accuracy gets worse. What is the most likely explanation?
Your few-shot classifier runs at 10M calls per month. Each call adds 600 tokens of examples. What is the main operational question this raises?
Why does setting temperature=0 matter specifically for few-shot classification prompts?