Phase 2: Prompt Engineering & LLM Patterns

Few-shot prompting with examples

Intermediate ~14 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're trying to teach a new puppy a trick, like "sit." If you just stand there and say "sit" over and over, the puppy might look at you blankly, or lie down, or bark. It's not because the puppy isn't smart, but because the word "sit" alone isn't quite enough for it to understand exactly what you mean in that moment. Computers, even really smart AI, can sometimes be a bit like that puppy – they understand words, but struggle with very specific details or styles you want.

This is where a clever trick comes in, a bit like how a good puppy trainer works. Instead of just telling the puppy "sit," you gently guide its bottom down, say "sit," and then give it a treat. You do this a few times: (guide, say "sit!") -> treat. The puppy isn't just listening to the word anymore; it's learning from your examples. It sees what "sit" looks like and what the reward is, connecting the action to the word. After a few tries, the puppy gets it!

In the world of computers, we do something similar. When you want an AI to do something very particular – maybe always answer questions in a certain style, or organize a list in a precise way – just giving it general instructions might not be enough. So, you give the computer a few perfect examples of what you want. For instance, if you want it to list fruits separated by commas, you could show it: "Input: List fruits. Output: apples, bananas, oranges." Then you show it another one: "Input: List vegetables. Output: carrots, peas, broccoli." By seeing these examples, the computer learns the exact pattern and style you prefer, much faster and more accurately than if you just told it "list items nicely."

This means when you're building cool things with AI, and you need it to behave in a super specific way – especially for tricky tasks where the "right" answer is easier to show than describe – you can give it these little demonstration examples. It's like you're gently guiding the computer, just like a puppy trainer, to get exactly the kind of response you're looking for, making your AI tools much more precise and helpful.

Few-shot prompting works because transformer-based language models perform a form of in-context learning. When the model processes your prompt, the attention mechanism lets later tokens attend to the example pairs you provided earlier. The model effectively uses those examples as a soft constraint on its probability distribution for the next token. It is not updating weights. It is reading a short specification at inference time and biasing generation toward outputs that are consistent with the demonstrated pattern. This is why example formatting matters so much: inconsistent label placement, mixed delimiters, or varied structure create conflicting signals the model has to resolve, and it sometimes resolves them wrong.

Consider a real scenario: you are building a customer support triage system that classifies incoming tickets into four queues (billing, technical, shipping, other) and extracts the urgency level (low/medium/high). Zero-shot gets maybe 80% accuracy on clean text, but it breaks on short ambiguous messages like "still waiting" or on non-native English. A fine-tuned classifier would be ideal, but you have 400 examples, not 40,000, and your team does not want to own a training pipeline. Few-shot prompting with 3-4 labeled examples per class (12-16 total) typically gets you to 90-93% without touching model weights. The right approach is to pick examples that cover edge cases -- tickets that could belong to two queues, tickets with implicit urgency -- not just the clean textbook cases.

The tradeoffs versus alternatives are real and worth internalizing. Against zero-shot: few-shot costs 200-600 extra tokens per call and adds 10-30ms of latency at typical rates, but buys you 5-15 percentage points of accuracy on ambiguous tasks. Against fine-tuning: few-shot needs no training data pipeline, no GPU time, no model hosting, and adapts instantly when your task definition changes -- but it consumes context window on every call and will never match a properly fine-tuned model on a well-defined, high-volume task. Against retrieval-augmented few-shot (dynamically selecting examples from a database): static few-shot is simpler and cheaper, while dynamic selection adds latency and an embedding cost but lets you handle a wider input distribution. Choose static few-shot first; only add retrieval when you can measure it helping.

At 10 users, none of this matters operationally. At 10,000 users, the token overhead becomes visible in your billing dashboard. If your few-shot block is 800 tokens and your average user message is 50 tokens, you are spending 94% of your input tokens on examples. At that scale, audit whether fine-tuning or a smaller model with better prompting would be cheaper. At 10 million users, the math almost always points toward fine-tuning for the stable, high-volume classification path and reserving few-shot for the long tail of edge cases. Cost per call compounds fast: 800 extra tokens times 10M calls per month at $0.15 per million input tokens is $1,200 per month in prompt overhead alone -- before you touch output tokens or latency SLAs.

A few construction rules that actually move the needle: first, keep your examples in the same format the model will see at inference time, including realistic noise like typos or short messages if those exist in production. Second, put your strongest, clearest examples first. Research suggests models weight earlier examples more heavily in long few-shot blocks, though this varies by model. Third, use a consistent and unambiguous delimiter between input and output -- a newline with a labeled prefix ("Sentiment:") is more robust than custom XML tags, which some models treat as markup to be continued rather than as a structure to be followed. Fourth, set temperature to 0 when you want deterministic, format-constrained output from few-shot prompts. You are doing classification, not creative generation.

Key Takeaways

  • Embed 2-5 labeled input-output pairs in your prompt to demonstrate the exact behavior you want.
  • Example quality beats example quantity -- diverse, representative, and consistently formatted examples outperform more examples.
  • Few-shot costs tokens, so measure whether accuracy gains justify the added latency and spend.
  • Order and label formatting both affect output quality -- experiment deliberately, not randomly.

Pro tips

  • Test your few-shot examples against a held-out sample of real production inputs before shipping. Examples built from imagined scenarios often miss the long tail of real user input, and you won't know until you measure.
  • When the model keeps producing almost-right output (correct class, wrong casing or punctuation), the fix is almost always in your example formatting, not in adding more examples. Make every example letter-perfect for the format you expect.
  • For multi-label or multi-field extraction tasks, show at least one example where a field is absent or null. If you only show examples where every field is populated, the model will hallucinate values rather than return null.
  • Dynamic few-shot selection -- embedding your example bank and retrieving the k-nearest examples for each input at runtime -- consistently outperforms a static set on high-variance inputs. It adds one embedding call latency, but the accuracy lift is often worth it for production systems.

Common pitfalls

  • Mistake: Using imagined, idealized examples that don't match real input distribution. Fix: Pull examples from actual production or staging data, including messy and edge-case inputs.
  • Mistake: Inconsistent formatting across examples (mixed delimiters, label positions, whitespace). Fix: Template your examples programmatically so every one follows the exact same structure.
  • Mistake: Adding more examples to fix a flawed pattern instead of fixing the examples themselves. Fix: Diagnose which input type is failing, then fix or replace the relevant example rather than appending more.
  • Mistake: Ignoring token cost of few-shot blocks at scale. Fix: Log prompt_tokens per call and model the monthly cost. Revisit fine-tuning if few-shot overhead exceeds a meaningful share of total spend.

When to use few-shot vs alternatives

Option Use when Avoid when
Zero-shot The task is well-defined, the model handles it reliably, and token budget is tight. Output format is strict or the task involves domain-specific labeling the model has not seen.
Static few-shot Input distribution is fairly uniform, you have 3-8 strong examples, and simplicity matters. Inputs vary widely enough that no single static example set covers the real distribution.
Dynamic few-shot (retrieval) You have a large example bank and input variance is high; accuracy is more important than latency. You cannot afford an extra embedding call or do not have enough labeled examples to build a bank.
Fine-tuning Task is stable, volume is high (millions of calls), and you have hundreds of labeled examples. Task definition changes frequently or you need to ship this week without a training pipeline.

Code Example

python
# openai>=1.0.0
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from environment

FEW_SHOT_PROMPT = """Classify the sentiment of each review as POSITIVE, NEGATIVE, or NEUTRAL.

Review: "The noise-cancelling is incredible, battery lasts forever."
Sentiment: POSITIVE

Review: "Broke after two weeks. Customer service never replied."
Sentiment: NEGATIVE

Review: "It's fine. Does what it says on the box."
Sentiment: NEUTRAL

Review: "{review}"
Sentiment:"""

def classify_review(review: str) -> str:
    prompt = FEW_SHOT_PROMPT.format(review=review)
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        max_tokens=10,
        temperature=0,
    )
    return response.choices[0].message.content.strip()

print(classify_review("Sounds great but the fit is terrible."))

How this code works

This code classifies product review sentiment (POSITIVE, NEGATIVE, NEUTRAL) using a "few-shot prompting" technique with an AI model. It starts by setting up the OpenAI client with client = OpenAI(), which conveniently reads the necessary API key from the environment without needing explicit configuration in the code itself. The core of this approach lies in the FEW_SHOT_PROMPT string. This multiline text provides the AI with several examples of reviews and their corresponding sentiments, effectively "teaching" the model the desired output format and classification style for new, unseen reviews.

The classify_review function takes a new review, inserts it into the FEW_SHOT_PROMPT template using .format(review=review), and sends this complete prompt to the AI. The client.chat.completions.create method handles the interaction with the specified model="gpt-4o-mini". A crucial detail is max_tokens=10, which limits the AI's response length. This ensures the model provides only the single sentiment word (like "POSITIVE") and doesn't generate additional text. Setting temperature=0 also makes the AI's classification very consistent and deterministic. Finally, the code extracts the sentiment from the AI's response using response.choices[0].message.content and prints the result.

Production-grade example

Adds retries with backoff, timeouts, token logging, unexpected-label validation, and graceful error return.

python
# openai>=1.0.0, tenacity>=8.0
import logging
import os
import time
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
from openai import OpenAI, RateLimitError, APITimeoutError, APIConnectionError

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=10.0)

EXAMPLES = [
    ("Broke after two weeks. No refund offered.", "NEGATIVE"),
    ("Fast shipping, exactly as described.", "POSITIVE"),
    ("Average. Nothing special.", "NEUTRAL"),
]

def build_prompt(examples: list[tuple[str, str]], review: str) -> str:
    lines = ["Classify the sentiment as POSITIVE, NEGATIVE, or NEUTRAL.\n"]
    for text, label in examples:
        lines.append(f'Review: "{text}"')
        lines.append(f"Sentiment: {label}\n")
    lines.append(f'Review: "{review}"')
    lines.append("Sentiment:")
    return "\n".join(lines)

@retry(
    retry=retry_if_exception_type((RateLimitError, APITimeoutError, APIConnectionError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
)
def classify_review(review: str, examples: list[tuple[str, str]] = EXAMPLES) -> dict:
    prompt = build_prompt(examples, review)
    start = time.monotonic()
    try:
        response = client.chat.completions.create(
            model="gpt-4o-mini",
            messages=[{"role": "user", "content": prompt}],
            max_tokens=10,
            temperature=0,
        )
    except Exception as exc:
        log.error("LLM call failed", extra={"error": str(exc), "review_snippet": review[:60]})
        return {"sentiment": "UNKNOWN", "error": str(exc), "latency_ms": None}

    latency_ms = round((time.monotonic() - start) * 1000)
    usage = response.usage
    label = response.choices[0].message.content.strip().upper()
    valid = {"POSITIVE", "NEGATIVE", "NEUTRAL"}
    if label not in valid:
        log.warning("Unexpected label", extra={"raw": label, "review_snippet": review[:60]})
        label = "UNKNOWN"

    log.info(
        "classify_review",
        extra={
            "label": label,
            "prompt_tokens": usage.prompt_tokens,
            "completion_tokens": usage.completion_tokens,
            "latency_ms": latency_ms,
        },
    )
    return {"sentiment": label, "prompt_tokens": usage.prompt_tokens, "latency_ms": latency_ms}

if __name__ == "__main__":
    result = classify_review("Sounds great but the fit is terrible.")
    print(result)

How this code works

This Python code classifies the sentiment of product reviews as "POSITIVE", "NEGATIVE", or "NEUTRAL" using a few-shot prompting technique with a large language model. It begins by defining EXAMPLES – a list of existing reviews and their correct sentiments. The build_prompt function takes these examples and a new review, formatting them into a clear, instructional message for the AI model. The core classify_review function then sends this prepared prompt to an OpenAI model like gpt-4o-mini, which learns the desired classification pattern from the provided demonstrations before processing the new review.

For robust communication with the AI, the classify_review function is decorated with @retry. This subtle but crucial feature automatically re-attempts the AI call multiple times if temporary issues like RateLimitError or network problems occur, making the system more resilient. Inside, client.chat.completions.create sends the structured prompt, specifying max_tokens=10 to encourage a concise sentiment label and temperature=0 for consistent responses. After receiving the AI's output, the code validates that the extracted sentiment label is one of the expected values. If the AI provides anything unexpected, the code gracefully defaults the sentiment to "UNKNOWN" to handle edge cases and logs a warning.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a few-shot prompt that extracts structured data from a raw product review: the product name, a sentiment label (POSITIVE/NEGATIVE/NEUTRAL), and a one-sentence summary. Test it on at least three different reviews and verify the output is consistently formatted as JSON. Pay attention to what happens when information is missing.

python
# openai>=1.0.0
from openai import OpenAI
import json

client = OpenAI()  # set OPENAI_API_KEY in your environment

# TODO: write 2-3 input-output example pairs that show the model
# how to produce JSON with keys: product_name, sentiment, summary.
# Use triple-quoted strings for readability.
FEW_SHOT_EXAMPLES = """
# your examples here
"""

def extract_review_data(review_text: str) -> dict:
    # TODO: build the full prompt by combining FEW_SHOT_EXAMPLES
    # with the new review_text, then call the API.
    prompt = ""
    # TODO: call client.chat.completions.create with temperature=0
    # and parse the response as JSON
    raw_output = ""
    # TODO: return a Python dict (use json.loads)
    return {}

# Test inputs
test_reviews = [
    "The AirPods Pro fit perfectly and the sound quality is outstanding.",
    "Keyboard stopped working after a month. Very disappointed.",
    "It arrived on time.",
]

for review in test_reviews:
    result = extract_review_data(review)
    print(json.dumps(result, indent=2))

Quick check

  1. You add a 5th example to fix a recurring misclassification but accuracy gets worse. What is the most likely explanation?

  2. Your few-shot classifier runs at 10M calls per month. Each call adds 600 tokens of examples. What is the main operational question this raises?

  3. Why does setting temperature=0 matter specifically for few-shot classification prompts?

Self-check: Without looking at your notes, explain why two prompts with identical instructions but different example formatting can produce different accuracy, and describe one concrete change you would make to examples when the model keeps hallucinating a field that is sometimes absent in real data.