Phase 1: Programming & AI Foundations

Base models vs instruction-tuned models

Beginner ~14 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you have a super-smart digital brain, like a robot pet, that has read almost every book, website, and conversation on the internet. It knows an unbelievable amount of stuff! Now, there are two main kinds of these digital brains. The first kind is like a very smart puppy. This puppy has watched and learned from everything around it – how people talk, how stories are written, even how to do math problems. Because it's seen so much, it's really good at guessing what comes next in a sentence or a story. If you start a sentence, it can often finish it in a way that makes sense, just like a puppy might try to copy what it sees.

But here's the trick: this "base model" puppy doesn't actually understand commands. If you give it a big article and say, "Summarize this for me:", it might just keep writing more parts of the article, or even write another article that sounds similar. Why? Because when it "read" the internet, it saw lots of times where a phrase like "Summarize this article:" was followed by... well, more articles! It's not trying to be unhelpful; it's just doing what it's best at: guessing what usually comes next based on everything it's observed. It's like the puppy just keeps playing when you say "sit" because it hasn't been taught what "sit" means as an instruction.

Now, the second kind of digital brain starts out as one of these smart puppies. But then, it gets special training! This is like taking our clever puppy to a dog trainer. The trainer (who are actually human experts) teaches it specific tricks and commands. They show it many examples of "Summarize this" prompts and then show it exactly what a good summary looks like. They also teach it what "write a poem," "answer this question," or "explain this" means. After lots and lots of practice, this digital brain learns to understand that when you give it an instruction, you're giving it a job to do, not just something to continue writing. It learns to actually follow your commands.

So, when you want an AI to be really helpful and do specific tasks for you – like acting as a chatbot, answering questions about a document, or even helping you write computer code – you'll always want to use one of these specially "trained" models. They are called "instruction-tuned" models because they've been fine-tuned to follow your instructions. It's like choosing between a brilliant but untrained puppy and a well-behaved service dog. Both are smart, but one is ready to understand and complete tasks for you right away!

The mental model: two different training objectives

Base models are trained on a simple objective called next-token prediction (technically, causal language modeling). Given a trillion-token corpus scraped from the web, books, and code repos, the model learns to assign high probability to the token that actually comes next. After training, the model is a very powerful text-completion engine. GPT-3 (the original), Llama 2 base, and Mistral 7B base are examples. They are genuinely useful for tasks that look like completion -- autocomplete, creative continuation, filling in blanks inside structured templates -- but they will not reliably obey a command phrased as a direct instruction. The probability distribution has been shaped by every kind of text that exists online, not by human demonstrations of "good assistant behavior."

Instruction-tuned models add a second training phase (and often a third). In supervised fine-tuning, the model is shown thousands to millions of examples of (prompt, ideal-response) pairs, assembled by human annotators or generated synthetically. This shifts the model's behavior dramatically: it learns that a certain input format means "generate a response," not "complete a document." The second phase -- RLHF or DPO -- further tunes for human preferences: helpfulness, harmlessness, and honesty. GPT-4o, Claude 3.5 Sonnet, and Llama 3 Instruct are all instruction-tuned. The "Instruct" or "Chat" suffix in a model name is usually the giveaway.

A real-world scenario: document extraction pipeline

Imagine you're building a pipeline that reads insurance claim PDFs, extracts structured fields (claim number, date of loss, coverage type), and writes them to a database. You send the OCR output to a model with the prompt: "Extract the claim number, date of loss, and coverage type. Return JSON."

With a base model, the output is unpredictable. It might continue the legal-sounding document prose, or hallucinate a second claim form. With gpt-4o-mini or claude-haiku, the model parses the text and returns the JSON you asked for. The instruction-tuned model has learned, from thousands of similar training examples, that this kind of prompt demands structured extraction. You can further constrain output with JSON mode or a response schema -- features that only exist because the model is instruction-aware.

Tradeoffs and when base models still win

Base models have niche advantages. Because they haven't been fine-tuned away from raw text patterns, they can be better starting points for domain-specific fine-tuning. If you're a medical AI company and you want to fine-tune on clinical notes, starting from a base checkpoint (like Llama 3 base) gives you more "room" to shape behavior without fighting the assistant persona baked in by RLHF. Base models also tend to be less "refusal-happy": instruction tuning installs safety behaviors that occasionally over-trigger. For research or controlled environments where refusals are a bigger problem than safety, this matters.

For green-field applications calling a hosted API, there's almost no reason to use a base model. Instruction-tuned variants are available at the same price and latency from every major provider.

What changes at scale

At 10 users, the difference barely matters in practice -- you'll feel it in prompt reliability, but it won't break your product. At 10,000 users, the behavioral consistency of instruction-tuned models becomes critical: base model unpredictability creates edge cases that flood your support queue. At 10 million users, you're likely running your own fine-tuned model, and you'll choose between starting from a base checkpoint (more control over final behavior) vs. starting from an instruct checkpoint (faster to deploy, less training data needed for task-specific behavior). At that scale, the model type also affects cost: running a 70B base model self-hosted versus a 70B instruct model self-hosted costs roughly the same per token, but the instruct model usually achieves the task in fewer turns, which reduces total token spend.

Prompt format is not optional

Instruction-tuned models expect a specific prompt format, baked in during SFT. Llama 3 Instruct uses a <|begin_of_text|> and <|eot_id|> token schema. Mistral Instruct uses [INST] and [/INST] tags. OpenAI models use a messages array with system, user, and assistant roles. When you call a hosted API like OpenAI or Anthropic, the SDK handles this formatting for you. But if you're running a self-hosted model with Ollama or vLLM and you call the raw model without applying the correct chat template, the model will behave more like a base model -- generating completions instead of responses. Always apply the model's documented chat template when running inference outside a managed API.

Key Takeaways

  • Base models predict the next token; they do not follow commands by design.
  • Instruction-tuned models are fine-tuned via SFT and RLHF to respond to prompts.
  • Use instruction-tuned variants for nearly all production applications.
  • Prompt structure differs significantly between base and instruction-tuned models.

Pro tips

  • When self-hosting with vLLM or Ollama, always pass the chat_template parameter explicitly or use the tokenizer's apply_chat_template() method. Skipping this is the single most common reason a self-hosted instruct model outputs garbage -- it's seeing raw text instead of the formatted prompt it was fine-tuned on.
  • If you're evaluating a new instruction-tuned model and its outputs feel weirdly literal or it keeps continuing your prompt instead of answering, check whether you accidentally hit the base model endpoint. Both variants are often listed in the same provider dashboard.
  • RLHF-tuned models are biased toward verbose, confident-sounding answers. When you need a model to say 'I don't know,' you usually have to state that explicitly in the system prompt, because training has rewarded fluency over epistemic honesty.
  • For fine-tuning on a custom task, start from an instruction-tuned checkpoint rather than base unless you have more than ~100k high-quality examples. The instruct model's existing instruction-following behavior massively reduces how much data you need to land on the specific behavior you want.

Common pitfalls

  • Mistake: Using base model completion-style prompts (no instruction framing) against an instruction-tuned API endpoint. Fix: Always use the messages array format with distinct system and user roles; the model was trained to expect this.
  • Mistake: Assuming all 'chat' models have the same safety calibration. Fix: Test refusal rates on your specific domain. Claude and GPT-4o have different thresholds; pick the model whose defaults match your use case.
  • Mistake: Starting fine-tuning from a base checkpoint when you only have a few thousand examples. Fix: Fine-tune from an instruct checkpoint instead; it needs far less data to learn a new task on top of existing instruction-following behavior.
  • Mistake: Treating temperature=0 as deterministic. Fix: It's near-deterministic but not guaranteed; for truly reproducible extraction pipelines, hash inputs and cache responses rather than relying on model determinism.

When to use a base model vs an instruction-tuned model

Option Use when Avoid when
Instruction-tuned (e.g., gpt-4o-mini, claude-haiku, llama-3-instruct) Building any application where users or code send natural-language commands: chatbots, Q&A, extraction, summarization, code generation. You need a starting checkpoint for heavy domain-specific fine-tuning and have > 100k training examples; the RLHF persona may constrain your target behavior.
Base model (e.g., llama-3-base, mistral-7b-base) Starting a fine-tuning run for a specialized domain (medical, legal, code) where you want full control over the final behavior without fighting built-in safety fine-tuning. Building anything user-facing without additional fine-tuning; raw outputs will be unreliable and unpredictable in response to instructions.
Hosted API (OpenAI, Anthropic, Google) Prototyping, early production, or when you need the strongest available models. No infra to manage. Data cannot leave your network, or per-token costs at your volume outweigh self-hosting operational cost.
Self-hosted instruct model (vLLM, Ollama, TGI) Cost optimization at scale, strict data residency requirements, or need for a model fine-tuned on proprietary data. You lack ML infra expertise or don't have volume to justify the operational overhead. Misconfigurations (chat templates, quantization) are easy to get wrong.

Code Example

python
# openai>=1.0.0
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from environment

# Instruction-tuned model: treats input as a command
response = client.chat.completions.create(
    model="gpt-4o-mini",  # instruction-tuned variant
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Summarize this in one sentence: The sky is blue because of Rayleigh scattering of sunlight."},
    ],
    max_tokens=60,
)
print(response.choices[0].message.content)
# Expected: a one-sentence summary, not a continuation of the input text

How this code works

This code demonstrates how an instruction-tuned large language model (LLM) like gpt-4o-mini processes a request, specifically following a given command. It connects to the OpenAI API using client = OpenAI(), which securely reads an OPENAI_API_KEY from the environment to authenticate the session. The core of the interaction happens within client.chat.completions.create(), where the model is explicitly set to gpt-4o-mini, identifying it as an instruction-tuned variant designed to understand and act on directives.

The messages parameter provides the conversational context. It includes a system role to guide the model's overall behavior ("You are a helpful assistant.") and a user role with the specific instruction: "Summarize this in one sentence: The sky is blue because of Rayleigh scattering of sunlight." This instruction-tuned model understands to perform the summary task rather than simply continuing the input text. A subtle but important detail is max_tokens=60. This option limits the length of the model's response to a maximum of 60 tokens, ensuring the output remains concise and helping manage API usage. Finally, print(response.choices[0].message.content) displays the summary generated by the model.

Production-grade example

Adds retries with backoff, timeouts, structured token logging, and specific exception handling vs the basic example.

python
# openai>=1.0.0, tenacity>=8.2.0
import os
import time
import logging
from openai import OpenAI, RateLimitError, APITimeoutError, APIStatusError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
logger = logging.getLogger(__name__)

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)

RETRYABLE = (RateLimitError, APITimeoutError)

@retry(
    retry=retry_if_exception_type(RETRYABLE),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
)
def extract_fields(document_text: str) -> dict:
    start = time.perf_counter()
    try:
        response = client.chat.completions.create(
            model="gpt-4o-mini",
            messages=[
                {"role": "system", "content": "Extract structured fields. Return valid JSON only."},
                {"role": "user", "content": f"Document:\n{document_text}\n\nExtract: claim_number, date_of_loss, coverage_type."},
            ],
            response_format={"type": "json_object"},
            max_tokens=256,
            temperature=0,
        )
    except APIStatusError as exc:
        logger.error("Non-retryable API error", extra={"status": exc.status_code, "body": exc.body})
        raise

    usage = response.usage
    latency_ms = (time.perf_counter() - start) * 1000
    logger.info(
        "LLM call complete",
        extra={
            "model": response.model,
            "prompt_tokens": usage.prompt_tokens,
            "completion_tokens": usage.completion_tokens,
            "latency_ms": round(latency_ms, 1),
        },
    )
    import json
    return json.loads(response.choices[0].message.content)


if __name__ == "__main__":
    sample = "Claim #CLM-2024-00192. Date of loss: March 3, 2024. Coverage: Comprehensive."
    result = extract_fields(sample)
    print(result)

How this code works

This Python script demonstrates how to extract structured information from unstructured text using an instruction-tuned large language model. Its job in the lesson is to showcase how models like gpt-4o-mini can follow precise instructions to perform tasks such as field extraction, illustrating their power for specific applications. The core extract_fields function sends a document_text to the OpenAI API with clear messages. It includes a "system" instruction to "Return valid JSON only" and a "user" instruction detailing which specific fields, like claim_number or date_of_loss, should be extracted. A vital part of this setup is response_format={"type": "json_object"}, which ensures the model's output is consistently valid and parsable JSON.

For production robustness, the code utilizes the @retry decorator from tenacity. This decorator automatically handles temporary API issues such as RateLimitError or APITimeoutError by retrying the call with exponential backoff, making the application more reliable. The temperature=0 setting ensures deterministic outputs, which is crucial for predictable data extraction rather than creative generation. After a successful call, the code logs performance metrics like latency_ms and token usage, then converts the model's JSON string output into a Python dictionary using json.loads. The explicit response_format is a subtle but crucial detail often overlooked by beginners, guaranteeing reliable structured data that can be immediately used by other parts of an application.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Using the OpenAI Python SDK, send the same document text to the API twice: once formatted as a raw completion prompt (simulating base model usage) and once as a proper instruction with a system message. Log both outputs. Observe how the instruction-tuned model's response differs from naive text continuation.

python
# openai>=1.0.0
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

DOCUMENT = "The quarterly revenue for Acme Corp was $4.2M, up 12% YoY. Operating costs were $3.1M."

# TODO: Call the API as if simulating a base-model completion prompt.
# Use the messages format but put the document directly in user content
# with NO instruction -- just the raw document text. Print the output.

# TODO: Call the API a second time with a proper system message
# instructing the model to extract revenue and operating cost as JSON.
# Print the output and note the structural difference.

def call_model(messages):
    # TODO: implement the API call with model="gpt-4o-mini", max_tokens=150
    pass

print("--- No instruction (base-style) ---")
# TODO: call call_model with just the document in user content

print("\n--- With instruction (instruct-style) ---")
# TODO: call call_model with system prompt + instruction in user content

Quick check

  1. A base model is prompted with 'List three causes of inflation:' and responds with more examples of economic quiz questions. Why?

  2. You're self-hosting Llama 3 Instruct with vLLM and getting incoherent outputs despite a well-formed prompt. What should you check first?

  3. When fine-tuning a model for a niche legal document task with only 5,000 examples, which starting checkpoint is the better choice?

Self-check: Without looking at notes, explain to a colleague why a base model fails to summarize a document on demand, what training step fixes this, and which type of checkpoint you'd choose as a starting point for fine-tuning with a 5,000-example dataset and why.