Phase 2: Prompt Engineering & LLM Patterns

System prompts, user prompts & multi-turn conversations

Intermediate ~14 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're setting up a super fun board game that an awesome robot is going to play with you. Before anyone rolls a dice or makes a move, you, the game designer, write down some secret rules just for the robot. This hidden rulebook is like a system prompt. It tells the robot things like, "Always be a helpful wizard," or "When you answer, only talk about space adventures," or "Your goal is to make the player laugh." The actual player doesn't see these rules, but they completely change how the robot acts and what kind of game you play. If you don't give the robot clear rules, it might just do anything, and the game could get really confusing or boring!

Now, when it's your turn in the game, you make a move or ask a question – this is your user prompt. For example, you might say, "What's my next quest?" or "Tell me a silly joke!" The robot looks at your question and remembers its secret rulebook to figure out what to do. So, if your secret rulebook said, "Only talk about space adventures," and you asked for a joke, the robot would tell you a space-themed joke, following both your request and its secret mission.

The tricky part comes when you have a multi-turn conversation, where you keep playing and talking back and forth. You might think the robot just "remembers" everything you've said before, like a friend does. But here's the secret: the robot doesn't have a magical memory inside its head for your game. Every single time you ask it a new question, you're actually sending the robot the entire story so far! It's like you're handing it a super long scroll with every question you've ever asked and every answer it's ever given, and then you add your new question to the very end of that scroll.

Why is this a big deal? Well, sending a short scroll is fast and easy. But if your game goes on for a long, long time, that scroll gets super long and heavy! Sending a huge scroll means the robot takes longer to read it all and figure out its next move, which can make the game slow or even expensive to play. So, when you're building games or apps that talk to robots, understanding this "scroll" means you can make smart choices about how much history to send, keeping the game fun, fast, and exactly how you want it!

The chat completions API accepts a list of messages, each with a role field that is one of system, user, or assistant. That list is the entire model context. The model does not know anything about previous API calls. It only knows what you put in that list right now. When you hear "the model has memory" in a demo, what is actually happening is that the application layer is maintaining a list and re-sending it. This mental model matters because it changes how you think about every architectural decision downstream.

The system message is processed first, with higher conceptual weight in most instruction-tuned models. Think of it as the outermost constraint layer. Anything you can resolve at the system level, you should resolve there rather than repeating it in each user message. Good candidates: output format rules (always return JSON, never include disclaimers), persona (terse senior engineer, friendly support agent), hard restrictions (never disclose pricing, always cite sources), and domain scope (you only answer questions about our product, politely refuse everything else). A poorly written system prompt is the most common reason developers escalate to bigger, more expensive models when they did not need to. If the output format keeps drifting, fix the system prompt before you switch from GPT-4o-mini to GPT-4o.

User messages represent the turn-by-turn input, and assistant messages are the model's previous responses that you replay back to it. This three-role structure lets you reconstruct a coherent dialogue. A real-world pattern you will use constantly: a customer support bot receives a user question, calls the LLM, gets a response, appends both sides to the in-memory list, then on the next HTTP request from the frontend, sends that full list again. The model sees the conversation so far and continues it naturally. The implementation is straightforward, but the tradeoffs compound quickly.

At 10 users with short sessions, you can store the full message list in memory or a simple dict keyed by session ID. At 10,000 concurrent users with conversations that go 50+ turns, you are looking at significant token overhead. A 50-turn conversation with an average of 200 tokens per message is 10,000 tokens of history before the user even types anything new. At that scale, you need a strategy. The three common approaches are: (1) sliding window, keep only the last N messages; (2) summary compression, periodically call the LLM to summarize earlier turns into a single system or user message and discard the originals; (3) retrieval-augmented history, embed past turns and retrieve only the semantically relevant ones before each call. Sliding window is simplest and fine for most cases, but it loses early context such as user preferences stated at the start. Summary compression preserves intent at the cost of an extra LLM call and some latency. At 10 million users you are likely moving to a database-backed history store with TTL-based expiry, and you are almost certainly not sending full histories.

One practical tradeoff that trips up developers: anthropic's Claude models handle the system prompt differently from OpenAI's models. Gemini models use a systemInstruction field outside the messages array. If you are building a multi-provider abstraction, your conversation serialization layer needs to account for these differences. Libraries like LiteLLM normalize this, but you still need to understand what each provider is doing under the hood because normalization leaks. Testing your system prompt on one provider and shipping to another without re-validation is a real source of production regressions. Version your system prompts the same way you version code. A system prompt change is a deployment.

Key Takeaways

  • System prompts set behavior for the whole session; write them like a strict job description.
  • The model has no memory. You send the full conversation history on every API call.
  • Token costs accumulate with each turn. Prune or summarize history at scale.
  • Role labels (system/user/assistant) are model-specific conventions, not universal HTTP standards.

Pro tips

  • Put hard constraints in the system prompt, not in every user message. If you are repeating 'always respond in JSON' in each user turn, you are wasting tokens and relying on the model to not forget it. One clear system-level instruction is cheaper and more reliable.
  • Inject dynamic context into the system prompt rather than the user message when it applies globally to the session. User role, account tier, and locale belong in system, not in a preamble you prepend to every user message, because that makes the history harder to parse and debug.
  • When debugging multi-turn drift, log the full messages array you are sending, not just the latest user message. The bug is almost always in what you accumulated in history, not in the user's input.
  • System prompt changes should go through the same review process as code changes. Write a regression test suite of 10-20 representative inputs and expected output shapes before you change a production system prompt. Silent regressions from prompt edits are one of the hardest production bugs to catch.

Common pitfalls

  • Mistake: Sending unbounded conversation history until you hit the context limit. Fix: Implement a sliding window or summary strategy from day one, not after your first context-length error in production.
  • Mistake: Assuming system prompt behavior is identical across providers. Fix: Test your system prompt on every provider you plan to use. Re-validate after every model upgrade.
  • Mistake: Putting secret instructions or sensitive business rules in the system prompt and assuming users cannot see them. Fix: Treat system prompts as guessable. Never put credentials, PII, or truly confidential logic there.
  • Mistake: Using the same conversation history object across concurrent users due to a shared mutable reference. Fix: Instantiate a fresh history list per session and store it in a session-scoped object or database row.

Where to put context: system vs user vs retrieved

Option Use when Avoid when
System prompt Content applies to every turn: persona, output format rules, hard restrictions, global domain scope. Content changes per request or per user action. Dynamic content in system prompts makes caching harder.
User message Content is specific to this turn: the user's actual question, a document to analyze, a code snippet to review. You are repeating the same boilerplate in every user message. That belongs in system.
Assistant message (replayed history) You need the model to continue a specific prior reasoning chain or maintain conversational coherence across turns. History is long and mostly irrelevant to the current question. Trim or summarize instead of blindly replaying.
Retrieved context injected into user message You have a large knowledge base and only a subset is relevant per turn. RAG pattern keeps total tokens manageable. Your entire knowledge base is small enough to fit in context and retrieval latency is a concern.

Code Example

python
# openai>=1.0.0
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from env

conversation = [
    {
        "role": "system",
        "content": "You are a Python code reviewer. Reply only with code and short inline comments. No prose."
    },
    {
        "role": "user",
        "content": "Review this: def add(a,b): return a+b"
    }
]

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=conversation
)

assistant_reply = response.choices[0].message.content
print(assistant_reply)

# To continue the conversation, append assistant reply then next user message
conversation.append({"role": "assistant", "content": assistant_reply})
conversation.append({"role": "user", "content": "Now add type hints."})

How this code works

The provided code demonstrates how to initiate and manage a multi-turn conversation with an AI model using the openai library. It sets up an AI persona as a Python code reviewer and sends an initial user request for code review. This is achieved by first importing OpenAI and creating an OpenAI() client object, which conveniently reads the OPENAI_API_KEY from environment variables by default. A conversation list is then defined, containing dictionaries where each specifies a role (like "system" to set the AI's behavior and "user" for the prompt) and content. This initial list is passed as the messages argument to client.chat.completions.create(), along with the chosen model, to get the AI's first response.

After receiving the initial assistant_reply, the code illustrates how to continue the interaction. To maintain the context for subsequent turns, the AI's previous response must be added back into the conversation list, effectively acting as the "assistant" role's turn. It's crucial that this assistant_reply is appended before the next user message is added. This sequential addition of both the AI's previous answer and the new user question to the conversation list ensures the model has access to the full chat history when generating its next reply, enabling it to remember past exchanges and maintain conversational flow across multiple turns.

Production-grade example

Adds retries with backoff, timeouts, token logging, history trimming, and structured error handling.

python
# openai>=1.0.0, tenacity>=8.0.0
import os
import logging
import time
from openai import OpenAI, APITimeoutError, RateLimitError, APIStatusError
from tenacity import retry, wait_exponential, stop_after_attempt, retry_if_exception_type

logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)

RETRYABLE = (RateLimitError, APITimeoutError)

@retry(
    retry=retry_if_exception_type(RETRYABLE),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4)
)
def chat(messages: list[dict], model: str = "gpt-4o-mini") -> str:
    start = time.monotonic()
    try:
        response = client.chat.completions.create(
            model=model,
            messages=messages,
            stream=False,
        )
    except APIStatusError as exc:
        logger.error("api_error status=%s body=%s", exc.status_code, exc.body)
        raise

    usage = response.usage
    elapsed = time.monotonic() - start
    logger.info(
        "llm_call model=%s prompt_tokens=%d completion_tokens=%d latency_s=%.2f",
        model, usage.prompt_tokens, usage.completion_tokens, elapsed
    )
    return response.choices[0].message.content


def build_session(system_prompt: str) -> list[dict]:
    return [{"role": "system", "content": system_prompt}]


def add_turn(history: list[dict], user_msg: str, assistant_msg: str | None = None) -> list[dict]:
    history.append({"role": "user", "content": user_msg})
    if assistant_msg:
        history.append({"role": "assistant", "content": assistant_msg})
    return history


def trim_history(history: list[dict], max_turns: int = 10) -> list[dict]:
    """Keep system prompt + last max_turns user/assistant pairs."""
    system = [m for m in history if m["role"] == "system"]
    dialogue = [m for m in history if m["role"] != "system"]
    return system + dialogue[-(max_turns * 2):]


if __name__ == "__main__":
    sys_prompt = "You are a terse Python tutor. Answer in 2 sentences max. No markdown."
    history = build_session(sys_prompt)
    for user_input in ["What is a list comprehension?", "Give a one-line example."]:
        history = trim_history(add_turn(history, user_input), max_turns=6)
        reply = chat(history)
        history = add_turn(history, user_input="", assistant_msg=reply)
        # remove the blank user placeholder we just added
        history = [m for m in history if m["content"]]
        print(f"Assistant: {reply}")

How this code works

This code provides a robust framework for managing multi-turn conversations with a large language model (LLM), crucial for building interactive AI applications. It handles communication, conversation history, and API resilience.

The core chat function sends messages to the gpt-4o-mini model. It’s made resilient by the @retry decorator, which automatically reattempts calls if temporary RateLimitError or APITimeoutError exceptions occur, preventing common interruptions. Extensive logging tracks prompt_tokens, completion_tokens, and latency_s for monitoring. Conversation flow is managed by several helper functions: build_session starts with a system_prompt, add_turn appends user and assistant messages, and trim_history keeps the conversation from growing indefinitely by retaining only the system prompt and the last max_turns of dialogue. This trimming prevents expensive, long context windows and token limit errors. In the main execution block, a loop demonstrates updating history with user input, calling chat for a reply, and then adding that reply to the history. A subtle but important cleanup step, history = [m for m in history if m["content"]], removes a temporary empty user message placeholder, ensuring clean conversation data.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a simple stateful CLI chat session in Python. The bot should act as a strict code reviewer who only speaks in bullet points. Keep the last 4 turns of history maximum. After each assistant reply, print the total number of messages currently in history so you can verify trimming is working.

python
# openai>=1.0.0
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

SYSTEM_PROMPT = """TODO: Write a system prompt that makes the model act
as a strict code reviewer who only responds in bullet points."""

def build_history(system_prompt: str) -> list[dict]:
    # TODO: return a list with just the system message
    pass

def trim(history: list[dict], max_turns: int = 4) -> list[dict]:
    # TODO: keep system message + last max_turns user/assistant pairs
    pass

def chat_turn(history: list[dict], user_input: str) -> str:
    # TODO: append user message, call the API, append assistant reply, return reply
    pass

if __name__ == "__main__":
    history = build_history(SYSTEM_PROMPT)
    print("Code Reviewer ready. Type 'quit' to exit.")
    while True:
        user_input = input("You: ")
        if user_input.lower() == "quit":
            break
        # TODO: call chat_turn, trim history, print reply and message count

Quick check

  1. You send a 10-turn conversation to the chat completions API. How many messages does the model have access to?

  2. Your chatbot's tone keeps drifting toward informal language after 15 turns. What is the most likely cause?

  3. Which role label should you use for content that applies globally to every turn of a session, like output format rules?

Self-check: Without referencing the lesson, explain what happens at the API level when a user sends message number 8 in a conversation. Then describe two strategies you would use to prevent context length errors in a production chatbot that expects very long sessions.