Phase 1: Programming & AI Foundations

Language fundamentals, virtual environments & package management

Beginner ~12 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you want to become an amazing chef who can create all sorts of incredible dishes, especially the kind that use super smart robots to help you (that's like AI!). Python is like the world's most popular cookbook for these robot chefs. It's not the fastest cookbook, but it has the most amazing collection of tools, special ingredients, and pre-made sauces that other chefs have shared. To cook truly groundbreaking robot dishes, you first need to learn the basics. This means knowing your main ingredients, like flour or sugar (we call these "data types" in Python), and understanding basic cooking steps, like chopping or mixing (these are like "functions"). If you try to jump straight to making a complicated robot cake without knowing these basics, your kitchen will quickly become a messy disaster, and your cake won't turn out right.

Beyond the basics, imagine you're cooking several different, complex dishes. One recipe might need a fancy high-speed blender, a specific type of rare herb, and a special baking pan. Another recipe, for a different meal, might need a different kind of blender, a different set of herbs, and a special steamer. If you just throw everything into one big kitchen, you’ll end up with all your tools and ingredients mixed up. You might use the wrong blender for the wrong recipe, or the rare herb from one recipe could accidentally get into another, ruining both. That's why professional chefs keep their workspaces super organized for different dishes.

This is where "virtual environments" come in. Think of a virtual environment as setting up a brand new, sparkling clean mini-kitchen just for one specific recipe or project. When you start a new robot-chef project, you set up a fresh mini-kitchen. Then, you use a special assistant called pip to order exactly the unique tools (like that high-speed blender) and special ingredients (like that rare herb) your recipe needs, and pip puts them only in that mini-kitchen. If your recipe needs a list of specific things, you'll write them down in a "requirements file" – it’s like a shopping list for pip to get everything perfectly for your kitchen. This keeps each project's kitchen neat and tidy, with only what it needs.

So, by first learning the basic cooking skills and ingredients, and then setting up these special, organized mini-kitchens, you make sure your robot chef projects always have the exact tools and ingredients they need, without any mix-ups. This means when you build a cool new AI feature – maybe a robot that can write stories or help design a new game – you won't have to worry about old tools or wrong ingredients getting in the way. Instead, you'll be a super-organized AI chef, ready to create amazing things with Python, knowing exactly what's in your digital pantry for each project.

Python's data model is built around objects. Every value, whether an integer, a string, a function, or a class instance, is an object with a type, an identity, and a value. In AI work, you'll spend most of your time with four core types: str (prompt text, API keys, model names), int/float (token counts, temperatures, probabilities), list (message histories, retrieved documents, embedding vectors), and dict (API request bodies, structured outputs, configuration). Understanding that Python lists are ordered and mutable while tuples are ordered and immutable matters when you're assembling a chat history you want to guarantee nobody modifies mid-flight.

Functions are the primary unit of reuse in Python, and in AI code they do a lot of heavy lifting. You'll write functions that format prompts, parse API responses, retry failed requests, and chunk documents. Python functions support default arguments, keyword arguments, args, and *kwargs, all of which you'll encounter when wrapping LLM client libraries. The single most impactful habit you can build early is adding type hints to every function signature. When a function says def embed(text: str, model: str = "text-embedding-3-small") -> list[float], you know exactly what it consumes and produces. This pays off immediately when reading AI library source code, which is heavily type-annotated.

List comprehensions and dictionary comprehensions are not syntactic sugar; they're idiomatic Python that shows up constantly in AI data pipelines. Transforming a list of document objects into a list of their text fields looks like [doc.page_content for doc in docs]. Filtering retrieved chunks by score looks like [c for c in chunks if c.score > 0.75]. If you're writing explicit for-loops to build new lists, you're missing a tool that makes the code both shorter and faster. Python's f-strings are equally fundamental for prompt construction: f"Answer the question based on this context: {context} Question: {question}".

Virtual environments solve a specific, painful problem: different projects need different versions of the same library. Say project A needs openai==0.28 (the old API) and project B needs openai>=1.0 (the new client). If you install both globally, one of them breaks. A virtual environment is a directory that contains a copy of the Python interpreter and its own site-packages folder, completely isolated from everything else. The standard workflow is: python -m venv .venv to create it, source .venv/bin/activate on macOS/Linux (or .venv\Scripts\activate on Windows) to activate it, and then every pip install lands inside that environment. Your shell prompt changes to show (.venv) so you always know which environment is active. When you're done, deactivate returns you to the system Python.

Pip is straightforward but has a few behaviors worth knowing. pip install openai installs the latest version. pip install openai==1.30.1 installs a specific version. pip install "openai>=1.0,<2.0" installs within a range. After installing everything a project needs, pip freeze > requirements.txt writes every installed package with its exact version to a file. Anyone cloning your repo can then run pip install -r requirements.txt and get an identical environment. For AI projects, this matters because model behavior can change subtly when a tokenizer library or HTTP client version changes underneath you. Some teams prefer pip-tools or poetry for more sophisticated dependency resolution, but for this course, venv plus pip freeze is sufficient and has no additional dependencies.

At scale, the tooling picture shifts. Ten users: requirements.txt and a shared .env file for secrets works fine. Ten thousand users in production: you're packaging your application into a Docker image with a pinned base image and a multi-stage build, and your dependency resolution is part of CI. Ten million users: you're likely using a compiled serving runtime (Triton, vLLM, or similar) and Python's role is orchestration, not inference. Understanding where Python fits in that stack, and where it hands off to lower-level runtimes, is what separates developers who build demo apps from those who build systems.

Key Takeaways

  • Use virtual environments per project; never install AI packages into the system Python.
  • Type hints on function signatures make LLM-generated code easier to validate and debug.
  • Pin exact dependency versions in requirements.txt to prevent silent breakage across machines.
  • Dictionaries and lists are the primary data structures for shuttling data to and from AI APIs.

Pro tips

  • When debugging a broken virtual environment, don't fix it: delete the .venv directory and recreate it from requirements.txt. Attempting to patch a corrupted env wastes more time than rebuilding it.
  • Python's dictionary .get(key, default) is safer than dict[key] when parsing LLM API responses because the API shape can change between model versions and optional fields are common.
  • Set PYTHONPATH=. in your .env file (loaded via python-dotenv) so local module imports resolve correctly regardless of which directory you run scripts from. This prevents a whole class of ModuleNotFoundError issues in AI projects.
  • Type hint your prompt-building functions with -> str and your response-parsing functions with -> dict[str, Any]. When an LLM generates code for you to review, mismatched types in these boundaries are where bugs hide.

Common pitfalls

  • Mistake: Installing packages with the system Python (sudo pip install). Fix: Always activate a project virtual environment first. System Python is used by OS tools and corrupting it breaks things that have nothing to do with your project.
  • Mistake: Committing .venv or venv directories to git. Fix: Add .venv/ to .gitignore immediately. Share the environment via requirements.txt, not by copying gigabytes of binary files.
  • Mistake: Using pip freeze in a cluttered global environment, which dumps every installed package not just your project's deps. Fix: Run pip freeze only inside a clean, project-specific virtual environment to get an accurate dependency list.
  • Mistake: Storing API keys as string literals in source files. Fix: Load them from environment variables using os.environ["API_KEY"] and keep secrets in a .env file that is excluded from version control.

When to use venv vs conda vs poetry

Option Use when Avoid when
venv + pip Pure Python projects, simple dependency trees, CI/CD environments, or when you want zero extra tooling overhead. You need non-Python binary dependencies (CUDA libs, GDAL) or cross-language environments.
conda You need to manage non-Python system libraries alongside Python packages, common in data science with GPU drivers or C extensions. You are deploying to production Docker containers; conda images are large and slower to build than pip-based ones.
poetry You're building a Python library others will install, or you want deterministic lock files and a single tool for packaging plus dependency resolution. The team is new to Python; poetry's resolver and config format add cognitive overhead on top of fundamentals.

Code Example

python
# Requires Python 3.10+ for match syntax; tested with standard library only

from typing import Any

def classify_response(response: dict[str, Any]) -> str:
    """Extract and classify an LLM response payload by finish reason."""
    finish_reason = response.get("choices", [{}])[0].get("finish_reason", "unknown")

    match finish_reason:
        case "stop":
            return response["choices"][0]["message"]["content"]
        case "length":
            return "[TRUNCATED] " + response["choices"][0]["message"]["content"]
        case "content_filter":
            return "[BLOCKED]"
        case _:
            return f"Unexpected finish reason: {finish_reason}"

# Example usage with a mock payload
mock_response = {
    "choices": [{"finish_reason": "stop", "message": {"content": "Paris"}}]
}
print(classify_response(mock_response))  # Paris

How this code works

This code defines a function, classify_response, designed to process and interpret a structured dictionary received from a Large Language Model (LLM). Its job in the lesson is to demonstrate handling different outcomes based on how an LLM finishes generating text, extracting useful information, or indicating issues. The function safely extracts a finish_reason from the complex response payload. It uses response.get("choices", [{}])[0].get("finish_reason", "unknown") to navigate the dictionary. A subtle but important detail here is the use of get with default values; [{}] ensures the code doesn't crash if "choices" is missing, and "unknown" catches cases where "finish_reason" itself isn't present, preventing KeyError exceptions and making the code robust against incomplete LLM responses.

Once the finish_reason is determined, a match statement (a Python 3.10+ feature for structural pattern matching) efficiently handles different scenarios. If the finish_reason is "stop", it directly returns the generated content. For "length", it indicates truncation, and "content_filter" results in a "[BLOCKED]" message. The case _ acts as a wildcard, catching any other unexpected finish_reason and returning an informative message, which is crucial for debugging and understanding unknown LLM behaviors. This pattern clearly separates logic for various LLM outcomes.

Production-grade example

Adds retries on transient errors, timeout, structured logging with token counts, and env-var auth.

python
# openai>=1.0.0, tenacity>=8.2.0
import os
import logging
import time
from typing import Any

import openai
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

logging.basicConfig(
    format="%(asctime)s %(levelname)s %(name)s %(message)s",
    level=logging.INFO,
)
log = logging.getLogger("llm_client")

client = openai.OpenAI(api_key=os.environ["OPENAI_API_KEY"])  # never hardcode

@retry(
    retry=retry_if_exception_type((openai.RateLimitError, openai.APITimeoutError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(4),
    reraise=True,
)
def complete(
    prompt: str,
    model: str = "gpt-4o-mini",
    max_tokens: int = 512,
    timeout: float = 15.0,
) -> str:
    """Call the chat completions endpoint with retries, logging, and timeout."""
    start = time.monotonic()
    try:
        response = client.chat.completions.create(
            model=model,
            messages=[{"role": "user", "content": prompt}],
            max_tokens=max_tokens,
            timeout=timeout,
        )
    except openai.AuthenticationError as exc:
        log.error("auth_error", extra={"detail": str(exc)})
        raise  # don't retry auth failures

    latency_ms = round((time.monotonic() - start) * 1000)
    usage = response.usage
    log.info(
        "llm_call_success",
        extra={
            "model": model,
            "prompt_tokens": usage.prompt_tokens,
            "completion_tokens": usage.completion_tokens,
            "latency_ms": latency_ms,
        },
    )
    finish = response.choices[0].finish_reason
    text = response.choices[0].message.content or ""
    if finish == "length":
        log.warning("response_truncated", extra={"max_tokens": max_tokens})
    return text

if __name__ == "__main__":
    print(complete("What is 2 + 2? Reply with only the number."))

How this code works

This Python code provides a robust and reliable way to interact with OpenAI's large language models (LLMs), which is fundamental for AI development. Its job is to send a text prompt to an LLM, receive a response, and gracefully handle common real-world challenges like temporary network issues or API rate limits.

The script begins by setting up logging to record execution details and then initializes an openai.OpenAI client. Crucially, the api_key is fetched securely from os.environ["OPENAI_API_KEY"] to prevent sensitive credentials from being hardcoded. The core interaction happens in the complete function, which calls client.chat.completions.create to send the prompt to the specified model. A key feature is the @retry decorator from the tenacity library, automatically re-attempting the LLM call if transient openai.RateLimitError or openai.APITimeoutError exceptions occur. This uses an exponential backoff strategy for up to stop_after_attempt(4) tries. A subtle but important detail is that openai.AuthenticationError is caught separately and not retried; this kind of error indicates a fundamental setup problem that retrying won't solve. After a successful call, llm_call_success logs useful metrics like latency_ms and prompt_tokens. It also logs a warning if the response is truncated because it hit max_tokens. The if __name__ == "__main__": block provides a basic example of how to use the complete function.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Create a virtual environment, install the openai package, and write a function called build_prompt that accepts a context string and a question string and returns a formatted prompt string. Then write a second function parse_answer that takes a mock API response dictionary and returns the content string safely, defaulting to an empty string if the key is missing.

python
# Step 1: In your terminal (outside this file):
# python -m venv .venv
# source .venv/bin/activate  (or .venv\Scripts\activate on Windows)
# pip install openai
# pip freeze > requirements.txt

# Step 2: Implement the two functions below

from typing import Any

def build_prompt(context: str, question: str) -> str:
    # TODO: Return an f-string that embeds context and question
    # in a readable template. Include labels like 'Context:' and 'Question:'
    pass

def parse_answer(response: dict[str, Any]) -> str:
    # TODO: Safely extract response['choices'][0]['message']['content']
    # Return '' if any key is missing
    pass

if __name__ == "__main__":
    prompt = build_prompt("The Eiffel Tower is in Paris.", "Where is the Eiffel Tower?")
    print(prompt)

    mock = {"choices": [{"message": {"content": "Paris"}}]}
    print(parse_answer(mock))
    print(parse_answer({}))  # should print empty string

Quick check

  1. You activate a virtual environment and run pip install numpy. Where does numpy get installed?

  2. A teammate clones your repo and runs python app.py but gets an ImportError for openai. What is the most likely cause?

  3. Which expression safely extracts a nested value from a dict without raising KeyError if a key is absent?

Self-check: Without looking at notes, explain to a colleague why you create a new virtual environment per project, and describe the exact three-command sequence to create, activate, and record dependencies for a new AI project.