Phase 3: RAG & Knowledge Systems

Training dataset preparation, validation & deduplication

Advanced ~17 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're teaching a super-smart robot chef how to bake a new, special kind of cookie. The robot learns by looking at lots of recipes and examples. Now, you could give it a giant pile of 50,000 messy, half-written, or even duplicate recipe cards, or you could give it 5,000 super clear, perfectly written, and delicious recipes. Which do you think will make a better cookie? That's right, the fewer, better recipes! It's like how a chef learns more from a few perfect cooking lessons than from hundreds of confusing ones. This "recipe" part is what we call a "training dataset" in the world of AI.

So, preparing your training dataset is all about making sure those recipes are absolutely perfect. This means cleaning them up – like washing all your ingredients before you start baking. You wouldn't want to bake with muddy flour, right? It also means making sure the recipes are in the right format. If the robot chef expects instructions written step-by-step, but you give it a story about baking, it won't understand. It'll just get confused and try to fight against its own basic cooking knowledge. We have to match the recipe style the robot already knows, so it can build on what it learned before, not get frustrated.

But how do you know if your robot chef is really learning and not just memorizing the recipes you gave it? That's where "validation" and "deduplication" come in. Imagine you keep a few special, secret cookies aside that the robot chef has never seen or tasted before. After it's done practicing, you ask it to bake one of these secret cookies. This is "validation" – a real taste test to see if it truly learned to bake, or just copied. "Deduplication" is making sure you didn't accidentally give the robot a recipe for one of those exact secret cookies during its practice. If it tasted a secret cookie before, it's not a fair test, and you can't trust if it's actually getting better. We want to be sure it's learning to bake new things, not just remembering old ones.

By being super careful with your recipes – making sure they're clean, in the right style, and that your secret taste tests are truly secret – you're building a "dataset pipeline." This means you can really trust that your robot chef is becoming an amazing baker who can make delicious new cookies for anyone! So, when you're building smart computer programs, knowing how to pick, clean, and check your training data helps you make much smarter and more helpful AI.

Every fine-tuning framework (Hugging Face TRL, Axolotl, OpenAI's fine-tuning API) expects data in a specific format. For instruction-tuned models, the canonical format is a JSONL file where each line is a JSON object with a "messages" array matching the ChatML template: system, user, and assistant turns. If you are fine-tuning a Llama 3 model with Axolotl and your dataset uses the Alpaca format (instruction/input/output keys), Axolotl can auto-convert, but you are adding a conversion layer that can silently misalign prompts. The safest approach is to format your data into the exact prompt template the base model uses during pretraining. For Llama 3, that means the <|begin_of_text|><|start_header_id|>... tags. For Mistral, it means [INST]...[/INST]. Get this wrong and the model's loss landscape changes in ways that look like a hyperparameter problem but are actually a data formatting problem.

The practical workflow for dataset construction starts with sourcing. You have three levers: real human-labeled data (highest signal, expensive), synthetic data generated by a strong model like GPT-4o or Claude 3.5 Sonnet (scalable, but requires careful quality filtering), and domain text repurposed into instruction pairs (fast, but the conversions are lossy). For most production fine-tunes targeting a specific task like medical triage triage routing, customer support intent classification, or code review summarization, the best datasets combine a small seed of 200-500 human-curated gold examples with 2,000-8,000 synthetically generated examples filtered by a quality classifier or LLM judge. The seed examples define the quality bar; the synthetic data provides coverage.

Quality filtering is a separate pipeline stage from collection. Common filters include: minimum and maximum token length (drop examples below 20 tokens or above your context window), format validation (the assistant turn must not be empty), language detection (remove examples that slipped into the wrong language), and LLM-as-judge scoring on a 1-5 quality rubric. You should log rejection rates at each filter stage. If you are rejecting more than 40% of examples at any single filter, that filter is either too aggressive or your upstream collection has a systemic problem. Fix the source, not the filter threshold.

Deduplication operates at two levels. Exact deduplication uses MD5 or SHA-256 hashing of the normalized text (lowercased, whitespace-collapsed). Near-duplicate detection requires either MinHash LSH (fast, approximate) or embedding-based clustering (slower, more accurate). The Hugging Face dataskimmer library and the text-dedup package both implement MinHash LSH at scale. The critical insight is that deduplication must be applied across splits, not just within them. If the same customer support ticket appears in both train and test, your test accuracy is measuring memorization. A practical rule: compute fingerprints across your entire raw dataset, then do a stratified train/val/test split on the deduplicated set, never the other way around.

Split strategy matters more than most developers expect. A random 80/10/10 split is fine for generic benchmarking but wrong for production fine-tuning. Your validation set should reflect your actual inference distribution, not a random sample of training data. If you are fine-tuning a model to handle customer support tickets for a SaaS product and you know 30% of real queries are about billing, your validation set should have approximately 30% billing examples. This means building splits by stratified sampling over intent or category labels, not by random index. Your test set should be kept completely blind until you have a candidate model you intend to deploy. Touching it for hyperparameter decisions is test set contamination, even if it feels innocuous.

At scale, dataset preparation becomes a data engineering problem. At 10,000 examples you can run deduplication in memory with a Python script. At 500,000 examples you need a distributed approach: Spark or Ray Data for the pipeline, MinHash LSH for fuzzy deduplication, and a columnar format like Parquet instead of JSONL for fast filtering queries. At millions of examples you are into data-warehouse territory and the dataset preparation pipeline starts to look more like a dbt + DuckDB workflow than a Python script. The cost dimension: synthetic data generation with GPT-4o at $5/1M output tokens for 100,000 examples averaging 200 output tokens each is roughly $100. That is cheap. The human review time to validate even 5% of those examples is not cheap. Budget for both.

Key Takeaways

  • Match your data format exactly to the base model's instruction template to avoid conflicting with pretrained priors.
  • Deduplicate across all splits (train/val/test) before training, not after seeing bad eval numbers.
  • Aim for 1,000–10,000 high-quality examples over massive noisy corpora for most task-specific fine-tunes.
  • Track per-example loss during training to surface outliers that are hurting generalization.

Pro tips

  • Run a per-example loss analysis after the first training epoch. Sort examples by loss descending. The top 1-2% highest-loss examples are almost always formatting errors, label noise, or adversarial outliers that will dominate gradient updates and degrade the whole model.
  • When generating synthetic training data with a frontier model, use a temperature between 0.7 and 1.0 for diversity, but then filter the outputs with a separate lower-temperature judge call. High-temperature generation plus quality filtering beats low-temperature generation every time for dataset variety.
  • Your validation loss curve is not the same as your task metric curve. A validation loss that plateaus does not mean your ROUGE score or accuracy has plateaued. Log both. Stopping on loss alone can leave task performance improvements on the table.
  • The token length distribution of your training data directly shapes the model's output length behavior. If all your training outputs are 50-100 tokens, the fine-tuned model will resist generating longer outputs even when prompted. Sample deliberately across length buckets.

Common pitfalls

  • Mistake: Splitting before deduplicating, letting the same example appear in train and test. Fix: Compute fingerprints across the full raw dataset, deduplicate globally, then split.
  • Mistake: Using the wrong prompt template for the base model, causing the model to treat special tokens as literal text. Fix: Load the tokenizer's chat_template and format all examples through tokenizer.apply_chat_template().
  • Mistake: Building the validation set with random sampling when inference has a known distribution. Fix: Use stratified sampling by intent or category labels to make the validation set representative of production traffic.
  • Mistake: Including examples longer than the model's training context window, which silently truncates assistant turns. Fix: Filter examples where len(tokenizer.encode(full_prompt)) > max_seq_len before training.

When to use real vs synthetic training data

Option Use when Avoid when
Human-labeled gold data Task requires nuanced judgment, subjective quality, or high-stakes correctness (legal, medical, safety). You need more than a few thousand examples and lack budget or annotation infrastructure.
Synthetic data from a frontier model (GPT-4o, Claude 3.5) You have a clear task definition, a small seed set to guide generation, and a quality filter to remove bad outputs. The task requires knowledge the frontier model lacks, or you cannot accept a model trained on another vendor's outputs for legal reasons.
Domain text converted to instruction pairs You have a large corpus of domain documents and want the model to internalize style, terminology, or reasoning patterns. You need precise input-output behavior control; repurposed text rarely produces tight instruction-following examples.
Hybrid (seed human + bulk synthetic) You want quality control plus scale. Use human examples to define the quality bar, synthetic to provide coverage. You cannot afford the compute or API cost to generate and filter a large synthetic corpus.

Code Example

python
# datasets==2.18.0, sentence-transformers not required for this snippet
import json
import hashlib
from pathlib import Path

def deduplicate_jsonl(input_path: str, output_path: str, key: str = "text") -> int:
    """Remove exact duplicates from a JSONL file using MD5 hashing."""
    seen = set()
    kept = []

    for line in Path(input_path).read_text().splitlines():
        record = json.loads(line)
        fingerprint = hashlib.md5(record[key].strip().lower().encode()).hexdigest()
        if fingerprint not in seen:
            seen.add(fingerprint)
            kept.append(record)

    Path(output_path).write_text("\n".join(json.dumps(r) for r in kept))
    return len(kept)

original_count = 1200
final_count = deduplicate_jsonl("raw_train.jsonl", "train_deduped.jsonl", key="text")
print(f"Kept {final_count}/{original_count} examples after deduplication")

How this code works

This Python script's primary role is to ensure a training dataset is clean and efficient by removing exact duplicate examples from a JSONL file. This prevents an AI model from repeatedly seeing the same information, which could lead to overfitting and poorer generalization. The script processes an input file, identifies redundant entries, and then saves only the unique examples to a new output file, reporting the count of kept items.

The deduplicate_jsonl function achieves this by iterating through each line of the input file and parsing it as a JSON record. For each record, it calculates a unique fingerprint using hashlib.md5 on the specified key's content, which defaults to "text". A subtle but critical step here is record[key].strip().lower().encode(). This ensures that content like " Hello World" and "hello world" generates the same fingerprint, effectively making the deduplication case-insensitive and whitespace-agnostic. It stores these fingerprints in a seen set. If a fingerprint is not yet in seen, the record is considered unique, added to a kept list, and its fingerprint is added to seen. Finally, all kept records are written to the output_path.

Production-grade example

Adds retries with backoff, cost logging, graceful degradation on judge failure, and reproducible splits.

python
# production dataset pipeline: dedup + quality filter + split + logging
# requirements: datasets==2.18.0, text-dedup==0.4.0, openai==1.30.0, tenacity==8.3.0

import hashlib, json, logging, os, random, time
from pathlib import Path
from typing import Iterator
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
import openai

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)

client = openai.OpenAI(api_key=os.environ["OPENAI_API_KEY"])

@retry(
    stop=stop_after_attempt(3),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    retry=retry_if_exception_type((openai.RateLimitError, openai.APITimeoutError)),
)
def score_example(instruction: str, output: str, timeout: float = 10.0) -> float:
    """Return a quality score 1-5 for a training example using GPT-4o-mini as judge."""
    prompt = (
        f"Rate the quality of this training example on a scale from 1 (bad) to 5 (excellent).\n"
        f"Instruction: {instruction[:300]}\nOutput: {output[:300]}\n"
        f"Reply with a single integer 1-5 and nothing else."
    )
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        max_tokens=2,
        timeout=timeout,
    )
    token_cost = resp.usage.total_tokens * 0.00000015  # illustrative, verify current pricing
    log.info("judge_call tokens=%d estimated_cost=$%.6f", resp.usage.total_tokens, token_cost)
    try:
        return float(resp.choices[0].message.content.strip())
    except ValueError:
        log.warning("Unexpected judge response: %s", resp.choices[0].message.content)
        return 3.0  # graceful degradation: treat as neutral quality

def normalize(text: str) -> str:
    return " ".join(text.lower().split())

def fingerprint(text: str) -> str:
    return hashlib.md5(normalize(text).encode()).hexdigest()

def build_dataset(
    raw_path: str,
    output_dir: str,
    quality_threshold: float = 3.5,
    val_ratio: float = 0.1,
    test_ratio: float = 0.1,
    score_sample_rate: float = 0.2,  # only judge 20% to control cost
) -> dict:
    records = [json.loads(l) for l in Path(raw_path).read_text().splitlines() if l.strip()]
    log.info("Loaded %d raw records", len(records))

    # --- exact deduplication on (instruction+output) composite key ---
    seen, deduped = set(), []
    for r in records:
        fp = fingerprint(r["instruction"] + r["output"])
        if fp not in seen:
            seen.add(fp)
            deduped.append(r)
    log.info("After dedup: %d records (removed %d)", len(deduped), len(records) - len(deduped))

    # --- quality filtering with LLM judge on a random sample ---
    passed, rejected = [], 0
    for r in deduped:
        if random.random() < score_sample_rate:
            try:
                score = score_example(r["instruction"], r["output"])
                if score < quality_threshold:
                    rejected += 1
                    log.debug("Rejected example (score=%.1f): %s", score, r["instruction"][:60])
                    continue
            except Exception as e:
                log.error("Judge failed, keeping example: %s", e)
        passed.append(r)
    log.info("After quality filter: %d records (rejected %d sampled)", len(passed), rejected)

    # --- stratified split (shuffle with fixed seed for reproducibility) ---
    random.seed(42)
    random.shuffle(passed)
    n = len(passed)
    n_test = max(1, int(n * test_ratio))
    n_val = max(1, int(n * val_ratio))
    splits = {
        "test": passed[:n_test],
        "val": passed[n_test:n_test + n_val],
        "train": passed[n_test + n_val:],
    }

    out = Path(output_dir)
    out.mkdir(parents=True, exist_ok=True)
    for split_name, examples in splits.items():
        path = out / f"{split_name}.jsonl"
        path.write_text("\n".join(json.dumps(e) for e in examples))
        log.info("Wrote %d examples to %s", len(examples), path)

    return {k: len(v) for k, v in splits.items()}

if __name__ == "__main__":
    stats = build_dataset("raw_data.jsonl", "dataset_v1/")
    log.info("Final split sizes: %s", stats)

How this code works

This Python script automates the preparation of a training dataset for fine-tuning, ensuring it's clean, high-quality, and properly split. It first loads raw instruction-output pairs, then performs exact deduplication using an MD5 fingerprint of the normalized combined text to remove identical entries. Next, it applies quality filtering by leveraging gpt-4o-mini via the score_example function, which acts as an LLM judge to rate example quality on a 1-5 scale. Examples falling below a quality_threshold are discarded.

A subtle but important detail is the score_sample_rate: only a fraction (e.g., 20%) of examples are actually sent to the LLM judge to control API costs. This means some potentially low-quality examples might pass if they aren't sampled. Finally, the script performs a fixed-ratio split of the remaining high-quality data into train, val, and test sets after shuffling, ensuring reproducibility with a fixed random.seed(42), before saving each split as a jsonl file.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a dataset preparation script for a customer support fine-tuning task. Load a JSONL file of raw examples (instruction + output fields), apply exact deduplication, filter out examples where the output is fewer than 10 tokens (use simple whitespace splitting as a proxy), then write train/val/test splits at an 80/10/10 ratio. Print a summary of counts at each stage.

python
import json
import hashlib
import random
from pathlib import Path

# TODO: load all records from "raw_support.jsonl"
records = []

# TODO: exact dedup using MD5 of normalized (instruction + output)
deduped = []

# TODO: filter records where output has fewer than 10 whitespace-delimited tokens
filtered = []

# TODO: shuffle with random.seed(42), then split 80/10/10
train, val, test = [], [], []

# TODO: write each split to train.jsonl, val.jsonl, test.jsonl

# TODO: print a summary like:
# Raw: 500 | After dedup: 480 | After filter: 460
# Train: 368 | Val: 46 | Test: 46

Quick check

  1. You split your dataset 80/10/10 randomly, then run deduplication within each split. What is the core problem?

  2. A fine-tuned model's validation loss is decreasing but your task-specific accuracy metric stops improving at epoch 3. What is the most likely explanation?

  3. You are fine-tuning Llama 3 8B with Axolotl and your instruction examples are formatted as plain 'instruction\n\noutput' strings without the Llama 3 chat tokens. What is the most likely consequence?

Self-check: Describe the order of operations for a dataset pipeline (collection, dedup, filter, split) and explain why changing the order of any two adjacent steps could corrupt your evaluation metrics or training signal.