# How to Do Synthetic Data for RAG Evaluation [2026 Pipeline]

> A leakage-first, adversarial-negative pipeline to generate, label, and validate a synthetic RAG evaluation set without building a benchmark that only tests your generator.

- Canonical: https://www.kunalganglani.com/blog/synthetic-data-rag-evaluation
- Author: Kunal Ganglani
- Published: 2026-09-30 · Updated: 2026-09-30
- Category: AI and Machine Learning · Tags: rag, synthetic-data, evals, testing, llmops

## TL;DR

Synthetic data for RAG evaluation is a way to test your RAG chatbot before you have enough real, labeled user questions. The trap is building a “benchmark” that only measures how good your question generator is, not how good your RAG system is. A good pipeline starts from frozen seed docs, generates realistic question intents, adds reference answers with citations, and then stresses retrieval with hard and adversarial negatives. Before trusting any scores, you run leakage checks to catch duplicated docs, near-identical chunks, and prompt echo. Finally, you score retrieval and answer quality separately and do a quick human spot check so you don’t fool yourself.

If you want **synthetic data for RAG evaluation** that actually predicts production quality, you need two things up front: (1) a pipeline that produces *hard* examples (not just easy Q&A), and (2) leakage checks before you trust a single metric. Most teams do the opposite. They generate 500 Q&As, run an LLM-as-a-judge, see a nice score, and ship. Then the bot faceplants the minute real users show up.

This post is the end-to-end pipeline I wish more teams ran. It’s leakage-first, it separates retrieval failures from generation failures, and it includes an adversarial negative cookbook that targets modern stacks (rerankers, long-context, and increasingly agentic retrieval).

Here’s the pipeline we’ll build:

1. Curate and snapshot **seed docs** (with provenance)
1. Generate **question intents** (not just questions)
1. Generate **reference answers + grounding spans**
1. Create **hard negatives** and **adversarial negatives**
1. Run **leakage checks** (doc overlap, near-duplicates, answer-string search, prompt echo)
1. Score with a **RAG evaluation dataset scorecard** (retrieval + generation + end-to-end)
1. Do a **30-minute human spot check** with escalation triggers
1. Version and maintain the set as docs and embeddings change
I’ll be direct about my bias: synthetic evals are great, but they’re also incredibly easy to fool yourself with. Treat them like tests in a distributed system. Assume flakiness, assume contamination, and assume Goodhart’s Law is coming for you.

## What is Synthetic Data for RAG Evaluation

**Synthetic data for RAG evaluation is an automatically generated test dataset (questions, expected answers, and grounding/citations) used to measure a Retrieval-Augmented Generation (RAG) system when you don’t yet have enough labeled real user queries.**

![The word DATA and a star symbol stenciled in dark dots on glass](https://cdn.sanity.io/images/vzekdneq/production/b5c81ac1fe275767b0a8573463ceece18af0a689-1200x675.webp)

You should use it when:

- You’re pre-launch or early launch and have <100 high-quality real queries.
- You’re changing chunking, embeddings, reranking, or prompt templates weekly.
- You need regression gates in CI for [AI in production](/pillars/ai-engineering-production).
You should not pretend it replaces real evaluation when:

- Your app has meaningful user traffic and your query distribution is drifting weekly.
- Your failures are dominated by UX, tool-use, or multi-turn behavior (synthetic single-turn Q&A won’t see that).
One practical rule I use: once you have **200–500** real queries with outcomes you trust, the synthetic set becomes your *regression harness*, not your *north star*.

## Step 1: Build seed docs that don’t poison the dataset

Everything starts with seed docs. If your seed corpus is messy, your synthetic dataset will be messy, and your RAG system will look “good” for the wrong reasons.

![Abstract blue waves of glowing particles](https://cdn.sanity.io/images/vzekdneq/production/a2a16728ed39e06e23e17473a1e4427473237d55-1200x675.webp)

**Snapshot your seed docs.** Pick a commit SHA, export timestamp, or content hash. If your corpus is a wiki or Google Drive, export a frozen dump. “Latest” is not a dataset.

Minimum provenance fields (I store these per document):

- `doc_id` (stable)
- `source` (e.g., Confluence space, S3 bucket, repo path)
- `version` (commit SHA or export timestamp)
- `owner` (team/contact)
- `pii_class` (public/internal/restricted)
If you’re building RAG over internal docs, you also need to think about [AI security](/blog/ai-security-complete-guide) and data governance. Synthetic evals are often generated by an external model API. That’s a data leak waiting to happen.

A concrete number: I aim for **50–200 seed docs** for the first synthetic pass. Less than 50 and your dataset will overfit to a few doc styles. More than 200 and you’ll spend your week debugging the generator instead of your RAG.

Internal links that pair well here:

- [RAG](/glossary/rag)
- retrieval-augmented generation
- [vector embeddings](/glossary/vector-embeddings)
- [vector database](/glossary/vector-database)
- [RAG context window limits](/blog/rag-context-window-limitations)
## Step 2: Generate realistic questions without baking in the rubric

The #1 failure mode in synthetic RAG test set generation is that you accidentally generate *questions that your own generator knows how to answer*. That’s not evaluation. That’s self-congratulation.

![Openai logo surrounded by abstract data network elements](https://cdn.sanity.io/images/vzekdneq/production/164458d795ec2b752acaf750c599d4c4493c3f7d-1200x675.webp)

I generate questions in two phases:

### Phase A: Generate intent buckets first

Create **6–8 intent buckets** that match real usage. Example buckets:

1. Direct fact lookup (“What’s the refund window?”)
1. Procedure (“How do I rotate API keys?”)
1. Troubleshooting (“Why is X failing after Y?”)
1. Policy / compliance (“Is this data allowed in logs?”)
1. Comparison (“A vs B, when to use which?”)
1. Edge case (“What happens if a user does Z?”)
1. Version/time-sensitive (“What changed after 2026-06?”)
This forces coverage. If you just ask an LLM to “generate 200 questions”, you’ll get 200 variations of the same bland FAQ.

### Phase B: Generate questions with model separation

To keep synthetic evals honest, don’t use the same model for:

- question generation
- answer generation
- judging
Use at least **two different model families** across the pipeline. If you can’t, at least randomize prompts and temperatures.

Also, do not let the generator see your scoring rubric verbatim. If you tell it “make questions that test faithfulness and citation quality”, it will obediently produce questions that are easy to cite.

This is where I like to borrow an idea from classic benchmarking culture: evaluation needs adversarial thinking. The Hugging Face team’s write-up on benchmark variance is a good reminder that “numbers move when prompts move” and that reproducibility is fragile. See [Thomas Wolf](https://huggingface.co/blog/evaluating-mmlu-leaderboard) and the Hugging Face evaluation analysis for how evaluation setups drift.

## Step 3: Generate reference answers *and* grounding spans

A synthetic Q without a grounded reference answer is just trivia.

Your dataset row should include:

- `question`
- `reference_answer`
- `supporting_chunks` (doc IDs + chunk IDs)
- `grounding_spans` (exact text spans or sentence indices)
- `answer_type` (extractive vs abstractive vs procedural)
Why spans matter: they let you score retrieval separately from generation. If your retrieved context contains the span, and the model still hallucinates, you’ve got a generation/prompt issue. If the span never shows up, it’s retrieval.

Two practical constraints that improve quality:

- Keep reference answers short. I target **40–120 tokens**. Long references let judges “agree” with vague similarity.
- Require **2 citations** for multi-part answers. If a question has two clauses, force grounding from two different chunks.
If you want a ready-made framework to run this kind of evaluation harness, OpenAI’s repository is still one of the better public reference points: [OpenAI Evals](https://github.com/openai/evals) (it’s a framework plus a registry of evals).

## Step 4: Adversarial and hard negatives (the cookbook)

Most RAG eval sets are all positives. That makes retrieval look great.

A real system fails because it retrieves the *wrong* plausible chunk, or the generator confidently answers from partial context.

Here’s the negative strategy that’s worked best for me.

### Hard negatives (semantic neighbors)

For each question, find chunks with high embedding similarity that are *not* the true supporting chunks.

Label them:

- `hard_negative_retrieval`: semantically similar but does not contain grounding spans
- `hard_negative_generation`: contains related info but contradicts the correct answer
Concrete ratio: per positive row, generate **3 hard negatives**. With fewer, your reranker won’t get stressed.

### Adversarial negatives (designed to break rerankers)

These are hand-crafted patterns you can still generate programmatically:

1. **Entity swaps**: same template, different product/feature name.
1. **Time/version traps**: “as of 2025” vs “as of 2026-08” docs.
1. **Policy exceptions**: chunk describes general rule; negative chunk contains exception.
1. **Multi-hop distractors**: two chunks individually look relevant but only one is correct when combined.
1. **Instructional injection**: chunks that contain “ignore previous instructions” style garbage to see if your pipeline is robust.
If you’re building agentic retrieval or tool-using flows, add negatives that target the *tool call step*, not just retrieval. I cover more of that testing style in [agent tool call failure testing](/blog/agent-tool-call-failure-testing) and [non-deterministic AI system testing](/blog/non-deterministic-ai-testing).

## Step 5: Leakage checks (train/test overlap, duplication, prompt echo)

This is the differentiator most tutorials skip. It’s also why their scores don’t transfer.

Run these checks before scoring.

### 1) Document overlap and near-duplicate chunks

You’re looking for:

- same doc in train and test (obvious)
- near-duplicate chunks across docs (common in wikis)
If your vector store contains duplicates, you can “improve” recall@k by accident because the same content appears 3 times.

A simple heuristic I’ve used: compute MinHash / SimHash fingerprints and flag pairs above **0.9 similarity** for review.

### 2) Answer-string search against the corpus

If your reference answer string exists verbatim in many chunks, your “generation quality” metric may really be “copy exact sentence”. That’s fine, but know what you’re measuring.

At minimum: search for the longest 8–12 word n-gram from the reference answer and flag if it appears in >**5** chunks.

### 3) Prompt echo detection

If your generator prompt includes phrases like “Answer with citations, be faithful”, you’ll see those phrases leak into outputs. Then your judge rewards the style it was told to prefer.

Fix: rotate prompt templates and add a detector that flags answers containing your instruction phrases.

### 4) Memorized answers (model contamination)

This is rare for private docs, common for public corpora. If your seed docs are public (docs, product manuals), the model may already “know” the answer.

Mitigation: include a small set of questions whose answers are intentionally absent from the corpus. If the model answers confidently anyway, you have a hallucination and contamination problem.

If you care about broader leakage thinking beyond evals, my [LLM data leakage playbook](/blog/llm-data-leakage-playbook) and [RAG data leakage test suite](/blog/rag-data-leakage-test-suite) are the mindset.

## Step 6: Score with a scorecard that separates retrieval vs generation

If your scorecard is one number, it’s useless.

You want a table that forces you to diagnose.

Here’s the template I use (edit it to fit your stack):

| Category | Metric | What it isolates | Target (example) | Failure meaning |
| --- | --- | --- | --- | --- |
| Retrieval | recall@k (k=5) | Did we fetch any supporting chunk? | ≥ 0.85 | Retriever/chunking/embeddings |
| Retrieval | MRR@10 | Ranking quality | ≥ 0.65 | Reranker or index noise |
| Retrieval | context precision | How much retrieved context is relevant | ≥ 0.70 | Over-retrieval, bad filters |
| Generation | faithfulness | Did answer stay within context? | ≥ 0.80 | Prompt, model, tool config |
| Generation | completeness | Did it cover required points? | ≥ 0.75 | Missing context or weak synthesis |
| End-to-end | exact/semantic match | User-visible correctness | ≥ 0.75 | Anything above |
| End-to-end | citation coverage | Claims backed by citations | ≥ 0.90 | Grounding/citation extractor |

For metric definitions, I’m not reinventing the wheel. Ragas has a solid catalog of retrieval and generation metrics, including context precision/recall and faithfulness. See the official docs: [Ragas documentation](https://docs.ragas.io/).

One important production note: set targets per intent bucket. A troubleshooting question is harder than an FAQ lookup. If you average them, you’ll hide pain.

A data anchor from my own work: on the Walmart conversational commerce chatbot I built at Firework (Zealsight), retrieval quality dominated answer quality at scale. We were handling **millions of queries daily** at **sub-second** response times, and model swaps moved metrics far less than fixing retrieval and chunking.

## Step 7: 30 minutes of human review (the sampling plan)

If you only do one human thing, do this.

Every evaluation run, sample **30 rows** total:

- 10 random
- 10 from the worst-scoring bucket
- 10 from adversarial negatives
Timebox: **30 minutes**. Two reviewers if you can.

Acceptance thresholds that catch issues fast:

- If >**3/30** have wrong citations, stop and fix grounding.
- If >**5/30** are “judge says pass, human says fail”, your judge is miscalibrated.
- If any “absent-answer” test is answered confidently, investigate hallucination controls.
This is the fastest way to detect the most dangerous failure mode: your LLM judge agreeing with your generator’s writing style.

If you want a broader weekly cadence for evals, the same rhythm applies to [agent evaluation roadmap for small teams](/blog/agent-evaluation-roadmap-teams) and [AI engineering evals regression gates](/blog/ai-engineering-evals-gates).

Here’s a practical walkthrough video that matches the “build a synthetic RAG test set” workflow (useful as a demo reference, not as a rigor reference):

[Watch: Boost Your RAG Chatbot’s Accuracy with Synthetic Test Data Generation!](https://www.youtube.com/watch?v=7rba4exIaOs)

## Step 8: Keep synthetic evals honest as your RAG stack evolves

RAG stacks in 2026 are not just “retrieve + stuff into a prompt”. They’re using rerankers, query rewriting, multi-step retrieval, and sometimes [AI agents](/pillars/ai-agents) that call tools.

That means your synthetic eval dataset needs maintenance like production code.

What I version:

- seed doc snapshot ID
- chunking config (size, overlap, separators)
- embedding model + dimension
- index params (HNSW ef/search_k, filters)
- reranker model
- prompt template versions
- judge model + rubric version
How often I refresh:

- regenerate questions monthly or after major doc style changes
- regenerate negatives whenever you change embeddings/reranker
- rerun leakage checks every time you rebuild the vector index
Concrete example: if you cache embeddings, you can accidentally evaluate against stale chunks. Your system looks stable, but your users are reading updated docs. Tie your eval run to the same embedding snapshot used in prod.

Two related internal posts worth keeping close:

- [RAG evaluation metrics for retrieval quality](/blog/rag-evaluation-metrics-retrieval-quality)
- [How to use MLflow LLM evaluation tracing](/blog/mlflow-evaluation-tracing-gates)
My prediction: as more teams adopt agentic retrieval, the most valuable synthetic evals won’t be Q&A at all. They’ll be **tool-call traces** with adversarial environment states. If your “RAG eval” doesn’t model the environment your agent operates in, you’re testing the wrong system.

Photo by PiggyBank on Unsplash.
