How to Do Synthetic Data for RAG Evaluation [2026 Pipeline]

A leakage-first, adversarial-negative pipeline to generate, label, and validate a synthetic RAG evaluation set without building a benchmark that only tests your generator.

Part of theAI in Production series
a man using a laptop computer on a table

If you want synthetic data for RAG evaluation that actually predicts production quality, you need two things up front: (1) a pipeline that produces hard examples (not just easy Q&A), and (2) leakage checks before you trust a single metric. Most teams do the opposite. They generate 500 Q&As, run an LLM-as-a-judge, see a nice score, and ship. Then the bot faceplants the minute real users show up.

This post is the end-to-end pipeline I wish more teams ran. It’s leakage-first, it separates retrieval failures from generation failures, and it includes an adversarial negative cookbook that targets modern stacks (rerankers, long-context, and increasingly agentic retrieval).

Here’s the pipeline we’ll build:

  1. Curate and snapshot seed docs (with provenance)
  2. Generate question intents (not just questions)
  3. Generate reference answers + grounding spans
  4. Create hard negatives and adversarial negatives
  5. Run leakage checks (doc overlap, near-duplicates, answer-string search, prompt echo)
  6. Score with a RAG evaluation dataset scorecard (retrieval + generation + end-to-end)
  7. Do a 30-minute human spot check with escalation triggers
  8. Version and maintain the set as docs and embeddings change

I’ll be direct about my bias: synthetic evals are great, but they’re also incredibly easy to fool yourself with. Treat them like tests in a distributed system. Assume flakiness, assume contamination, and assume Goodhart’s Law is coming for you.

What is Synthetic Data for RAG Evaluation

Synthetic data for RAG evaluation is an automatically generated test dataset (questions, expected answers, and grounding/citations) used to measure a Retrieval-Augmented Generation (RAG) system when you don’t yet have enough labeled real user queries.

The word DATA and a star symbol stenciled in dark dots on glass

You should use it when:

  • You’re pre-launch or early launch and have <100 high-quality real queries.
  • You’re changing chunking, embeddings, reranking, or prompt templates weekly.
  • You need regression gates in CI for AI in production.

You should not pretend it replaces real evaluation when:

  • Your app has meaningful user traffic and your query distribution is drifting weekly.
  • Your failures are dominated by UX, tool-use, or multi-turn behavior (synthetic single-turn Q&A won’t see that).

One practical rule I use: once you have 200–500 real queries with outcomes you trust, the synthetic set becomes your regression harness, not your north star.

Step 1: Build seed docs that don’t poison the dataset

Everything starts with seed docs. If your seed corpus is messy, your synthetic dataset will be messy, and your RAG system will look “good” for the wrong reasons.

Abstract blue waves of glowing particles

Snapshot your seed docs. Pick a commit SHA, export timestamp, or content hash. If your corpus is a wiki or Google Drive, export a frozen dump. “Latest” is not a dataset.

Minimum provenance fields (I store these per document):

  • doc_id (stable)
  • source (e.g., Confluence space, S3 bucket, repo path)
  • version (commit SHA or export timestamp)
  • owner (team/contact)
  • pii_class (public/internal/restricted)

If you’re building RAG over internal docs, you also need to think about AI security and data governance. Synthetic evals are often generated by an external model API. That’s a data leak waiting to happen.

A concrete number: I aim for 50–200 seed docs for the first synthetic pass. Less than 50 and your dataset will overfit to a few doc styles. More than 200 and you’ll spend your week debugging the generator instead of your RAG.

Internal links that pair well here:

Step 2: Generate realistic questions without baking in the rubric

The #1 failure mode in synthetic RAG test set generation is that you accidentally generate questions that your own generator knows how to answer. That’s not evaluation. That’s self-congratulation.

Openai logo surrounded by abstract data network elements

I generate questions in two phases:

Phase A: Generate intent buckets first

Create 6–8 intent buckets that match real usage. Example buckets:

  1. Direct fact lookup (“What’s the refund window?”)
  2. Procedure (“How do I rotate API keys?”)
  3. Troubleshooting (“Why is X failing after Y?”)
  4. Policy / compliance (“Is this data allowed in logs?”)
  5. Comparison (“A vs B, when to use which?”)
  6. Edge case (“What happens if a user does Z?”)
  7. Version/time-sensitive (“What changed after 2026-06?”)

This forces coverage. If you just ask an LLM to “generate 200 questions”, you’ll get 200 variations of the same bland FAQ.

Phase B: Generate questions with model separation

To keep synthetic evals honest, don’t use the same model for:

  • question generation
  • answer generation
  • judging

Use at least two different model families across the pipeline. If you can’t, at least randomize prompts and temperatures.

Also, do not let the generator see your scoring rubric verbatim. If you tell it “make questions that test faithfulness and citation quality”, it will obediently produce questions that are easy to cite.

This is where I like to borrow an idea from classic benchmarking culture: evaluation needs adversarial thinking. The Hugging Face team’s write-up on benchmark variance is a good reminder that “numbers move when prompts move” and that reproducibility is fragile. See Thomas Wolf and the Hugging Face evaluation analysis for how evaluation setups drift.

Step 3: Generate reference answers *and* grounding spans

A synthetic Q without a grounded reference answer is just trivia.

Your dataset row should include:

  • question
  • reference_answer
  • supporting_chunks (doc IDs + chunk IDs)
  • grounding_spans (exact text spans or sentence indices)
  • answer_type (extractive vs abstractive vs procedural)

Why spans matter: they let you score retrieval separately from generation. If your retrieved context contains the span, and the model still hallucinates, you’ve got a generation/prompt issue. If the span never shows up, it’s retrieval.

Two practical constraints that improve quality:

  • Keep reference answers short. I target 40–120 tokens. Long references let judges “agree” with vague similarity.
  • Require 2 citations for multi-part answers. If a question has two clauses, force grounding from two different chunks.

If you want a ready-made framework to run this kind of evaluation harness, OpenAI’s repository is still one of the better public reference points: OpenAI Evals (it’s a framework plus a registry of evals).

Step 4: Adversarial and hard negatives (the cookbook)

Most RAG eval sets are all positives. That makes retrieval look great.

A real system fails because it retrieves the wrong plausible chunk, or the generator confidently answers from partial context.

Here’s the negative strategy that’s worked best for me.

Hard negatives (semantic neighbors)

For each question, find chunks with high embedding similarity that are not the true supporting chunks.

Label them:

  • hard_negative_retrieval: semantically similar but does not contain grounding spans
  • hard_negative_generation: contains related info but contradicts the correct answer

Concrete ratio: per positive row, generate 3 hard negatives. With fewer, your reranker won’t get stressed.

Adversarial negatives (designed to break rerankers)

These are hand-crafted patterns you can still generate programmatically:

  1. Entity swaps: same template, different product/feature name.
  2. Time/version traps: “as of 2025” vs “as of 2026-08” docs.
  3. Policy exceptions: chunk describes general rule; negative chunk contains exception.
  4. Multi-hop distractors: two chunks individually look relevant but only one is correct when combined.
  5. Instructional injection: chunks that contain “ignore previous instructions” style garbage to see if your pipeline is robust.

If you’re building agentic retrieval or tool-using flows, add negatives that target the tool call step, not just retrieval. I cover more of that testing style in agent tool call failure testing and non-deterministic AI system testing.

Step 5: Leakage checks (train/test overlap, duplication, prompt echo)

This is the differentiator most tutorials skip. It’s also why their scores don’t transfer.

Run these checks before scoring.

1) Document overlap and near-duplicate chunks

You’re looking for:

  • same doc in train and test (obvious)
  • near-duplicate chunks across docs (common in wikis)

If your vector store contains duplicates, you can “improve” recall@k by accident because the same content appears 3 times.

A simple heuristic I’ve used: compute MinHash / SimHash fingerprints and flag pairs above 0.9 similarity for review.

2) Answer-string search against the corpus

If your reference answer string exists verbatim in many chunks, your “generation quality” metric may really be “copy exact sentence”. That’s fine, but know what you’re measuring.

At minimum: search for the longest 8–12 word n-gram from the reference answer and flag if it appears in >5 chunks.

3) Prompt echo detection

If your generator prompt includes phrases like “Answer with citations, be faithful”, you’ll see those phrases leak into outputs. Then your judge rewards the style it was told to prefer.

Fix: rotate prompt templates and add a detector that flags answers containing your instruction phrases.

4) Memorized answers (model contamination)

This is rare for private docs, common for public corpora. If your seed docs are public (docs, product manuals), the model may already “know” the answer.

Mitigation: include a small set of questions whose answers are intentionally absent from the corpus. If the model answers confidently anyway, you have a hallucination and contamination problem.

If you care about broader leakage thinking beyond evals, my LLM data leakage playbook and RAG data leakage test suite are the mindset.

Step 6: Score with a scorecard that separates retrieval vs generation

If your scorecard is one number, it’s useless.

You want a table that forces you to diagnose.

Here’s the template I use (edit it to fit your stack):

CategoryMetricWhat it isolatesTarget (example)Failure meaning
Retrievalrecall@k (k=5)Did we fetch any supporting chunk?≥ 0.85Retriever/chunking/embeddings
RetrievalMRR@10Ranking quality≥ 0.65Reranker or index noise
Retrievalcontext precisionHow much retrieved context is relevant≥ 0.70Over-retrieval, bad filters
GenerationfaithfulnessDid answer stay within context?≥ 0.80Prompt, model, tool config
GenerationcompletenessDid it cover required points?≥ 0.75Missing context or weak synthesis
End-to-endexact/semantic matchUser-visible correctness≥ 0.75Anything above
End-to-endcitation coverageClaims backed by citations≥ 0.90Grounding/citation extractor

For metric definitions, I’m not reinventing the wheel. Ragas has a solid catalog of retrieval and generation metrics, including context precision/recall and faithfulness. See the official docs: Ragas documentation.

One important production note: set targets per intent bucket. A troubleshooting question is harder than an FAQ lookup. If you average them, you’ll hide pain.

A data anchor from my own work: on the Walmart conversational commerce chatbot I built at Firework (Zealsight), retrieval quality dominated answer quality at scale. We were handling millions of queries daily at sub-second response times, and model swaps moved metrics far less than fixing retrieval and chunking.

Step 7: 30 minutes of human review (the sampling plan)

If you only do one human thing, do this.

Every evaluation run, sample 30 rows total:

  • 10 random
  • 10 from the worst-scoring bucket
  • 10 from adversarial negatives

Timebox: 30 minutes. Two reviewers if you can.

Acceptance thresholds that catch issues fast:

  • If >3/30 have wrong citations, stop and fix grounding.
  • If >5/30 are “judge says pass, human says fail”, your judge is miscalibrated.
  • If any “absent-answer” test is answered confidently, investigate hallucination controls.

This is the fastest way to detect the most dangerous failure mode: your LLM judge agreeing with your generator’s writing style.

If you want a broader weekly cadence for evals, the same rhythm applies to agent evaluation roadmap for small teams and AI engineering evals regression gates.

Here’s a practical walkthrough video that matches the “build a synthetic RAG test set” workflow (useful as a demo reference, not as a rigor reference):

Step 8: Keep synthetic evals honest as your RAG stack evolves

RAG stacks in 2026 are not just “retrieve + stuff into a prompt”. They’re using rerankers, query rewriting, multi-step retrieval, and sometimes AI agents that call tools.

That means your synthetic eval dataset needs maintenance like production code.

What I version:

  • seed doc snapshot ID
  • chunking config (size, overlap, separators)
  • embedding model + dimension
  • index params (HNSW ef/search_k, filters)
  • reranker model
  • prompt template versions
  • judge model + rubric version

How often I refresh:

  • regenerate questions monthly or after major doc style changes
  • regenerate negatives whenever you change embeddings/reranker
  • rerun leakage checks every time you rebuild the vector index

Concrete example: if you cache embeddings, you can accidentally evaluate against stale chunks. Your system looks stable, but your users are reading updated docs. Tie your eval run to the same embedding snapshot used in prod.

Two related internal posts worth keeping close:

My prediction: as more teams adopt agentic retrieval, the most valuable synthetic evals won’t be Q&A at all. They’ll be tool-call traces with adversarial environment states. If your “RAG eval” doesn’t model the environment your agent operates in, you’re testing the wrong system.

Photo by PiggyBank on Unsplash.

Continue reading

A digital dashboard displaying marketing metrics including CTR and quality score on a screen

How to Do Non Deterministic AI System Testing [2026]

A release-gating playbook for LLMs and agents: golden sets, metamorphic tests, variance control, tolerance bands, cohort diffs, and eval-regression postmortems.

a computer screen with a bunch of data on it

RAG Evaluation Metrics for Retrieval Quality: My Production Playbook

If your RAG app got worse after an embeddings or chunking change, grading answers won’t tell you why. Here’s a retrieval-first eval workflow: recall@k, MRR, citation accuracy, leakage tests, and a frozen offline corpus so regressions are real—not vibes.

a clipboard with a checklist on it next to a cup of coffee and

Agent Evaluation Roadmap for Small Teams [2026]: The 30-Min/Week Plan

A pragmatic map of offline vs online vs HITL evals, what to measure beyond “response quality,” and the smallest program that actually prevents agent regressions.

a person typing on a laptop keyboard on a desk

AI Engineering Evals: Regression Gates for Prompts, Tools, RAG [2026]

Stop letting prompt tweaks and model upgrades silently break production. Here’s a CI-style regression gate system for prompts, tool calling, and RAG with golden sets, schemas, shadow evals, and failure budgets.

Cite this article
Kunal Ganglani (2026, September 30). How to Do Synthetic Data for RAG Evaluation [2026 Pipeline]. Kunal Ganglani. Retrieved September 30, 2026, from https://www.kunalganglani.com/blog/synthetic-data-rag-evaluation