An evaluation dataset is not a benchmark you download from a paper. It is a purpose-built artifact tied to your application. A customer-support bot for a telecom company needs eval cases drawn from actual support tickets, not from MMLU or SQuAD. The mental model is this: your dataset samples from the distribution of inputs your system will actually receive in production. If that distribution shifts -- new product lines, seasonal queries, changing user demographics -- your dataset must evolve with it. Think of it less like a static test suite and more like a living sample of your users.
The anatomy of an eval case is (id, input, expected_output, metadata). The id field is non-negotiable -- you need it to track individual case performance over time and correlate it back to a source ticket, user session, or annotation round. The input is the exact prompt context you feed to the model (including system prompt if it's fixed). The expected_output is the gold-standard answer, which can be a verbatim string, a structured object, a reference answer for semantic comparison, or a set of required facts. Metadata is where you record difficulty tier, topic category, data source, and annotator ID -- all of which become essential when you need to slice scores by segment.
Building a rubric is where teams consistently underinvest. A rubric is a mapping from output properties to numeric scores. The key design constraint is that each criterion must be independently and unambiguously scorable. "Quality" is not a criterion. "The response correctly states the account balance mentioned in the context" is a criterion. For most application evals you will combine several criterion types: binary pass/fail (factual accuracy, safety compliance), ordinal scale (1-5 for relevance or fluency), and reference-free judgments (does the output contain a specific required element regardless of phrasing). Write your rubric criteria as if you were training a new human annotator who has never seen your product; if two annotators would disagree on what score to assign, the criterion needs to be tightened.
The workflow for constructing the initial dataset follows a specific sequence. Start with a production log sample if your system is live, or seed cases from domain experts if you are pre-launch. Annotate expected outputs using subject matter experts for high-stakes criteria (medical, legal, financial), or use a senior engineer's judgment for lower-stakes criteria. Compute inter-annotator agreement (Cohen's kappa or Krippendorff's alpha) on a 50-case overlap set. If agreement is below roughly 0.7, your rubric needs refinement before you scale annotation. Once you have 150 to 300 well-annotated cases, split them: 70% for iterative development (safe to look at), 30% held-out (only run against this when you are about to ship). Overfitting prompts to the development split is a real failure mode -- every time you run your held-out set you are spending a limited budget of unbiased signal.
Tradeoffs versus alternative approaches are worth being explicit about. Hand-crafted datasets with human annotation give you high-quality signal but cost real money and time -- budget roughly 2 to 5 minutes per case for simple binary scoring, 10 to 20 minutes per case for detailed rubric annotation. Synthetic generation using an LLM (GPT-4o generates eval cases from your docs) is faster and cheaper but introduces the same biases your generation model has, which means you are not independently auditing your pipeline. The right answer for most teams is a hybrid: use LLM generation for volume and diversity, then human-review a stratified sample of 15-20% to catch systematic errors in the generated ground truth.
At scale, the dataset management problem becomes infrastructure. At 10 users you keep everything in a JSON file in the repo. At 10k users you have enough production data to run automated sampling pipelines that pull edge cases and failures into a staging dataset for annotation. At 10M users the bottleneck is annotation throughput, not data volume -- teams at this scale build internal annotation tools, tiered annotation queues (easy cases auto-scored, ambiguous cases routed to humans), and version-controlled dataset stores (tools like Argilla, Label Studio, or proprietary internal platforms). What does not change at any scale is the fundamental principle: every case must have a traceable provenance, every rubric change must be versioned, and your held-out set must stay truly held-out.
Key Takeaways
- Build eval datasets from real production queries, not synthetic ones you invented at your desk.
- Every rubric criterion must be independently scorable; ambiguous criteria produce noisy, useless scores.
- Separate your 'development' and 'held-out' eval sets to avoid overfitting your prompts to known cases.
- Store datasets and rubrics in version control alongside prompt files so changes stay traceable.
Pro tips
- Track inter-annotator agreement before scaling annotation. If two humans scoring the same 50 cases agree less than ~70% of the time (kappa < 0.7), your rubric criterion is ambiguous and will produce garbage signal at scale.
- Keep a 'failure taxonomy' doc alongside your dataset. Every time a case fails, tag the failure type (hallucination, format violation, off-topic, wrong entity, etc.). After 20 failures you will see patterns that reveal whether the problem is the prompt, the model, or the data.
- Version your eval dataset in git with the same discipline as your code. A dataset that changes without a commit message makes it impossible to know whether a score improvement came from a better prompt or from easier eval cases.
- Never debug a failing eval case by tweaking the prompt until that case passes. That is overfitting. Instead, understand the failure class, fix the underlying prompt pattern, and verify the fix generalizes to held-out cases you have not looked at.
Common pitfalls
- Mistake: Building eval cases that are too easy or all from the happy path. Fix: Deliberately include adversarial inputs, edge cases, and common user mistakes -- aim for ~20% of cases to be failure-inducing.
- Mistake: Using the same dataset for prompt iteration and final reporting, so scores inflate over time. Fix: Maintain a held-out split you run only before shipping; never iterate prompts against it.
- Mistake: Writing rubric criteria like 'response is helpful' with no operationalization. Fix: Define each criterion as a testable question: 'Does the response include the exact account number from the context?'
- Mistake: Generating 100% of ground truth with the same LLM you are evaluating. Fix: Human-review at least a 15% stratified sample to catch systematic errors the generator baked in.
When to use different ground-truth strategies
| Option | Use when | Avoid when |
|---|---|---|
| Verbatim match | Expected output is a fixed value: a code snippet, SQL query, specific date, or account number. | Output is natural language where multiple valid phrasings exist; will produce many false negatives. |
| Reference-fact checklist | You need the response to contain several specific facts but don't care about exact phrasing or order. | Facts are hard to isolate from prose or require complex reasoning to verify programmatically. |
| Human annotation with rubric | Quality dimensions are subjective (tone, empathy, brand voice) or domain expertise is required to judge accuracy. | Budget or time is constrained and the criteria can be operationalized programmatically instead. |
| LLM-as-judge against rubric | You need scale and automation and your rubric criteria are well-defined enough to prompt a judge model. | The task requires real-world factual verification the judge model cannot reliably perform; use human experts there. |
| Semantic similarity (embedding cosine) | Paraphrase detection or semantic equivalence matters and you have a single canonical reference answer. | Factual precision is critical; semantically similar sentences can still contain wrong facts. |
Code Example
# openai>=1.0.0, pydantic>=2.0
import json
from pydantic import BaseModel
class EvalCase(BaseModel):
id: str
input: str
expected_output: str
metadata: dict = {}
def load_dataset(path: str) -> list[EvalCase]:
with open(path) as f:
raw = json.load(f)
return [EvalCase(**item) for item in raw]
# Simple binary rubric: does the response contain the expected answer?
def score_exact_match(response: str, expected: str) -> float:
return 1.0 if expected.strip().lower() in response.strip().lower() else 0.0
if __name__ == "__main__":
cases = load_dataset("eval_dataset.json")
scores = []
for case in cases:
# Placeholder: replace with real LLM call
response = "The capital of France is Paris."
score = score_exact_match(response, case.expected_output)
scores.append(score)
print(f"{case.id}: {score}")
print(f"Mean score: {sum(scores)/len(scores):.2f}")How this code works
This code demonstrates a fundamental process for evaluating Large Language Models (LLMs) using an evaluation dataset and a scoring rubric. It begins by defining a EvalCase Pydantic model to structure each test example, ensuring every case has an id, an input (the prompt sent to the LLM), and an expected_output for comparison. The load_dataset function then reads a JSON file containing these test cases, converting the raw JSON entries into a list of structured EvalCase objects. This makes it easy to access specific parts of each test. Finally, the score_exact_match function acts as a simple rubric, determining if an LLM's response perfectly contains the expected_output.
The main part of the script, enclosed in if __name__ == "__main__":, puts these components into action. It loads the cases and then loops through each one. Inside the loop, it currently uses a placeholder response but would ideally make a real LLM call with case.input. The score_exact_match function is then called to compare this response against case.expected_output, yielding a score of 1.0 for a match or 0.0 otherwise. A subtle but important detail is how score_exact_match uses .strip().lower() on both the response and expected output. This ensures the comparison ignores differences in capitalization and accidental whitespace, making the matching more robust and preventing common types of scoring errors. The script concludes by calculating and printing the Mean score across all cases.
Production-grade example
Adds retries with backoff, per-case latency and token logging, structured logs, validation, and timeout.
# openai>=1.0.0, tenacity>=8.2, pydantic>=2.0, structlog>=23.0
import os
import time
import json
import structlog
from pydantic import BaseModel, ValidationError
from openai import OpenAI, RateLimitError, APITimeoutError, APIConnectionError
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
log = structlog.get_logger()
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=15.0)
class EvalCase(BaseModel):
id: str
input: str
expected_output: str
metadata: dict = {}
class ScoredResult(BaseModel):
case_id: str
response: str
score: float
latency_ms: float
prompt_tokens: int
completion_tokens: int
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError, APIConnectionError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
)
def call_model(prompt: str, system: str) -> tuple[str, int, int]:
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "system", "content": system}, {"role": "user", "content": prompt}],
temperature=0,
max_tokens=512,
)
usage = resp.usage
return resp.choices[0].message.content, usage.prompt_tokens, usage.completion_tokens
def score_contains(response: str, expected: str) -> float:
return 1.0 if expected.strip().lower() in response.strip().lower() else 0.0
def run_eval(dataset_path: str, system_prompt: str) -> list[ScoredResult]:
with open(dataset_path) as f:
raw = json.load(f)
try:
cases = [EvalCase(**item) for item in raw]
except ValidationError as e:
log.error("dataset_validation_failed", error=str(e))
raise
results = []
for case in cases:
t0 = time.monotonic()
try:
response, pt, ct = call_model(case.input, system_prompt)
except Exception as e:
log.error("model_call_failed", case_id=case.id, error=str(e))
continue
latency_ms = (time.monotonic() - t0) * 1000
score = score_contains(response, case.expected_output)
result = ScoredResult(
case_id=case.id, response=response, score=score,
latency_ms=latency_ms, prompt_tokens=pt, completion_tokens=ct,
)
log.info("eval_case_scored", **result.model_dump())
results.append(result)
scores = [r.score for r in results]
total_tokens = sum(r.prompt_tokens + r.completion_tokens for r in results)
log.info("eval_run_complete", mean_score=sum(scores)/len(scores), total_tokens=total_tokens, n_cases=len(results))
return resultsHow this code works
This code creates a robust system for evaluating Large Language Models (LLMs) against a set of predefined test cases. Its primary job is to take a dataset of EvalCases, each with an input and an expected_output, send the input to an LLM, and then score the LLM's response based on how well it matches the expected_output. The run_eval function orchestrates this, collecting detailed ScoredResults including score, latency_ms, and prompt_tokens/completion_tokens for each interaction, before logging overall performance statistics like mean_score.
The process begins by loading evaluation cases using pydantic.BaseModel for validation. For each case, call_model interacts with the OpenAI API, sending a system prompt and the case.input. A subtle but critical detail is the @retry decorator on call_model: it automatically re-attempts API calls that fail due to common issues like RateLimitError or APITimeoutError, ensuring the evaluation doesn't prematurely halt on transient errors by waiting wait_exponential times between attempts. After getting a response, the score_contains function acts as a simple rubric, checking if the expected_output is present in the response. All results are then logged and aggregated.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a minimal eval harness for a question-answering bot. Load a JSON dataset of 5 cases, run each input through the OpenAI chat completions API with a fixed system prompt, score each response using a 'required facts' rubric (all listed facts must appear in the response), and print per-case scores plus the mean.
# eval_harness.py
# openai>=1.0.0
import json, os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
SYSTEM_PROMPT = "You are a concise Q&A assistant. Answer factually."
# TODO: Load this file from disk using json.load()
DATASET = [
{"id": "q1", "input": "What is the boiling point of water?", "required_facts": ["100", "celsius"]},
{"id": "q2", "input": "Who wrote Romeo and Juliet?", "required_facts": ["shakespeare"]},
{"id": "q3", "input": "What gas do plants absorb?", "required_facts": ["carbon dioxide", "co2"]},
{"id": "q4", "input": "How many sides does a hexagon have?", "required_facts": ["6", "six"]},
{"id": "q5", "input": "What is the speed of light?", "required_facts": ["299", "km/s"]},
]
def call_model(user_input: str) -> str:
# TODO: Call client.chat.completions.create and return content string
pass
def score_required_facts(response: str, required_facts: list[str]) -> float:
# TODO: Return fraction of required_facts present in response (case-insensitive)
pass
if __name__ == "__main__":
scores = []
for case in DATASET:
# TODO: call call_model, then score_required_facts, print result
pass
# TODO: print mean scoreQuick check
You iterate your system prompt against your eval dataset 30 times and the score reaches 0.92. What is the main validity concern?
A rubric criterion reads: 'The response is high quality.' Why is this problematic?
You need to evaluate whether a customer-support bot's responses contain the correct account balance. Which ground-truth strategy fits best?