Mental model: what you are actually measuring
When you evaluate a fine-tuned model, you are answering one question: does the model's behavior on the target task improve enough, relative to the baseline, to justify the full lifecycle cost of running it? That means you need three things locked down before training ends: (1) a held-out test set that was never seen during training or validation, (2) a baseline that represents your current best option without fine-tuning, and (3) one or two primary metrics that map directly to what matters in production. Everything else is secondary signal.
The baseline choice is more nuanced than it looks. If you are fine-tuning to make a model follow a domain-specific format consistently, your baseline should be the strongest prompting approach you can construct -- multi-shot examples, a tight system prompt, maybe constrained decoding. If you skip this step and benchmark against a zero-shot GPT-3.5 call, you will almost always show a win, but you have not proven that fine-tuning beats a well-engineered prompt on GPT-4o, which is the honest comparison. Under-baselining is one of the most common ways teams convince themselves fine-tuning worked when it did not.
Metric selection by task type
Generic NLP benchmarks are largely useless for product decisions. ROUGE-L and BLEU were designed for translation and summarization research; they correlate poorly with whether a customer support bot gives a useful answer. Use them as sanity checks, not primary signals.
For classification and extraction tasks (intent detection, entity extraction, slot filling), precision, recall, and F1 at the label level are straightforward and reliable. For generative tasks -- summarization, Q&A, instruction following -- you need a combination of approaches. First, define a rubric: what does a correct answer actually look like? Then run automated scoring against that rubric using an LLM judge (you covered this in aidev-evaluation-llm-judge). Prompt GPT-4o to rate responses on correctness, format adherence, and conciseness on a 1-5 scale, and run it across your full test set. LLM judges are noisy but scalable. Complement them with human evaluation on a 50-100 example stratified sample -- humans catch hallucinations and tone issues that automated scorers miss consistently.
Business-aligned custom metrics are the most valuable signal of all. If you are fine-tuning a model to generate SQL queries, measure execution accuracy -- does the query run without errors and return the right rows? If you are fine-tuning for customer email triage, measure whether the predicted category matches the human-assigned category downstream. These task-specific metrics directly answer whether the product works, and they are usually straightforward to compute.
A real-world scenario: coding assistant specialization
Imagine you are fine-tuning Mistral-7B-Instruct using LoRA on 8,000 examples of internal Python code review comments. Your baseline is GPT-4o with a system prompt describing your team's style guide and five few-shot examples. You hold out 400 examples as a test set, stratified by comment category (naming, documentation, security, performance).
After training, you run three evaluations in parallel: (1) exact-match accuracy for the category label, (2) LLM-judge scoring on comment quality and specificity using GPT-4o as the judge, and (3) manual review of 50 randomly sampled examples by two senior engineers using a shared rubric. The fine-tuned 7B model scores 81% category accuracy versus 78% for the GPT-4o baseline. The LLM judge rates it 4.1 vs 4.3 out of 5 on quality. Human reviewers prefer the baseline in 28 of 50 comparisons.
Conclusion: the fine-tuned model is cheaper to run per token (self-hosted via vLLM) and slightly better at classification, but worse on free-form quality. The right call depends on your cost constraints. If you are running 10 million code reviews per month, the cost delta probably wins. At 50,000 reviews per month, the quality gap is probably not worth the hosting overhead.
Tradeoffs vs alternative approaches
The honest alternative to fine-tuning for most tasks is a better-engineered prompt on a stronger base model. GPT-4o with structured output and a detailed system prompt beats a LoRA-tuned 7B model on most generative tasks. The reason to fine-tune is usually one of: cost at scale, latency requirements (on-premises or edge deployment), data privacy (cannot send examples to an external API), or consistency of format that prompting genuinely cannot achieve reliably.
RAG is another alternative that often goes unevaluated against fine-tuning. If the task is knowledge-intensive (the model needs to know facts it was not trained on), RAG will almost always beat fine-tuning on freshness and accuracy. Fine-tuning teaches style and behavior, not factual recall. Conflating these is a common mistake.
What changes at scale
At 10 users, you can run manual eval on all outputs. At 10,000 users, you need automated metrics and LLM-judge pipelines running continuously, with statistical significance checks on metric deltas before you ship. At 10 million users, even tiny metric regressions represent large volumes of bad outputs, so you need shadow mode evaluation -- running the fine-tuned and baseline model in parallel on live traffic before a full cutover -- and real-time dashboards tracking production metrics like user satisfaction signals, downstream conversion, or explicit feedback. Evaluation infrastructure becomes a first-class engineering concern at this scale, not a script you run once before launch.
Key Takeaways
- Define your baseline and test set before fine-tuning starts, not after you see the results.
- Combine task-specific automated metrics with LLM-as-a-judge and human spot-checks.
- A marginal metric uplift rarely justifies the operational cost of hosting a custom model.
- Treat evaluation as a loop: results should drive the next round of data and training decisions.
Pro tips
- Run your baseline evaluation and fine-tuned evaluation in the same script, same API call parameters (temperature=0, same max_tokens), on the same day. Any environmental difference -- model version, token budget, sampling settings -- will pollute the delta and make results unreproducible.
- Always report confidence intervals or run significance tests (bootstrap resampling works well) before claiming a win. A 1-point ROUGE-L improvement on 50 examples is almost certainly noise. On 500+ examples, it starts to mean something.
- Log every prediction from both models to a persistent store (S3, BigQuery, a simple JSONL file) before you compute any aggregate metric. Aggregates are lossy -- you will want to slice by category, length, or failure mode after the fact, and you cannot go back if you only saved the average.
- Define your test set composition before you look at any training results. If you keep adjusting the test set after seeing numbers, you are leaking information. Write the test set spec in a doc and commit it to git before training starts.
Common pitfalls
- Mistake: Benchmarking the fine-tuned model against a weak baseline (zero-shot, no examples). Fix: Use the strongest realistic prompt you would actually deploy -- multi-shot, chain-of-thought, structured output -- as the baseline.
- Mistake: Using the validation loss from training as proof that the model improved. Fix: Validation loss measures token prediction; it does not measure task performance. Always evaluate on held-out examples using task metrics.
- Mistake: Reporting only aggregate metrics and ignoring failure mode analysis. Fix: Manually inspect 30-50 examples where the fine-tuned model scored worse than the baseline to understand systematic error patterns before shipping.
- Mistake: Including test examples that appeared in the training dataset in any form. Fix: Hash all training and test examples before training starts and assert zero overlap. Deduplication is covered in
aidev-fine-tuning-datasetsbut must be verified at eval time too.
Which primary metric to use when evaluating a fine-tuned model
| Option | Use when | Avoid when |
|---|---|---|
| Task-specific business metric (e.g., SQL execution accuracy, category F1) | You can define and compute a ground-truth correct answer programmatically. | The task is open-ended generation where there is no single correct answer. |
| LLM-as-a-judge scoring (GPT-4o rating on a rubric) | Outputs are generative and subjective; you have a large test set where human eval would be too slow. | The judge model is from the same family as the model being evaluated -- bias toward similar outputs inflates scores. |
| Human evaluation on a stratified sample | You need high-confidence ground truth for a final go/no-go decision before production deployment. | You are running evals continuously as part of a training loop; the cost and latency are prohibitive. |
| ROUGE-L / BLEU | You are doing a sanity check or comparing against published benchmarks for a standard NLP task. | You are making a product shipping decision -- these scores correlate poorly with user-perceived quality. |
Code Example
# Uses: openai>=1.30, rouge-score>=0.1.2
import json
from openai import OpenAI
from rouge_score import rouge_scorer
client = OpenAI() # reads OPENAI_API_KEY from env
test_examples = [
{"input": "Summarize: The experiment failed due to a memory leak.",
"reference": "The experiment failed because of a memory leak."},
{"input": "Summarize: Users reported slow response times on the dashboard.",
"reference": "Users experienced slow dashboard responses."},
]
scorer = rouge_scorer.RougeScorer(["rougeL"], use_stemmer=True)
def score_model(model_id: str, examples: list) -> float:
scores = []
for ex in examples:
resp = client.chat.completions.create(
model=model_id,
messages=[{"role": "user", "content": ex["input"]}],
max_tokens=64,
)
prediction = resp.choices[0].message.content.strip()
s = scorer.score(ex["reference"], prediction)
scores.append(s["rougeL"].fmeasure)
return sum(scores) / len(scores)
baseline_score = score_model("gpt-4o-mini", test_examples)
print(f"Baseline ROUGE-L: {baseline_score:.3f}")
# Replace with your fine-tuned model ID to compareHow this code works
This code's primary job is to establish a performance baseline for a language model, specifically evaluating its summarization capabilities. It does this by measuring how well a standard model, like gpt-4o-mini, performs on a set of test_examples before comparing fine-tuned versions against this benchmark.
The process begins by setting up the OpenAI client and defining test_examples, each containing an input prompt and its expected reference summary. A rouge_scorer.RougeScorer is then configured to calculate the "rougeL" F-measure, a metric that quantifies the overlap between two texts. A subtle but important detail here is use_stemmer=True, which normalizes words (e.g., "running" becomes "run") before comparison. This helps ensure that the score reflects semantic similarity rather than just exact word matches, making the evaluation more robust. The score_model function iterates through the test_examples, sends each input to the specified model_id via client.chat.completions.create, and receives a prediction. It then uses the configured scorer to compare the reference against the prediction, collecting the "rougeL" fmeasure for each. Finally, it returns the average F-measure. The script concludes by calling score_model with "gpt-4o-mini" to compute and print the baseline_score.
Production-grade example
Adds retries on rate limits, per-example error isolation, token logging, latency tracking, and structured JSON output for dashboards.
# Uses: openai>=1.30, rouge-score>=0.1.2, tenacity>=8.2
import os
import json
import logging
import time
from dataclasses import dataclass, asdict
from typing import Optional
from openai import OpenAI, RateLimitError, APITimeoutError, APIError
from rouge_score import rouge_scorer
from tenacity import retry, wait_exponential, stop_after_attempt, retry_if_exception_type
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], timeout=30.0)
scorer = rouge_scorer.RougeScorer(["rougeL"], use_stemmer=True)
@dataclass
class EvalResult:
model_id: str
example_id: int
rouge_l: float
prompt_tokens: int
completion_tokens: int
latency_ms: float
error: Optional[str] = None
@retry(
retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
wait=wait_exponential(multiplier=1, min=2, max=30),
stop=stop_after_attempt(4),
)
def call_model(model_id: str, prompt: str) -> tuple[str, int, int]:
start = time.monotonic()
resp = client.chat.completions.create(
model=model_id,
messages=[{"role": "user", "content": prompt}],
max_tokens=128,
temperature=0.0,
)
latency_ms = (time.monotonic() - start) * 1000
text = resp.choices[0].message.content.strip()
return text, resp.usage.prompt_tokens, resp.usage.completion_tokens, latency_ms
def evaluate_model(model_id: str, test_set: list[dict]) -> list[EvalResult]:
results = []
for i, ex in enumerate(test_set):
try:
prediction, pt, ct, latency = call_model(model_id, ex["input"])
rouge_l = scorer.score(ex["reference"], prediction)["rougeL"].fmeasure
result = EvalResult(model_id, i, rouge_l, pt, ct, latency)
log.info(json.dumps(asdict(result)))
except APIError as exc:
log.error("Example %d failed: %s", i, exc)
result = EvalResult(model_id, i, 0.0, 0, 0, 0.0, error=str(exc))
results.append(result)
return results
def summarize(results: list[EvalResult]) -> dict:
valid = [r for r in results if r.error is None]
if not valid:
return {"error": "all examples failed"}
return {
"model": valid[0].model_id,
"n": len(valid),
"mean_rouge_l": round(sum(r.rouge_l for r in valid) / len(valid), 4),
"total_prompt_tokens": sum(r.prompt_tokens for r in valid),
"total_completion_tokens": sum(r.completion_tokens for r in valid),
"mean_latency_ms": round(sum(r.latency_ms for r in valid) / len(valid), 1),
}
if __name__ == "__main__":
test_set = [
{"input": "Summarize: Memory leak caused test failure.",
"reference": "A memory leak caused the test to fail."},
]
for model_id in ["gpt-4o-mini", os.environ.get("FINETUNED_MODEL_ID", "gpt-4o-mini")]:
stats = summarize(evaluate_model(model_id, test_set))
print(json.dumps(stats, indent=2))How this code works
This code is designed to evaluate and compare the performance of different language models, specifically a baseline model against a fine-tuned one, within the context of a developer lesson on model evaluation. Its primary job is to measure how well models respond to prompts, track resource usage, and quantify performance using a common metric, providing key data to assess if fine-tuning improved results.
The process begins by defining an EvalResult dataclass to structure each example's evaluation data. The call_model function sends prompts to the OpenAI API, measuring latency and capturing token usage. A subtle but crucial detail here is the @retry decorator, which automatically re-attempts API calls if there are temporary issues like rate limits, preventing evaluation interruptions and ensuring robust data collection. The evaluate_model function then loops through a test_set, uses rouge_scorer to calculate a rougeL score by comparing the model's prediction to a reference answer, and logs the detailed EvalResult. Finally, the summarize function aggregates these results, calculating mean ROUGE-L scores, total tokens consumed, and average latency for a comprehensive performance overview, useful for comparing baseline and fine-tuned models.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Write a script that evaluates two models -- a baseline (gpt-4o-mini with a system prompt) and a stand-in fine-tuned model (use ft:gpt-4o-mini or another model ID from your env) -- on a small test set of at least 5 examples. Compute mean ROUGE-L and mean latency for each. Print a side-by-side comparison summary.
# Uses: openai>=1.30, rouge-score>=0.1.2
import os
from openai import OpenAI
from rouge_score import rouge_scorer
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
scorer = rouge_scorer.RougeScorer(["rougeL"], use_stemmer=True)
TEST_SET = [
# TODO: add at least 5 dicts with 'input' and 'reference' keys
]
BASELINE_MODEL = "gpt-4o-mini"
FINETUNED_MODEL = os.environ.get("FINETUNED_MODEL_ID", "gpt-4o-mini") # replace with real ID
def run_eval(model_id: str, system_prompt: str = "") -> dict:
# TODO: iterate over TEST_SET, call each model, score with ROUGE-L,
# track latency, and return a dict with mean_rouge_l and mean_latency_ms
pass
if __name__ == "__main__":
baseline_stats = run_eval(BASELINE_MODEL, system_prompt="You are a helpful assistant.")
finetuned_stats = run_eval(FINETUNED_MODEL)
# TODO: print a side-by-side comparison of both result dicts
passQuick check
Your fine-tuned model scores 2 points higher on ROUGE-L than the baseline. What should you do before declaring a win?
A fine-tuned model shows a large improvement on your eval set but underperforms on production traffic. What is the most likely cause?
Why is training validation loss a poor substitute for held-out task evaluation when comparing a fine-tuned model to a baseline?