Phase 2: Prompt Engineering & LLM Patterns

Faithfulness, relevance & helpfulness metrics

Intermediate ~14 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're trying to bake a magnificent chocolate cake, and you ask a super-smart chef for advice. This chef knows everything about cooking, but sometimes, when they give instructions, things can go a little bit wrong. Maybe their advice isn't quite right, or it doesn't really help you. To figure out what kind of wrong it is, we need to check three important things about their answers. This is how we make sure the computer programs that answer questions are giving us the best help possible.

First, we check for faithfulness. This is like asking: "Is what the chef saying actually true, or did they just make it up?" For example, if you ask how to make a cake, and the chef tells you to add a whole bottle of pickle juice, you’d probably think, "Wait, is that a real baking step, or are they just guessing?" If it’s not true to how cakes are really made, or to the recipe books they have, then their advice isn't faithful. It's like they're telling a fib about food!

Next, we look at relevance. This means: "Is the chef actually answering my question?" So, if you asked for cake instructions, but the chef started explaining how to perfectly scramble an egg, their advice might be totally true (scrambling eggs is a real thing!), but it’s not relevant to your cake problem. It doesn't help you with what you specifically asked. The chef needs to stick to the topic!

Finally, we consider helpfulness. This asks: "Is the chef’s advice clear, complete, and easy for me to actually use to solve my problem?" Imagine the chef says, "To make a cake, you just bake it." Well, that's true and about cake, but it's not helpful at all! You still don’t know what ingredients to use, what temperature to set the oven, or for how long. The best advice is faithful, relevant, and helpful enough to get you to that delicious chocolate cake.

By checking these three things separately, people who build these smart computer programs can pinpoint exactly what went wrong. Instead of just saying, "the computer gave a bad answer," they can say, "the computer made up a fact!" or "it talked about the wrong thing!" or "its answer was true and relevant, but totally unclear!" This helps them understand and fix the problem, making the computer smarter and more reliable, so you always get perfect cake-baking advice.

Faithfulness, relevance, and helpfulness sound like vague quality adjectives until you pin down what each one actually measures and how to compute it programmatically.

Faithfulness is about grounding. The model's response should only assert claims that are entailed by the provided context or by verifiable world knowledge. In a RAG system, you retrieve a set of documents and pass them as context. A faithful response draws conclusions only from those documents. The canonical way to score this is to decompose the response into atomic claims, then for each claim ask: does the context support this? RAGAS, the most widely adopted open-source eval framework for RAG, does exactly this using an LLM to extract claims and a second LLM call to verify each one against the retrieved chunks. A score of 0.0 means every claim is hallucinated; 1.0 means every claim is fully supported. Note that faithfulness does NOT measure whether the context itself is correct -- that's a retrieval quality problem, not a generation quality problem. Keep the two separate.

Relevance measures how tightly the response addresses the actual question. A common implementation generates several hypothetical questions that the response could plausibly answer, then computes the cosine similarity between those hypothetical questions and the original question in embedding space. If the response answers a slightly different question than the one asked, that similarity drops. This is answer relevancy in RAGAS. In practice, a response can score 1.0 on faithfulness (everything said is true) and 0.3 on relevance (it answered something adjacent to the question). This frequently happens when retrieved context is good but the generation prompt is too permissive -- the model uses the context to answer a related but different question.

Helpfulness is trickier because it's inherently holistic. You can't decompose helpfulness into atomic operations the way you can faithfulness. Most practitioners score helpfulness with a 1-5 Likert rubric evaluated by either a human annotator or an LLM judge. The rubric typically asks: Is the response clear? Is it complete enough to act on? Does it avoid unnecessary verbosity? Does it address the user's underlying intent, not just the literal words of the query? A financially motivated user asking "should I buy TSLA?" needs a different kind of helpful response than a developer asking "how do I parse JSON in Python." Baking the user's intent category into your evaluation rubric is what separates a useful helpfulness score from a meaningless one.

A real-world scenario: you're building a support chatbot for a SaaS product. Your RAG pipeline retrieves documentation chunks and generates answers. After launch, users complain the bot is "not helpful." You break that down into the three metrics. Faithfulness is 0.95 -- almost no hallucinations. Relevance is 0.71 -- the bot often answers a nearby question instead of the exact one asked. Helpfulness is 2.8/5 -- responses are accurate and somewhat on-topic, but they paste raw documentation without synthesizing it into an actionable step-by-step answer. Now you have a concrete diagnosis: improve retrieval precision (so the right chunks surface) and update the system prompt to require numbered steps. You re-evaluate after the change. Faithfulness stays at 0.94, relevance climbs to 0.88, helpfulness reaches 3.9/5. That's a story you can tell in a sprint review.

The tradeoff landscape matters at scale. RAGAS uses multiple LLM calls per data point -- roughly 3-5 calls to score faithfulness and relevance on a single QA pair. At 10 users you run evals manually. At 10k users you're evaluating hundreds of production samples per day, and those LLM calls start adding up. Most teams adopt a tiered strategy: run cheap deterministic checks (exact match, token overlap, length guards) on every request, run RAGAS-style metrics on a 2-5% sample, and run expensive human review on edge cases flagged by automated scoring. At 10M users you're also dealing with distribution shift -- the user population changes over time, so you need a sliding-window eval strategy, not a one-time benchmark. Instrument your pipeline so every request logs the question, retrieved context, and response to a store like S3 or BigQuery, then run nightly eval jobs over a stratified sample.

Cost and latency implications are non-trivial. If you use GPT-4o as your judge, scoring 1000 samples per day at ~1000 tokens per eval call costs roughly $1-2/day -- illustrative, verify current pricing. If you use a hosted RAGAS setup with a cheaper model like GPT-4o-mini, you can cut that by 10x with some precision loss. Many teams run a calibration experiment: score 200 samples with a strong judge model and 200 with a cheap model, compute the rank correlation, and accept the cheap model if Spearman's rho is above 0.85. This lets you scale eval economically without abandoning quality signal.

Key Takeaways

  • Faithfulness measures grounding: the response must only assert what the context or facts support.
  • Relevance is independent of faithfulness; a response can be accurate but completely miss the question.
  • Helpfulness is a composite metric; optimize faithfulness and relevance first, then holistic utility.
  • Automate metric scoring early using rubrics and LLM-as-a-judge so evaluation scales with your dataset.

Pro tips

  • Faithfulness and context recall measure different failure modes. Faithfulness catches generation hallucinations; context recall (did retrieval surface the right chunks?) catches retrieval failures. Running only faithfulness evals will miss a huge class of RAG bugs that live upstream in the retriever.
  • Decompose helpfulness into sub-dimensions before you score it: clarity, completeness, actionability, tone match. A single 1-5 helpfulness score is almost impossible to act on because you can't tell which sub-dimension to fix. Score each separately, even if you aggregate later.
  • When using LLM-as-a-judge for any of these metrics, always log the judge's reasoning, not just its numeric score. The reasoning trace is where you find systematic prompt failures and edge cases that a single number will never surface.
  • Establish a human-calibrated baseline before you automate anything. Score 50-100 samples by hand, then measure how well your automated metric correlates with your human scores. A metric that doesn't correlate with humans at r > 0.7 is worse than no metric because it creates false confidence.

Common pitfalls

  • Mistake: Treating helpfulness as a proxy for faithfulness and skipping grounding checks. Fix: Score faithfulness independently; a fluent, seemingly helpful response can still hallucinate facts the context never stated.
  • Mistake: Evaluating only on your golden test set and ignoring production traffic samples. Fix: Sample 2-5% of real requests daily and run your eval pipeline on them to catch distribution drift.
  • Mistake: Using the same LLM for both generation and judging without a stronger reference model. Fix: Use a model one tier stronger (e.g., GPT-4o as judge for GPT-4o-mini outputs) to avoid the model validating its own failure modes.
  • Mistake: Aggregating scores into one "quality" number before understanding what drives it. Fix: Keep faithfulness, relevance, and helpfulness as separate columns in your eval results table; aggregation hides actionable signal.

Which metric to prioritize for your use case

Option Use when Avoid when
Faithfulness Your app makes factual claims from retrieved documents (RAG, Q&A over docs, legal/medical summaries). Your app is creative writing or brainstorming where grounding to a specific context is not required.
Answer Relevancy Users ask specific questions and expect focused answers; tangential responses erode trust. Your app intentionally provides broad exploratory responses, such as a research assistant surfacing related topics.
Helpfulness (holistic rubric) You need a single stakeholder-facing quality signal and your audience cares about actionability and clarity, not just accuracy. You need a debuggable signal for engineering; helpfulness is too coarse to tell you what to fix in the pipeline.
All three independently You're running a systematic eval pipeline and want to triage which part of the system (retrieval, generation, prompt) is responsible for failures. You need a quick sanity check with minimal LLM calls; scoring all three multiplies cost 3-5x per sample.

Code Example

python
# ragas>=0.1.0 — pip install ragas
from ragas.metrics import faithfulness, answer_relevancy
from ragas import evaluate
from datasets import Dataset

# Minimal example: one question, one RAG answer, one retrieved context
data = {
    "question": ["What is the capital of France?"],
    "answer": ["The capital of France is Paris."],
    "contexts": [["France is a country in Western Europe. Its capital city is Paris."]],
    "ground_truth": ["Paris"]
}

dataset = Dataset.from_dict(data)
result = evaluate(dataset, metrics=[faithfulness, answer_relevancy])
print(result)  # {'faithfulness': 1.0, 'answer_relevancy': 0.97}

How this code works

This code illustrates how to evaluate the quality of a Retrieval-Augmented Generation (RAG) system's output using the ragas library. Its primary job is to measure two crucial metrics: faithfulness, which assesses if the generated answer is factually supported by the information retrieved, and answer_relevancy, which checks how well the answer addresses the original question. By running this evaluation, the system provides objective, numerical scores for these aspects, helping developers understand and improve their RAG pipeline's performance.

The process starts by defining a data dictionary containing a single example: a question, the RAG system's answer, the contexts (the information retrieved from a knowledge base), and a ground_truth reference. This structured data is then converted into a Dataset object using Dataset.from_dict, which is the required input format for ragas. The core of the evaluation happens with the evaluate function, where the dataset is passed along with a list of desired metrics. A subtle but important point for beginners is that faithfulness primarily checks the answer against the provided contexts to ensure the answer doesn't hallucinate beyond what was retrieved, even though ground_truth is also present for a complete comparison set.

Production-grade example

Adds retries with backoff, timeouts, structured logging, alerting threshold, and graceful degradation on failure.

python
# ragas>=0.1.0, openai>=1.0.0, tenacity>=8.0.0
import os
import time
import logging
import json
from typing import Any
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy
from datasets import Dataset
from openai import RateLimitError, APITimeoutError

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)

OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]  # never hardcode

@retry(
    retry=retry_if_exception_type((RateLimitError, APITimeoutError)),
    wait=wait_exponential(multiplier=1, min=2, max=30),
    stop=stop_after_attempt(5)
)
def run_eval_with_retry(dataset: Dataset) -> dict[str, Any]:
    start = time.monotonic()
    result = evaluate(dataset, metrics=[faithfulness, answer_relevancy], timeout=30)
    elapsed = time.monotonic() - start
    scores = result.to_pandas()[["faithfulness", "answer_relevancy"]].mean().to_dict()
    log.info(json.dumps({
        "event": "eval_complete",
        "n_samples": len(dataset),
        "elapsed_s": round(elapsed, 2),
        **{k: round(v, 4) for k, v in scores.items()}
    }))
    if scores["faithfulness"] < 0.7:
        log.warning(json.dumps({"event": "faithfulness_alert", "score": scores["faithfulness"]}))
    return scores

def evaluate_sample(questions, answers, contexts, ground_truths):
    dataset = Dataset.from_dict({
        "question": questions,
        "answer": answers,
        "contexts": contexts,
        "ground_truth": ground_truths
    })
    try:
        return run_eval_with_retry(dataset)
    except Exception as exc:
        log.error(json.dumps({"event": "eval_failed", "error": str(exc)}))
        return {"faithfulness": None, "answer_relevancy": None}  # graceful degradation

How this code works

This code provides a robust way to evaluate the quality of an LLM's responses, specifically for Retrieval-Augmented Generation (RAG) systems. Its main job is to measure "faithfulness," checking if the answer aligns with the provided context, and "answer relevancy," ensuring the answer directly addresses the question. This helps assess how well an LLM utilizes its source material and responds appropriately to user queries.

The evaluate_sample function prepares input data, including questions, answers, contexts, and ground truths, structuring them into a Dataset object as required by the ragas library. It then calls run_eval_with_retry, which performs the actual evaluation using ragas.evaluate with the specified faithfulness and answer_relevancy metrics. A subtle but critical feature is the @retry decorator from tenacity on run_eval_with_retry. This automatically re-attempts the evaluation if external API issues like RateLimitError or APITimeoutError occur, making the evaluation process more resilient without crashing. The code logs the results, warns if faithfulness scores drop below 0.7, and gracefully handles complete failures by returning None for the scores.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a small evaluation script that scores three QA pairs from a fictional product FAQ using faithfulness and answer relevancy via RAGAS. At least one answer should contain a hallucinated claim not present in the context. Print the per-row scores and identify which row has the lowest faithfulness score.

python
# pip install ragas datasets
from ragas.metrics import faithfulness, answer_relevancy
from ragas import evaluate
from datasets import Dataset

# TODO: Define at least 3 QA pairs.
# One answer should hallucinate a fact not in its context.
questions = []
answers = []
contexts = []  # list of lists — each item is a list of retrieved chunks
ground_truths = []

# TODO: Build the Dataset from the lists above.
dataset = None

# TODO: Run evaluate() with faithfulness and answer_relevancy metrics.
result = None

# TODO: Convert result to a DataFrame and print per-row scores.
# TODO: Print the index and faithfulness score of the lowest-scoring row.

Quick check

  1. A RAG response cites a statistic that isn't in any retrieved chunk but happens to be factually correct. How does this affect its faithfulness score?

  2. You see faithfulness=0.95 but answer_relevancy=0.55 in your eval results. What is the most likely root cause?

  3. Why is a single aggregate 'quality' score less useful than separate faithfulness, relevance, and helpfulness scores during development?

Self-check: Describe a scenario where a RAG response scores high on faithfulness and low on answer relevancy. Then explain what change you would make to the pipeline first, and which metric you would check to confirm the fix worked.