Phase 3: RAG & Knowledge Systems

Re-ranking models for improved precision

Advanced ~15 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're trying to bake a super yummy chocolate cake, and you have a new, complicated recipe. The first thing you need to do is gather all your ingredients. So, you rush to a giant supermarket, looking for anything that could possibly be in a cake. You quickly grab lots of things – flour, sugar, eggs, milk, chocolate, sprinkles, baking powder, even some yeast or salt. You end up with a huge shopping cart full, maybe 20 to 100 different items, because you want to make sure you didn't miss anything important for your cake. This first grab is super fast, just like a computer's first search to find anything that might be related to your question.

Now you get home, and you have this enormous pile of stuff on your kitchen counter. You can't just throw everything into the bowl, right? Some things might be similar but not quite right for this specific chocolate cake. For example, you grabbed yeast, but this cake needs baking powder, not yeast. Or you grabbed sweetener, but the recipe calls for real sugar. So, you slow down. You carefully read your recipe and compare it to each ingredient, one by one. You put the recipe right next to the flour bag to make sure it's the right kind of flour, or you compare the sugar against the recipe to see if it's the exact type needed. This careful, slower check helps you pick out only the very best, perfect ingredients, maybe just 3 to 10 items, that truly belong in your cake.

This careful sorting process is exactly what "re-ranking" does for a computer that's trying to answer a question, especially in something called a Retrieval Augmented Generation (RAG) system. When you ask an AI (Artificial Intelligence) a question, it first quickly scans through tons of information, like your big supermarket trip. It grabs all sorts of documents or pieces of text that might be useful. But just like your shopping cart, not everything will be exactly what's needed. Re-ranking is like your careful check at home. It takes those many initial results and uses a smarter, slower "judge" to look at each one very carefully alongside your original question. It asks: "Does this specific piece of information truly and perfectly answer this exact question?"

By doing this, the computer can filter out the almost-right answers or confusing information. It ends up with a much smaller, super-precise list of facts. This means that when you build an AI system using this technique, it will give much clearer, more accurate answers to tricky questions, rather than getting confused by similar but incorrect information. It makes the AI smarter at finding the needle in the haystack, so you get reliable and helpful responses every time.

The mental model for re-ranking is a two-stage funnel. Stage one (retrieval) runs fast approximate nearest-neighbor search over millions of vectors in under 100ms. It casts a wide net and prioritizes recall. Stage two (re-ranking) takes the 20–100 candidates from stage one and runs a much more expensive model over each one. The re-ranker doesn't need to handle millions of documents, just dozens, so the compute is tractable. The output is a sorted list where the top items have been verified by a model that genuinely understands the relationship between query and document, not just their embedding proximity.

The bi-encoder vs cross-encoder distinction is the key technical insight. A bi-encoder like text-embedding-3-small encodes the query into a 1536-dimensional vector, encodes each document into its own vector, then computes cosine similarity. There is no direct attention between query tokens and document tokens. A cross-encoder like cross-encoder/ms-marco-MiniLM-L-6-v2 or Cohere's Rerank API takes the input [CLS] query [SEP] document [SEP] and runs a full transformer over the concatenated sequence. Every query token can attend to every document token. This is why cross-encoders catch relevance signals that bi-encoders miss: negation, specificity, paraphrase with wrong context, numeric precision. The tradeoff is that cross-encoders can't be precomputed offline, so you pay inference cost at query time for every candidate.

In a real production scenario, consider a customer support knowledge base with 50,000 articles. A user asks: "How do I cancel my subscription without losing my billing history?" Vector search might retrieve 50 articles about cancellation, subscription management, and billing. Some of those articles will be about cancellation but not mention billing history retention. Others will be about exporting billing history but not about cancellation. The cross-encoder re-ranker, seeing the full query next to each article, can identify which articles actually address both concerns together and push them to the top. The LLM then receives 5 highly relevant articles instead of 50 mixed-quality ones, producing a more accurate and concise answer.

Tradeoffs against alternatives: You could skip re-ranking and just pass more documents to the LLM (long-context stuffing). Models like Gemini 1.5 Pro accept 1M token contexts, so why not send 100 documents? A few reasons: cost scales linearly with tokens; latency increases significantly; models suffer from "lost in the middle" attention degradation on long contexts, meaning they attend strongly to the beginning and end but less to the middle; and you're paying for a lot of tokens that carry no useful information. Re-ranking lets you stay with a 4k–8k context window but fill it with the best possible content. Another alternative is maximal marginal relevance (MMR), which re-orders by diversity rather than relevance. MMR is useful when your query might have multiple facets but it doesn't improve relevance precision the way a cross-encoder does.

At scale, the economics change significantly. At 10 users, you can run a local cross-encoder on CPU and re-rank 50 documents in roughly 200–400ms. At 10,000 users, that latency becomes a bottleneck and you need either a GPU-backed inference server for your local model or a managed API like Cohere Rerank. Cohere charges per API call (check current pricing; costs vary), and at high volume you'll want to batch requests and cache re-rank scores for repeated queries. At 10 million users, you're likely building custom re-ranking infrastructure, possibly fine-tuning a smaller cross-encoder on your domain's query-document pairs using training data you've collected from user feedback. You'd also add observability: log the pre-rerank and post-rerank document sets, track whether the re-ranker actually changes the order, and monitor for distribution shift where your re-ranker starts underperforming on new query types.

Latency budget is the practical constraint most engineers underestimate. A typical RAG pipeline target is 1–2 seconds end-to-end. Vector search takes 20–80ms. LLM generation takes 500ms–2s depending on model and output length. That leaves 100–500ms for re-ranking. Running a local ms-marco-MiniLM-L-6-v2 model on CPU re-ranking 50 candidates takes roughly 300–600ms depending on document length. On a single GPU (T4 class), that drops to under 50ms. Cohere Rerank API adds roughly 100–300ms of network and compute latency depending on batch size and geography. If you're over budget, reduce your candidate pool size before re-ranking rather than skipping re-ranking entirely.

Key Takeaways

  • Use re-ranking to improve precision after retrieval, not as a substitute for good retrieval.
  • Cross-encoders score query+document jointly; bi-encoders score them independently — that difference matters.
  • Always over-retrieve (top-50 to top-100) before re-ranking; you can't recover documents not in the candidate set.
  • Cohere Rerank and cross-encoder models from sentence-transformers are the two dominant production options.

Pro tips

  • Set your initial retrieval to return 3–5x more candidates than you'll pass to the LLM. If you want 5 documents in context, retrieve 25–50 first. The re-ranker can only promote documents that retrieval found; if retrieval misses the best document, re-ranking cannot recover it.
  • Fine-tuning a cross-encoder on domain-specific query-document pairs from your application's query logs dramatically outperforms out-of-the-box models. Even 500–1000 labeled pairs can move precision metrics significantly on specialized domains like legal, medical, or technical documentation.
  • Log the pre-rerank and post-rerank ranked lists for a sample of production queries. Compute the percentage of queries where the top-1 document changed. If re-ranking rarely changes the order, your retriever is already doing well, and you may not need re-ranking at all. If it changes order 70%+ of the time, your retriever needs tuning.
  • Cohere Rerank and similar APIs have input length limits (typically 512 tokens per document). If your chunks are longer, the re-ranker silently truncates them. Either chunk to fit within the limit or use the first 512 tokens of each candidate as the re-rank input and pass the full chunk to the LLM if selected.

Common pitfalls

  • Mistake: Passing all documents directly to the re-ranker without an initial retrieval step. Fix: Always use vector or hybrid search first to reduce candidates to a manageable pool (20–100). Re-ranking every document in a large corpus is computationally infeasible.
  • Mistake: Using the re-ranker's output score as a relevance threshold without calibration. Fix: Scores from different models are not comparable. Calibrate thresholds empirically on your dataset, or use relative ranking rather than absolute score cutoffs.
  • Mistake: Running a local cross-encoder synchronously in a web request handler, blocking the thread. Fix: Use async workers or a dedicated inference service. Even a fast local model adds 200–500ms per request on CPU, which compounds under concurrent load.
  • Mistake: Skipping re-ranking when using hybrid search because "hybrid already improves precision." Fix: Hybrid search improves recall by combining signals; re-ranking improves precision independently. The two techniques are complementary, not substitutes.

When to use Cohere Rerank API vs a local cross-encoder

Option Use when Avoid when
Cohere Rerank API You need fast integration, don't want to manage GPU infra, and query volume is low-to-medium (under ~1M queries/day). Data cannot leave your infrastructure (compliance), or per-query API cost at your volume exceeds self-hosted model cost.
Local cross-encoder (sentence-transformers) You need data sovereignty, have GPU infrastructure available, or want to fine-tune on domain-specific data. You have no GPU and need sub-200ms re-ranking latency; CPU inference on 50 candidates is too slow for your SLA.
No re-ranking (pass top-K directly) Retrieval quality is already high, latency budget is extremely tight, or your LLM has a very large context window and cost isn't a concern. Users are seeing irrelevant answers; precision of retrieval is noticeably poor; production quality matters.
MMR (Maximal Marginal Relevance) Queries have multiple sub-topics and you want diverse coverage across retrieved documents rather than maximum relevance to one interpretation. Your main problem is relevance precision rather than diversity; MMR does not improve relevance, only variety.

Code Example

python
# sentence-transformers==2.7.0
from sentence_transformers import CrossEncoder

model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")

query = "What is the capital of France?"
candidates = [
    "Paris is the capital and largest city of France.",
    "France is a country in Western Europe known for its wine.",
    "The Eiffel Tower is a famous landmark located in Paris.",
    "Berlin is the capital of Germany.",
]

# Score each (query, document) pair jointly
pairs = [[query, doc] for doc in candidates]
scores = model.predict(pairs)

# Sort candidates by re-rank score descending
ranked = sorted(zip(scores, candidates), reverse=True)
for score, doc in ranked:
    print(f"{score:.4f}  {doc[:60]}")

How this code works

This code demonstrates how to use a re-ranking model to improve the precision of document retrieval, a core component of advanced RAG systems. Its job is to take a user's query and a list of candidates (potential answers or documents), then intelligently sort these candidates to show the most relevant ones first.

The process begins by loading a specialized neural network model using CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2"). This particular model is pre-trained to understand the relationship between a query and a document, evaluating how well they match. To do this, the code prepares pairs of [query, document] for each candidate. The model then calculates a scores array by calling model.predict(pairs), assigning each pair a numerical relevance score. A subtle but crucial aspect is that CrossEncoder models evaluate the query and document together for a more precise relevance judgment, unlike simpler models that might just compare their independent embeddings. Finally, zip(scores, candidates) links each score back to its original document, and sorted(..., reverse=True) arranges them so the documents with the highest (most relevant) scores appear at the top.

Production-grade example

Adds retries on rate limits, graceful degradation on bad requests, and structured latency/cost logging.

python
# cohere==5.x, tenacity==8.x
import os
import time
import logging
from typing import List
import cohere
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

co = cohere.Client(api_key=os.environ["COHERE_API_KEY"])

@retry(
    stop=stop_after_attempt(3),
    wait=wait_exponential(multiplier=1, min=1, max=10),
    retry=retry_if_exception_type(cohere.errors.TooManyRequestsError),
)
def rerank_documents(
    query: str,
    documents: List[str],
    top_n: int = 5,
    model: str = "rerank-english-v3.0",
) -> List[dict]:
    if not documents:
        logger.warning("rerank called with empty document list, returning empty")
        return []

    start = time.perf_counter()
    try:
        response = co.rerank(
            query=query,
            documents=documents,
            top_n=top_n,
            model=model,
        )
    except cohere.errors.BadRequestError as exc:
        logger.error("Cohere rerank bad request", extra={"query_len": len(query), "doc_count": len(documents), "error": str(exc)})
        # Graceful degradation: return original order truncated to top_n
        return [{"index": i, "document": documents[i], "relevance_score": None} for i in range(min(top_n, len(documents)))]

    latency_ms = (time.perf_counter() - start) * 1000
    logger.info(
        "rerank completed",
        extra={
            "model": model,
            "input_docs": len(documents),
            "top_n": top_n,
            "latency_ms": round(latency_ms, 1),
            "billed_units": response.meta.billed_units.search_units if response.meta else None,
        },
    )

    return [
        {"index": r.index, "document": documents[r.index], "relevance_score": r.relevance_score}
        for r in response.results
    ]


if __name__ == "__main__":
    query = "How do I reset my password?"
    docs = [
        "To reset your password, click Forgot Password on the login page.",
        "Our refund policy allows returns within 30 days.",
        "Password requirements: 8 characters, one uppercase, one number.",
        "Contact support at [email protected] for account issues.",
    ]
    results = rerank_documents(query, docs, top_n=2)
    for r in results:
        print(f"Score: {r['relevance_score']:.4f}  Doc: {r['document'][:60]}")

How this code works

This code significantly improves RAG systems by re-ranking documents to present the most relevant information first. The rerank_documents function takes a user query and a list of documents, then leverages Cohere's rerank-english-v3.0 model to sort them by relevance. It connects to Cohere's API using an api_key loaded securely from os.environ. The function includes a top_n parameter to specify how many of the highest-scoring documents to return, crucial for efficiency. It also handles edge cases, such as an empty document list, by returning an empty result immediately.

A subtle but important feature for production-grade robustness is the @retry decorator from tenacity. This decorator automatically reattempts the Cohere API call up to three times (stop_after_attempt) if it encounters a cohere.errors.TooManyRequestsError, waiting progressively longer (wait_exponential) to avoid overwhelming the server. If a cohere.errors.BadRequestError occurs (e.g., due to invalid input), the code gracefully degrades by returning the original documents, truncated to top_n, without a relevance score. logging is used throughout to record important details like latency_ms and billed_units for monitoring and debugging. The if __name__ == "__main__": block demonstrates calling the function and printing the re-ranked documents with their relevance_score.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a minimal re-ranking comparison script. Given a fixed query and 8 candidate documents, retrieve the top-3 using cosine similarity on embeddings, then re-rank all 8 using a local cross-encoder and take the top-3. Print both result sets side by side and note which documents moved up or down after re-ranking.

python
# sentence-transformers==2.7.0
from sentence_transformers import SentenceTransformer, CrossEncoder, util

query = "How do I cancel my subscription?"
candidates = [
    "To cancel, go to Account Settings and click Cancel Subscription.",
    "Subscription plans start at $9.99 per month.",
    "You can pause your subscription for up to 3 months.",
    "Our cancellation policy requires 30 days notice.",
    "To delete your account, contact [email protected].",
    "Upgrade or downgrade your plan at any time.",
    "Cancellation takes effect at the end of the billing period.",
    "Refunds are not issued for partial billing periods after cancellation.",
]

# TODO: Load a bi-encoder model (e.g., 'all-MiniLM-L6-v2')
# TODO: Encode the query and all candidates
# TODO: Compute cosine similarity scores and get top-3 by embedding similarity

# TODO: Load a cross-encoder model (e.g., 'cross-encoder/ms-marco-MiniLM-L-6-v2')
# TODO: Score all (query, candidate) pairs and get top-3 by re-rank score

# TODO: Print both top-3 lists and highlight any differences

Quick check

  1. Why can a cross-encoder detect relevance signals that a bi-encoder misses?

  2. You're building a RAG system. Initial retrieval returns 5 documents. Adding re-ranking barely changes the results. What does this most likely indicate?

  3. A colleague proposes skipping re-ranking and instead passing all 100 retrieved documents to a long-context LLM. What is the strongest argument against this approach?

Self-check: Explain in your own words why bi-encoders are used for initial retrieval but cross-encoders are used for re-ranking. Then describe one scenario where skipping re-ranking would be a defensible engineering decision.