The mental model for re-ranking is a two-stage funnel. Stage one (retrieval) runs fast approximate nearest-neighbor search over millions of vectors in under 100ms. It casts a wide net and prioritizes recall. Stage two (re-ranking) takes the 20–100 candidates from stage one and runs a much more expensive model over each one. The re-ranker doesn't need to handle millions of documents, just dozens, so the compute is tractable. The output is a sorted list where the top items have been verified by a model that genuinely understands the relationship between query and document, not just their embedding proximity.
The bi-encoder vs cross-encoder distinction is the key technical insight. A bi-encoder like text-embedding-3-small encodes the query into a 1536-dimensional vector, encodes each document into its own vector, then computes cosine similarity. There is no direct attention between query tokens and document tokens. A cross-encoder like cross-encoder/ms-marco-MiniLM-L-6-v2 or Cohere's Rerank API takes the input [CLS] query [SEP] document [SEP] and runs a full transformer over the concatenated sequence. Every query token can attend to every document token. This is why cross-encoders catch relevance signals that bi-encoders miss: negation, specificity, paraphrase with wrong context, numeric precision. The tradeoff is that cross-encoders can't be precomputed offline, so you pay inference cost at query time for every candidate.
In a real production scenario, consider a customer support knowledge base with 50,000 articles. A user asks: "How do I cancel my subscription without losing my billing history?" Vector search might retrieve 50 articles about cancellation, subscription management, and billing. Some of those articles will be about cancellation but not mention billing history retention. Others will be about exporting billing history but not about cancellation. The cross-encoder re-ranker, seeing the full query next to each article, can identify which articles actually address both concerns together and push them to the top. The LLM then receives 5 highly relevant articles instead of 50 mixed-quality ones, producing a more accurate and concise answer.
Tradeoffs against alternatives: You could skip re-ranking and just pass more documents to the LLM (long-context stuffing). Models like Gemini 1.5 Pro accept 1M token contexts, so why not send 100 documents? A few reasons: cost scales linearly with tokens; latency increases significantly; models suffer from "lost in the middle" attention degradation on long contexts, meaning they attend strongly to the beginning and end but less to the middle; and you're paying for a lot of tokens that carry no useful information. Re-ranking lets you stay with a 4k–8k context window but fill it with the best possible content. Another alternative is maximal marginal relevance (MMR), which re-orders by diversity rather than relevance. MMR is useful when your query might have multiple facets but it doesn't improve relevance precision the way a cross-encoder does.
At scale, the economics change significantly. At 10 users, you can run a local cross-encoder on CPU and re-rank 50 documents in roughly 200–400ms. At 10,000 users, that latency becomes a bottleneck and you need either a GPU-backed inference server for your local model or a managed API like Cohere Rerank. Cohere charges per API call (check current pricing; costs vary), and at high volume you'll want to batch requests and cache re-rank scores for repeated queries. At 10 million users, you're likely building custom re-ranking infrastructure, possibly fine-tuning a smaller cross-encoder on your domain's query-document pairs using training data you've collected from user feedback. You'd also add observability: log the pre-rerank and post-rerank document sets, track whether the re-ranker actually changes the order, and monitor for distribution shift where your re-ranker starts underperforming on new query types.
Latency budget is the practical constraint most engineers underestimate. A typical RAG pipeline target is 1–2 seconds end-to-end. Vector search takes 20–80ms. LLM generation takes 500ms–2s depending on model and output length. That leaves 100–500ms for re-ranking. Running a local ms-marco-MiniLM-L-6-v2 model on CPU re-ranking 50 candidates takes roughly 300–600ms depending on document length. On a single GPU (T4 class), that drops to under 50ms. Cohere Rerank API adds roughly 100–300ms of network and compute latency depending on batch size and geography. If you're over budget, reduce your candidate pool size before re-ranking rather than skipping re-ranking entirely.
Key Takeaways
- Use re-ranking to improve precision after retrieval, not as a substitute for good retrieval.
- Cross-encoders score query+document jointly; bi-encoders score them independently — that difference matters.
- Always over-retrieve (top-50 to top-100) before re-ranking; you can't recover documents not in the candidate set.
- Cohere Rerank and cross-encoder models from sentence-transformers are the two dominant production options.
Pro tips
- Set your initial retrieval to return 3–5x more candidates than you'll pass to the LLM. If you want 5 documents in context, retrieve 25–50 first. The re-ranker can only promote documents that retrieval found; if retrieval misses the best document, re-ranking cannot recover it.
- Fine-tuning a cross-encoder on domain-specific query-document pairs from your application's query logs dramatically outperforms out-of-the-box models. Even 500–1000 labeled pairs can move precision metrics significantly on specialized domains like legal, medical, or technical documentation.
- Log the pre-rerank and post-rerank ranked lists for a sample of production queries. Compute the percentage of queries where the top-1 document changed. If re-ranking rarely changes the order, your retriever is already doing well, and you may not need re-ranking at all. If it changes order 70%+ of the time, your retriever needs tuning.
- Cohere Rerank and similar APIs have input length limits (typically 512 tokens per document). If your chunks are longer, the re-ranker silently truncates them. Either chunk to fit within the limit or use the first 512 tokens of each candidate as the re-rank input and pass the full chunk to the LLM if selected.
Common pitfalls
- Mistake: Passing all documents directly to the re-ranker without an initial retrieval step. Fix: Always use vector or hybrid search first to reduce candidates to a manageable pool (20–100). Re-ranking every document in a large corpus is computationally infeasible.
- Mistake: Using the re-ranker's output score as a relevance threshold without calibration. Fix: Scores from different models are not comparable. Calibrate thresholds empirically on your dataset, or use relative ranking rather than absolute score cutoffs.
- Mistake: Running a local cross-encoder synchronously in a web request handler, blocking the thread. Fix: Use async workers or a dedicated inference service. Even a fast local model adds 200–500ms per request on CPU, which compounds under concurrent load.
- Mistake: Skipping re-ranking when using hybrid search because "hybrid already improves precision." Fix: Hybrid search improves recall by combining signals; re-ranking improves precision independently. The two techniques are complementary, not substitutes.
When to use Cohere Rerank API vs a local cross-encoder
| Option | Use when | Avoid when |
|---|---|---|
| Cohere Rerank API | You need fast integration, don't want to manage GPU infra, and query volume is low-to-medium (under ~1M queries/day). | Data cannot leave your infrastructure (compliance), or per-query API cost at your volume exceeds self-hosted model cost. |
| Local cross-encoder (sentence-transformers) | You need data sovereignty, have GPU infrastructure available, or want to fine-tune on domain-specific data. | You have no GPU and need sub-200ms re-ranking latency; CPU inference on 50 candidates is too slow for your SLA. |
| No re-ranking (pass top-K directly) | Retrieval quality is already high, latency budget is extremely tight, or your LLM has a very large context window and cost isn't a concern. | Users are seeing irrelevant answers; precision of retrieval is noticeably poor; production quality matters. |
| MMR (Maximal Marginal Relevance) | Queries have multiple sub-topics and you want diverse coverage across retrieved documents rather than maximum relevance to one interpretation. | Your main problem is relevance precision rather than diversity; MMR does not improve relevance, only variety. |
Code Example
# sentence-transformers==2.7.0
from sentence_transformers import CrossEncoder
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
query = "What is the capital of France?"
candidates = [
"Paris is the capital and largest city of France.",
"France is a country in Western Europe known for its wine.",
"The Eiffel Tower is a famous landmark located in Paris.",
"Berlin is the capital of Germany.",
]
# Score each (query, document) pair jointly
pairs = [[query, doc] for doc in candidates]
scores = model.predict(pairs)
# Sort candidates by re-rank score descending
ranked = sorted(zip(scores, candidates), reverse=True)
for score, doc in ranked:
print(f"{score:.4f} {doc[:60]}")How this code works
This code demonstrates how to use a re-ranking model to improve the precision of document retrieval, a core component of advanced RAG systems. Its job is to take a user's query and a list of candidates (potential answers or documents), then intelligently sort these candidates to show the most relevant ones first.
The process begins by loading a specialized neural network model using CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2"). This particular model is pre-trained to understand the relationship between a query and a document, evaluating how well they match. To do this, the code prepares pairs of [query, document] for each candidate. The model then calculates a scores array by calling model.predict(pairs), assigning each pair a numerical relevance score. A subtle but crucial aspect is that CrossEncoder models evaluate the query and document together for a more precise relevance judgment, unlike simpler models that might just compare their independent embeddings. Finally, zip(scores, candidates) links each score back to its original document, and sorted(..., reverse=True) arranges them so the documents with the highest (most relevant) scores appear at the top.
Production-grade example
Adds retries on rate limits, graceful degradation on bad requests, and structured latency/cost logging.
# cohere==5.x, tenacity==8.x
import os
import time
import logging
from typing import List
import cohere
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
co = cohere.Client(api_key=os.environ["COHERE_API_KEY"])
@retry(
stop=stop_after_attempt(3),
wait=wait_exponential(multiplier=1, min=1, max=10),
retry=retry_if_exception_type(cohere.errors.TooManyRequestsError),
)
def rerank_documents(
query: str,
documents: List[str],
top_n: int = 5,
model: str = "rerank-english-v3.0",
) -> List[dict]:
if not documents:
logger.warning("rerank called with empty document list, returning empty")
return []
start = time.perf_counter()
try:
response = co.rerank(
query=query,
documents=documents,
top_n=top_n,
model=model,
)
except cohere.errors.BadRequestError as exc:
logger.error("Cohere rerank bad request", extra={"query_len": len(query), "doc_count": len(documents), "error": str(exc)})
# Graceful degradation: return original order truncated to top_n
return [{"index": i, "document": documents[i], "relevance_score": None} for i in range(min(top_n, len(documents)))]
latency_ms = (time.perf_counter() - start) * 1000
logger.info(
"rerank completed",
extra={
"model": model,
"input_docs": len(documents),
"top_n": top_n,
"latency_ms": round(latency_ms, 1),
"billed_units": response.meta.billed_units.search_units if response.meta else None,
},
)
return [
{"index": r.index, "document": documents[r.index], "relevance_score": r.relevance_score}
for r in response.results
]
if __name__ == "__main__":
query = "How do I reset my password?"
docs = [
"To reset your password, click Forgot Password on the login page.",
"Our refund policy allows returns within 30 days.",
"Password requirements: 8 characters, one uppercase, one number.",
"Contact support at [email protected] for account issues.",
]
results = rerank_documents(query, docs, top_n=2)
for r in results:
print(f"Score: {r['relevance_score']:.4f} Doc: {r['document'][:60]}")How this code works
This code significantly improves RAG systems by re-ranking documents to present the most relevant information first. The rerank_documents function takes a user query and a list of documents, then leverages Cohere's rerank-english-v3.0 model to sort them by relevance. It connects to Cohere's API using an api_key loaded securely from os.environ. The function includes a top_n parameter to specify how many of the highest-scoring documents to return, crucial for efficiency. It also handles edge cases, such as an empty document list, by returning an empty result immediately.
A subtle but important feature for production-grade robustness is the @retry decorator from tenacity. This decorator automatically reattempts the Cohere API call up to three times (stop_after_attempt) if it encounters a cohere.errors.TooManyRequestsError, waiting progressively longer (wait_exponential) to avoid overwhelming the server. If a cohere.errors.BadRequestError occurs (e.g., due to invalid input), the code gracefully degrades by returning the original documents, truncated to top_n, without a relevance score. logging is used throughout to record important details like latency_ms and billed_units for monitoring and debugging. The if __name__ == "__main__": block demonstrates calling the function and printing the re-ranked documents with their relevance_score.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a minimal re-ranking comparison script. Given a fixed query and 8 candidate documents, retrieve the top-3 using cosine similarity on embeddings, then re-rank all 8 using a local cross-encoder and take the top-3. Print both result sets side by side and note which documents moved up or down after re-ranking.
# sentence-transformers==2.7.0
from sentence_transformers import SentenceTransformer, CrossEncoder, util
query = "How do I cancel my subscription?"
candidates = [
"To cancel, go to Account Settings and click Cancel Subscription.",
"Subscription plans start at $9.99 per month.",
"You can pause your subscription for up to 3 months.",
"Our cancellation policy requires 30 days notice.",
"To delete your account, contact [email protected].",
"Upgrade or downgrade your plan at any time.",
"Cancellation takes effect at the end of the billing period.",
"Refunds are not issued for partial billing periods after cancellation.",
]
# TODO: Load a bi-encoder model (e.g., 'all-MiniLM-L6-v2')
# TODO: Encode the query and all candidates
# TODO: Compute cosine similarity scores and get top-3 by embedding similarity
# TODO: Load a cross-encoder model (e.g., 'cross-encoder/ms-marco-MiniLM-L-6-v2')
# TODO: Score all (query, candidate) pairs and get top-3 by re-rank score
# TODO: Print both top-3 lists and highlight any differences
Quick check
Why can a cross-encoder detect relevance signals that a bi-encoder misses?
You're building a RAG system. Initial retrieval returns 5 documents. Adding re-ranking barely changes the results. What does this most likely indicate?
A colleague proposes skipping re-ranking and instead passing all 100 retrieved documents to a long-context LLM. What is the strongest argument against this approach?