A second-stage retrieval refinement that takes the top results from initial retrieval (often semantic search) and reorders them using a more accurate but slower model — usually a cross-encoder.
Reranking exists because the first-stage retriever (vector search, BM25) optimizes for speed and recall over many documents, while a reranker optimizes for precision over the top 20-100. Cross-encoders consider query and document together, producing higher-quality relevance scores than the bi-encoder used for initial retrieval. The tradeoff: cross-encoders are 100-1000× slower per query, so they only run on top-N candidates. Cohere Rerank, BGE Reranker, and ms-marco-MiniLM are common choices.
Boosting RAG accuracy 20-40% by reranking the top-50 vector-search results with Cohere Rerank before passing the top-5 to the LLM.
Reranking is often the single highest-ROI addition to a RAG pipeline — it's cheaper than fine-tuning and gives measurable retrieval quality gains.
Need help implementing this in your business?
Get Started