jnachi
Learning Hub
Python Development9 min readAdvanced

Advanced RAG: Hybrid Search, Reranking & Query Transformations

Elevate RAG accuracy from 70% to 95%+ using Hybrid Search (BM25 + Dense Vectors with RRF), Cross-Encoder Rerankers (Cohere), and HyDE query transformations.

Works with:Hybrid Search (BM25 + Vector)Cohere RerankReciprocal Rank Fusion (RRF)HyDE

Key Takeaways

  • Naive dense vector search often misses exact keyword queries (SKUs, error codes, part numbers); Hybrid Search combines BM25 keyword matching with dense vectors
  • Reciprocal Rank Fusion (RRF) merges disparate search score lists into a unified, balanced relevance ranking
  • Cross-Encoder Rerankers (Cohere Rerank, BGE-Reranker) score query-document pairs simultaneously, dramatically boosting top-3 precision
  • Hypothetical Document Embeddings (HyDE) prompts an LLM to generate a hypothetical answer before embedding, bridging the gap between question and answer vector spaces

The Diagnostic Context

Basic RAG fails in production when users search for exact technical terms (e.g., "Error ERR-90214") or ask abstract questions. Advanced RAG combines sparse lexical search (BM25), dense semantic search, neural rerankers, and query rewriting to achieve production-grade 95%+ retrieval precision.

The Core Technique

The Advanced RAG Pipeline (Hybrid + Reranking)

DIAGRAM / WORKFLOW
graph TD
    UserQuery["User Query: 'How to fix error code ERR-8012?'"] --> QueryRewrite["Query Transformation / HyDE"]
    
    QueryRewrite --> BM25["Sparse BM25 Search<br/>(Finds exact keyword 'ERR-8012')"]
    QueryRewrite --> Dense["Dense Vector Search<br/>(Finds semantic concepts 'fix error')"]
    
    BM25 --> RRF["Reciprocal Rank Fusion (RRF)<br/>Combines Top 50 Chunks"]
    Dense --> RRF
    
    RRF --> Reranker["Cross-Encoder Reranker (Cohere / BGE)<br/>Deep Query-Document Re-scoring"]
    Reranker --> TopK["Top 3 High-Precision Chunks"]
    TopK --> LLM["LLM Synthesis"]
    LLM --> FinalResponse["Precise Fact-Checked Answer"]

Two-Stage Retrieval: Fetch Wide (50) $\rightarrow$ Rerank Deep (3)

  1. Stage 1: Broad Retrieval (Hybrid Search):

    • Dense vector search excels at broad concepts; BM25 keyword search excels at exact strings and serial numbers.
    • Combine the top 25 results from both algorithms using Reciprocal Rank Fusion (RRF): $$\text{RRF Score}(d) = \sum_{m \in M} \frac{1}{k + r_m(d)} \quad (k \approx 60)$$
  2. Stage 2: Precision Neural Reranking:

    • Vector embeddings compare queries and documents independently (Bi-Encoder).
    • A Cross-Encoder Reranker passes the query and chunk together through transformer attention layers, evaluating word-by-word semantic alignment.
    • Reranking 50 candidates down to the top 3 passages increases generation accuracy while reducing prompt token bloat by 80%.
5-Minute Activation Challenge

Try This Right Now

Implement Reciprocal Rank Fusion (RRF) in Python: Write a function that takes two ranked lists of document IDs (from BM25 and Vector search) and calculates combined RRF scores to produce a single unified top-ranked list.

Tip: Knowledge only becomes capability once you run the prompt yourself.

Comprehension Check

Test Your Instincts (3 Questions)

1

Why does naive vector embedding search often fail on queries containing specific product codes or error numbers (e.g., "Fix error ERR-4091")?

2

What is the architectural role of a Cross-Encoder Reranker in an Advanced RAG system?

3

How does Hypothetical Document Embeddings (HyDE) improve retrieval for abstract or open-ended questions?