jnachi
Learning Hub
Agentic AI & RAG8 min readAdvanced

Semantic Chunking & Cross-Encoder Rerankers

Eliminate chunk fragmentation and boost RAG precision using semantic distance splitting, parent-child document retrieval, and Cross-Encoder reranking.

Works with:Cohere RerankBGE-Reranker-LargeSentenceTransformersLangChain TextSplitters

Key Takeaways

  • Fixed-size character chunking splits paragraphs mid-thought, severing context and reducing embedding quality
  • Semantic chunking measures embedding similarity between consecutive sentences, creating split points only when topic shifts occur
  • Bi-Encoder models are fast for retrieving top-100 candidates; Cross-Encoder models perform full joint attention to rerank the top-5 chunks with extreme precision
  • Parent-Child retrieval indexes small sub-chunks for accurate vector matching but feeds the full parent section to the LLM

The Diagnostic Context

The quality of any RAG pipeline is bottlenecked by the accuracy of retrieved context. Feeding irrelevant or fragmented chunks forces the generator to hallucinate. Two techniques dramatically improve context quality: semantic chunking and Cross-Encoder reranking.

The Core Technique

Deploying a Cross-Encoder Reranker

Bi-encoders encode query and passage separately ($q \cdot p$), missing nuanced cross-token relationships. Cross-encoders process $(q, p)$ simultaneously with full transformer self-attention:

PYTHON
from sentence_transformers import CrossEncoder

# Load specialized enterprise reranker
reranker = CrossEncoder('BAAI/bge-reranker-large')

query = "What is the policy for secondary data retention on AWS S3?"
candidate_chunks = [
    "AWS S3 bucket versioning preserves historical file revisions.",
    "Secondary data archives must be permanently purged after 90 days per compliance policy SEC-402.",
    "Data transmission over TLS 1.3 is enforced for all REST endpoints.",
    "Customer billing invoices are retained for 7 years in cold Glacier storage."
]

# Pair query with each candidate chunk
query_passage_pairs = [[query, chunk] for chunk in candidate_chunks]

# Compute precise relevance scores
scores = reranker.predict(query_passage_pairs)

# Rank descending
ranked_results = sorted(
    zip(candidate_chunks, scores), 
    key=lambda x: x[1], 
    reverse=True
)

for rank, (chunk, score) in enumerate(ranked_results, 1):
    print(f"Rank {rank} (Score: {score:.4f}): {chunk}")

Parent-Child Document Indexing

To solve the trade-off between search precision and synthesis context:

  1. Break a 2,000-word document into small 100-word child chunks.
  2. Vectorize and index child chunks with a pointer metadata tag:
    CODE / PROMPT
    parent_id
    .
  3. When a child chunk is retrieved, fetch the complete 2,000-word parent document to supply to the LLM prompt.
5-Minute Activation Challenge

Try This Right Now

Inspect a 5-page enterprise PDF. Test splitting it with standard 500-char fixed splitting vs recursive markdown header splitting, and observe how tables and headers stay intact!

Tip: Knowledge only becomes capability once you run the prompt yourself.

Comprehension Check

Test Your Instincts (1 Questions)

1

Why is a Cross-Encoder reranker typically used ONLY on the top 20-50 candidates rather than the entire million-vector database?