Semantic Chunking & Cross-Encoder Rerankers
Eliminate chunk fragmentation and boost RAG precision using semantic distance splitting, parent-child document retrieval, and Cross-Encoder reranking.
Key Takeaways
- Fixed-size character chunking splits paragraphs mid-thought, severing context and reducing embedding quality
- Semantic chunking measures embedding similarity between consecutive sentences, creating split points only when topic shifts occur
- Bi-Encoder models are fast for retrieving top-100 candidates; Cross-Encoder models perform full joint attention to rerank the top-5 chunks with extreme precision
- Parent-Child retrieval indexes small sub-chunks for accurate vector matching but feeds the full parent section to the LLM
The Diagnostic Context
The quality of any RAG pipeline is bottlenecked by the accuracy of retrieved context. Feeding irrelevant or fragmented chunks forces the generator to hallucinate. Two techniques dramatically improve context quality: semantic chunking and Cross-Encoder reranking.
The Core Technique
Deploying a Cross-Encoder Reranker
Bi-encoders encode query and passage separately ($q \cdot p$), missing nuanced cross-token relationships. Cross-encoders process $(q, p)$ simultaneously with full transformer self-attention:
from sentence_transformers import CrossEncoder
# Load specialized enterprise reranker
reranker = CrossEncoder('BAAI/bge-reranker-large')
query = "What is the policy for secondary data retention on AWS S3?"
candidate_chunks = [
"AWS S3 bucket versioning preserves historical file revisions.",
"Secondary data archives must be permanently purged after 90 days per compliance policy SEC-402.",
"Data transmission over TLS 1.3 is enforced for all REST endpoints.",
"Customer billing invoices are retained for 7 years in cold Glacier storage."
]
# Pair query with each candidate chunk
query_passage_pairs = [[query, chunk] for chunk in candidate_chunks]
# Compute precise relevance scores
scores = reranker.predict(query_passage_pairs)
# Rank descending
ranked_results = sorted(
zip(candidate_chunks, scores),
key=lambda x: x[1],
reverse=True
)
for rank, (chunk, score) in enumerate(ranked_results, 1):
print(f"Rank {rank} (Score: {score:.4f}): {chunk}")
Parent-Child Document Indexing
To solve the trade-off between search precision and synthesis context:
- Break a 2,000-word document into small 100-word child chunks.
- Vectorize and index child chunks with a pointer metadata tag: .CODE / PROMPT
parent_id
- When a child chunk is retrieved, fetch the complete 2,000-word parent document to supply to the LLM prompt.
Try This Right Now
Inspect a 5-page enterprise PDF. Test splitting it with standard 500-char fixed splitting vs recursive markdown header splitting, and observe how tables and headers stay intact!
Tip: Knowledge only becomes capability once you run the prompt yourself.