Context Engineering: Why Retrieval Quality Decides Almost Everything About Your RAG System

Part 2 of the AI Engineer Series. Naive RAG pipelines fail at retrieval roughly 40% of the time. Chunking, hybrid search, reranking, and query rewriting determine quality more than any prompt change.

The bottleneck is retrieval, not the model

The standard RAG tutorial shows you four steps: embed documents, store them in a vector database, retrieve the top-k, generate. Then you ship it, run real queries against real users, and discover that roughly 40% of the time, your model produces a confident, well-structured answer grounded in the wrong documents.

Production benchmarks from 2026 consistently show the same thing. The retrieval step, not generation, is where most RAG systems fail. This is Part 2 of the AI Engineer Series, and it covers the layer that decides almost everything else.

If you want a broader treatment of why most RAG demos die in production, I have a deeper write-up in Part 1 of the RAG Series covering the systemic reasons. This post focuses specifically on the retrieval quality decisions you make when you sit down to build.

Chunking is the unit of retrieval

The chunk is what gets embedded, scored, and retrieved. Get this wrong and nothing else matters. The most common chunking strategies in 2026:

  • Fixed-size: split every N tokens. Fast, cheap, and a fine baseline.
  • Recursive: split by hierarchy (paragraphs, then sentences, then tokens). This is the LangChain default for a reason.
  • Structural: split by markdown headers, function definitions, table rows. Best for content that already has structure.
  • Semantic: split where embedding similarity between adjacent sentences drops. More expensive at index time but better recall on heterogeneous content.
  • Late chunking: embed the whole document first, then split the embedding sequence. Works when context windows are large enough to embed full documents.

Practical starting point: recursive chunking with 512-token chunks and 50-token overlap. Move to structural chunking when your content has clear structure. Only invest in semantic chunking when measurement says naive chunking is the bottleneck.

Hybrid search: BM25 plus dense vectors

Dense vector embeddings are great at semantic similarity. They handle "how do I cancel my subscription" matching a document titled "account termination policy." But they compress an entire chunk into a fixed-size vector, which loses information.

This is where they fail: exact terms, product codes, function names, rare words. A user searching for "TPC-3402 error" wants the document that mentions TPC-3402, even if it shares no semantic neighbors. Dense vectors will miss it. BM25, the term-frequency-based ranker that powers Elasticsearch, Solr, and most production search systems, will not.

Run both retrievers and fuse them with Reciprocal Rank Fusion:

from collections import defaultdict

def reciprocal_rank_fusion(rankings, k=60):
    """Combine multiple ranked lists into one.
    
    rankings: list of lists, each containing ranked doc IDs.
    Returns: list of (doc_id, score) sorted by fused score.
    """
    scores = defaultdict(float)
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking):
            scores[doc_id] += 1.0 / (k + rank)
    return sorted(scores.items(), key=lambda x: -x[1])

# Usage
dense_results = vector_db.search(query_embedding, top_k=50)
sparse_results = bm25_index.search(query_text, top_k=50)
fused = reciprocal_rank_fusion([
    [d.id for d in dense_results],
    [d.id for d in sparse_results],
])

RRF is the strongest default. It does not require score normalization, it does not require tuning weights between rankers, and it is robust to skewed score distributions.

Reranking with cross-encoders

Embedding models are bi-encoders. They encode the query and each document independently, then compare them with cosine similarity. The query never gets to interact with the document during encoding. Cross-encoders fix that. They take query and document together as a single input and produce a relevance score that has seen both.

The cost is that cross-encoders cannot be precomputed. You have to run them at query time, one document at a time. So the pattern is: retrieve a broad candidate set with cheap bi-encoders (top 20 to 50), then rerank that smaller set with the cross-encoder, and pass the top 5 to the model.

from sentence_transformers import CrossEncoder

reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")

def retrieve_and_rerank(query: str, k_retrieve: int = 30, k_final: int = 5):
    # Broad recall
    candidates = vector_db.search(query, top_k=k_retrieve)
    
    # Precision via reranking
    pairs = [(query, doc.text) for doc in candidates]
    scores = reranker.predict(pairs)
    
    ranked = sorted(zip(candidates, scores), key=lambda x: -x[1])
    return [doc for doc, _ in ranked[:k_final]]

Production-grade reranker options in 2026: BAAI/bge-reranker-v2-m3 for self-hosting, Cohere Rerank 3.5, and Jina Reranker v2. Expect 10-20% improvement on retrieval metrics. If you do not see that lift, the bottleneck is upstream in your chunking or embedding choice.

Lost in the middle

Even with perfect retrieval, models do not attend uniformly across the context window. They attend most to the beginning and end. Putting your most relevant document in position 5 of 10 measurably hurts answer quality.

So order retrieved chunks deliberately. Highest relevance first, second highest last, weaker ones in the middle. This single change can lift answer quality without touching anything else in your stack.

Query rewriting

Users write bad queries. "What was that thing about pricing?" does not retrieve well. The fix is to rewrite the query with a small, cheap LLM call before retrieval.

Three patterns that work:

  • Expansion: rewrite the query into 2-3 variations covering different phrasings, expand abbreviations, resolve pronouns from chat history.
  • HyDE (Hypothetical Document Embeddings): ask the model to write a hypothetical answer to the query, then embed that answer and use it as the search vector. Often outperforms embedding the question directly because answers and documents look more alike than questions and documents.
  • Decomposition: split a compound question into sub-questions, retrieve for each, and merge results.
def hyde_search(query: str):
    hypothetical = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=200,
        messages=[{
            "role": "user",
            "content": f"Write a hypothetical paragraph that would answer: {query}"
        }],
    ).content[0].text
    
    # Search using the hypothetical answer, not the question
    return vector_db.search(embed(hypothetical), top_k=10)

Embedding model choice

Bigger is not always better. Things that actually matter:

  • Domain match. A general embedding model can underperform a domain-tuned one (medical, legal, code) by a wide margin.
  • Dimension size. A 1536d embedding takes 4x more storage and ANN search time than a 384d one. Matryoshka Representation Learning (MRL) embeddings let you truncate to smaller dimensions without retraining.
  • Asymmetric prefixes. Some models expect query: and passage: prefixes. Using the wrong prefix silently degrades quality.
  • Multilingual. If your corpus or queries are not English-only, do not use an English-only model. BGE-M3 and multilingual-e5 are strong defaults.

The MTEB leaderboard on Hugging Face is the standard reference for picking an embedding model. Filter by your task type (retrieval, not classification) and the language you need.

Evaluate the retriever separately from the generator

The cardinal sin of RAG debugging is mixing retrieval and generation quality. Build two eval sets:

  • Retrieval evals: a set of queries with known relevant document IDs. Measure recall@k (did the right doc end up in your top-k?) and context precision (what fraction of retrieved docs were actually relevant?).
  • Generation evals: pass the retrieved context to the model and measure answer quality with an LLM judge or human reviewer.

Without this separation, you cannot tell if your changes are helping or hurting. The RAGAS library packages this evaluation cleanly and is the most adopted choice in 2026. We will cover evals in much more depth in Part 4.

What to build this week

Take last week's harness and add a retrieval layer:

  • Pick a non-trivial corpus (your company docs, arXiv papers, a Wikipedia subset)
  • Build two baseline pipelines: naive (dense-only, no reranking) and full (hybrid + reranking + query rewriting)
  • Construct a 30-50 question eval set with known answers
  • Measure recall@10 and answer quality on both pipelines
  • Report the delta

You will almost certainly see a 15-30 point improvement on retrieval metrics. That is your case for putting the full pipeline in production.

What is next

Part 3: structured outputs and fallback chains. Once your model has the right context, the next failure mode is the output shape itself. We will look at JSON mode, tool calling, constrained decoding, and the fallback patterns that keep the rest of your code stable when the model returns something unexpected.

Previous in the series

References

  1. Contextual Retrieval from Anthropic. The technique that combines BM25 and dense retrieval with contextual chunk embeddings for state-of-the-art recall.
  2. Lost in the Middle by Liu et al. The paper that documented how transformer models attend non-uniformly to long contexts.
  3. Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE) by Gao et al. The original paper on hypothetical document embeddings.
  4. MTEB leaderboard. The standard reference for picking an embedding model, filtered by task and language.
  5. RAGAS documentation. The most adopted framework for evaluating RAG pipelines, with metrics for both retrieval and generation.
  6. Why 80% of RAG Demos Die in Production. Part 1 of my RAG Series, covering the systemic reasons retrieval pipelines fail.

Subscribe to Vivek Wisdom

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe