The Retrieval Multiverse — Vectors, Graphs, and Why Hybrid Is the Only Way to Go
Part 2 of the RAG Series: Production RAG in 2026 is never vector-only. Learn why hybrid search, GraphRAG, and late chunking each solve problems that pure vector search cannot — with LangChain and LangGraph code throughout.
This is Part 2 of a 5-part series on building production-grade RAG systems in 2026. Part 1 covered the Knowledge Runtime — clean ingestion, metadata enrichment, and evaluation. Now we go deeper into retrieval itself.
The Vector Store Trap
Here's a scenario that plays out constantly in production RAG systems: a user types the exact product SKU from an internal catalogue — say, PRD-7742-B — and asks for its warranty terms. The vector retriever confidently returns chunks about similar products, because embeddings are trained to find semantic neighbours. The exact string match? Buried, or missing entirely.
Vector search is remarkably good at finding what something means. It's terrible at finding what something is — exact identifiers, part numbers, version strings, names, dates. This is the vector store trap: people reach for embeddings as the universal answer to retrieval, and then wonder why their system fails on the queries that matter most to real users.
In 2026, mature production RAG systems use at least two retrieval strategies in concert. Let's map the landscape.
Strategy 1: Hybrid Search — Best of Both Worlds
Hybrid search combines dense vector retrieval (semantic similarity via embeddings) with sparse BM25 retrieval (keyword matching, the algorithm behind classic search engines). The two complement each other perfectly: vectors catch meaning, BM25 catches exact terms.
Think of it this way: dense search finds "what the user is really asking about," while sparse search finds "the exact words the user typed." When fused via Reciprocal Rank Fusion (RRF) — a reranking method that combines ranked lists from both retrievers — the result is consistently better than either alone.
LangChain's EnsembleRetriever makes this almost trivially easy to implement:
from langchain_community.retrievers import BM25Retriever
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings
from langchain.retrievers import EnsembleRetriever
from langchain_core.documents import Document
# Assume `chunks` is your list of enriched LangChain Documents from Part 1
# 1. Dense retriever — semantic similarity via embeddings
vectorstore = Chroma.from_documents(chunks, OpenAIEmbeddings())
dense_retriever = vectorstore.as_retriever(search_kwargs={"k": 5})
# 2. Sparse retriever — BM25 keyword matching
bm25_retriever = BM25Retriever.from_documents(chunks)
bm25_retriever.k = 5
# 3. Ensemble — fuses results via Reciprocal Rank Fusion
# weights: [BM25 contribution, vector contribution]
# 0.4 / 0.6 slightly favours semantic — tune based on your query distribution
hybrid_retriever = EnsembleRetriever(
retrievers=[bm25_retriever, dense_retriever],
weights=[0.4, 0.6],
)
# Query — now catches both semantic meaning AND exact terms
results = hybrid_retriever.invoke("warranty terms for PRD-7742-B")
for doc in results:
print(doc.metadata.get("source"), "|", doc.page_content[:100])
The weights parameter is your main tuning lever. If your users tend to search with exact terminology (e.g., internal tools, legal documents, code references), weight BM25 higher — try [0.6, 0.4]. If queries are more conversational and conceptual, lean on the dense retriever.
When to use hybrid: Almost always. The cost of adding BM25 to a vector pipeline is near-zero — it runs in memory, requires no additional infrastructure, and typically improves recall by 15–30% on real-world queries. The only exception is when your entire query workload is purely semantic (e.g., conceptual Q&A over essays with no proper nouns).
Strategy 2: GraphRAG — When Relationships Matter More Than Similarity
Vector search finds chunks that are similar to your query. It cannot answer questions that require traversing relationships between entities.
Consider: "Which of our suppliers have contracts expiring in Q3 that are also flagged for compliance review?" A vector store will find chunks about supplier contracts, chunks about compliance flags, and chunks about Q3 timelines — all separately. It cannot connect them. A knowledge graph can, because it stores entities and their relationships explicitly, allowing multi-hop traversal: Supplier → Contract → ExpiryDate AND Supplier → ComplianceFlag → Status.
This is the GraphRAG insight: some queries need structured traversal, not nearest-neighbour search.
from langchain_neo4j import Neo4jGraph, GraphCypherQAChain
from langchain_openai import ChatOpenAI
# Connect to Neo4j (can also use in-memory graphs for smaller use cases)
graph = Neo4jGraph(
url="bolt://localhost:7687",
username="neo4j",
password="your-password",
)
# GraphCypherQAChain: LLM generates a Cypher query, executes it, generates answer
# The LLM translates natural language → Cypher → structured result → answer
llm = ChatOpenAI(model="gpt-4o", temperature=0)
chain = GraphCypherQAChain.from_llm(
llm=llm,
graph=graph,
verbose=True, # Shows generated Cypher — useful for debugging
validate_cypher=True,
)
# Multi-hop query — impossible with pure vector search
result = chain.invoke({
"query": "Which suppliers have contracts expiring in Q3 and are flagged for compliance review?"
})
print(result["result"])
# Under the hood, the LLM generates something like:
# MATCH (s:Supplier)-[:HAS_CONTRACT]->(c:Contract)
# WHERE c.expiry_quarter = 'Q3'
# AND EXISTS { MATCH (s)-[:HAS_FLAG]->(:ComplianceFlag {status: 'active'}) }
# RETURN s.name, c.expiry_date
When to use GraphRAG: Relationship-heavy domains where entities connect to other entities in ways that matter. The canonical use cases are fraud detection (transaction → account → network → other flagged accounts), legal discovery (document → cited case → jurisdiction → regulation), and org chart / project management queries (person → team → project → dependency). If your queries frequently involve "which X relates to Y via Z," you need a graph.
When NOT to use GraphRAG: General prose Q&A, creative content, or anything where the answer lives in a single contiguous chunk. Graph construction is expensive (you need entity extraction to build it), maintenance is non-trivial, and the added complexity is only justified when multi-hop traversal is genuinely needed.
Strategy 3: Late Chunking — Preserving Global Context
Both hybrid search and GraphRAG operate on pre-chunked data. Late chunking inverts this: instead of chunking documents first and then embedding the chunks, you pass the entire document to a long-context embedding model first, and chunk afterwards. The resulting embeddings carry global document context that early-chunking approaches lose entirely.
The intuition: a chunk saying "as mentioned above" or "this exception applies only to the rule in Section 2" is semantically orphaned when embedded in isolation. Late chunking means the embedding model saw the full document when it encoded that chunk, so the embedding reflects what "mentioned above" actually refers to.
from langchain_openai import OpenAIEmbeddings
from langchain_core.documents import Document
from langchain_community.vectorstores import Chroma
def late_chunk_embed(full_document: str, chunks: list[Document]) -> list[Document]:
"""
Late chunking: embed full document first to create context-aware embeddings.
Uses the full document as prefix context for each chunk embedding.
Note: requires a long-context embedding model (e.g. text-embedding-3-large
supports up to 8191 tokens). For very long docs, use sliding window approach.
"""
embeddings_model = OpenAIEmbeddings(model="text-embedding-3-large")
# Embed each chunk WITH the full document as prefix context
# This ensures every chunk embedding reflects its position in the whole
contextualised_texts = [
f"Document context:\n{full_document[:2000]}\n\n---\n\nChunk:\n{chunk.page_content}"
for chunk in chunks
]
# Get embeddings — these are now context-aware
chunk_embeddings = embeddings_model.embed_documents(contextualised_texts)
# Attach contextual embedding as metadata for downstream use
for chunk, embedding in zip(chunks, chunk_embeddings):
chunk.metadata["contextual_embedding"] = embedding
return chunks
# Usage — apply after chunking, before indexing
contextualised_chunks = late_chunk_embed(full_document_text, chunks)
# Index as normal — the metadata-stored embeddings can be used for custom retrieval
vectorstore = Chroma.from_documents(contextualised_chunks, OpenAIEmbeddings())
When to use late chunking: Documents with heavy cross-references, complex arguments that build over multiple sections, legal contracts, or technical specifications. If your chunks frequently say things like "as stated in Section 3" or "this overrides the previous rule," early chunking is losing critical context and late chunking will meaningfully improve retrieval quality.
Putting It Together: A Retrieval Decision Framework
The three strategies aren't mutually exclusive — the best production systems layer them:
| Query Type | Best Strategy |
|---|---|
| Conceptual / conversational ("explain our leave policy") | Dense vector only |
| Exact terms mixed with context ("SKU X72 return conditions") | Hybrid (BM25 + vector) |
| Multi-hop relationship ("supplier A's contracts linked to compliance flags") | GraphRAG (Neo4j + Cypher) |
| Cross-referencing documents ("this clause, per Section 3.2...") | Late chunking + hybrid |
| All of the above, dynamically | Adaptive RAG (Part 3) |
Up Next: Architectures of Choice
Now that we have a rich retrieval toolkit, the question becomes: how does the system decide which strategy to use at runtime?
In Part 3, we build the decision-making layer — Adaptive RAG, Corrective RAG (CRAG), and Modular RAG — using LangGraph's state machine primitives. We'll build a pipeline that routes queries to the right retrieval strategy automatically, self-corrects when retrieved documents are irrelevant, and stays composable so you can swap components without rewriting the whole system.
If you're building RAG in production and found this useful, share it — someone on your team probably needs this too.