Architectures of Choice — Adaptive, Corrective, and Modular RAG with LangGraph
Part 3 of the RAG Series: Stop building RAG systems that do the same thing for every query. Adaptive RAG, Corrective RAG, and Modular RAG — built with LangGraph state machines — give you systems that route, self-correct, and stay composable.
This is Part 3 of a 5-part series on building production-grade RAG systems in 2026. Part 1 covered the Knowledge Runtime. Part 2 covered the Retrieval Multiverse. Now we build the decision-making layer on top.
The Problem with Fixed Pipelines
Every RAG system we've built so far does the same thing regardless of what the user asks: retrieve, then generate. A user typing "what's 2+2?" gets the same retrieval process as a user asking "compare the contractual obligations across our three supplier agreements." One of those queries needs retrieval. The other doesn't. Running the same expensive pipeline for both wastes compute, adds latency, and dilutes output quality.
This is the fundamental flaw of static RAG pipelines: they're one-size-fits-all systems in a world where queries are wildly heterogeneous.
LangGraph changes this. By modelling the RAG pipeline as a state machine with conditional edges, you can build systems that inspect each query and route it to the right strategy, or loop back and self-correct if the first attempt wasn't good enough.
In this post we'll implement three patterns: Adaptive RAG (routing), Corrective RAG (self-correction), and Modular RAG (composability). All three use LangGraph.
LangGraph Primer: Thinking in Graphs
Before the patterns, a quick mental model. In LangGraph, your pipeline is a directed graph where:
- Nodes are Python functions that transform state (e.g., "retrieve documents", "grade relevance", "generate answer")
- Edges are transitions between nodes, which can be conditional, "if the grader says documents are irrelevant, go to rewrite; otherwise go to generate"
- State is a typed dictionary that flows through every node, accumulating context
This mental model turns a RAG pipeline from a linear chain into a programmable workflow with loops, branches, and decision points.
Pattern 1: Adaptive RAG, Route Before You Retrieve
Adaptive RAG classifies every incoming query into one of three paths before doing anything else:
- Direct answer, the LLM can answer from training knowledge (no retrieval needed)
- Local RAG, the answer is in your vector store
- Web search, the answer requires current information beyond your index
The router is an LLM with a structured output schema, it costs one cheap LLM call per query and saves you expensive retrieval on every simple question.
from typing import Literal
from pydantic import BaseModel, Field
from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from typing import TypedDict, List
from langchain_core.documents import Document
# --- State definition ---
class RAGState(TypedDict):
question: str
route: str
documents: List[Document]
answer: str
generation_count: int
# --- Router schema ---
class RouteQuery(BaseModel):
"""Route a user query to the most appropriate data source."""
datasource: Literal["vectorstore", "web_search", "direct_answer"] = Field(
description="Route to vectorstore for internal docs, web_search for current events, direct_answer for simple factual questions."
)
# --- Router node ---
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0) # Cheap model for routing
structured_llm_router = llm.with_structured_output(RouteQuery)
router_prompt = ChatPromptTemplate.from_messages([
("system", """You are an expert at routing questions.
Route to 'vectorstore' for questions about internal policies, products, or documents.
Route to 'web_search' for questions requiring current events or real-time data.
Route to 'direct_answer' for simple factual questions the LLM can answer directly."""),
("human", "{question}"),
])
query_router = router_prompt | structured_llm_router
def route_question(state: RAGState) -> RAGState:
result = query_router.invoke({"question": state["question"]})
return {**state, "route": result.datasource}
def decide_route(state: RAGState) -> str:
"""Conditional edge: branches based on router output."""
return state["route"] # Returns "vectorstore", "web_search", or "direct_answer"
# --- Build the graph ---
workflow = StateGraph(RAGState)
workflow.add_node("router", route_question)
workflow.add_node("retrieve", retrieve_from_vectorstore) # your retrieval node
workflow.add_node("web_search", search_the_web) # Tavily or similar
workflow.add_node("generate", generate_answer) # LLM generation node
workflow.set_entry_point("router")
workflow.add_conditional_edges(
"router",
decide_route,
{
"vectorstore": "retrieve",
"web_search": "web_search",
"direct_answer": "generate",
}
)
workflow.add_edge("retrieve", "generate")
workflow.add_edge("web_search", "generate")
workflow.add_edge("generate", END)
adaptive_rag = workflow.compile()
result = adaptive_rag.invoke({"question": "What is our refund policy?", "generation_count": 0})
The key insight is the routing model: you use gpt-4o-mini (cheap, fast) for routing, and only invoke the more expensive retrieval + generation pipeline when it's actually needed. On a workload with 30% simple queries, this alone cuts costs meaningfully.
Pattern 2: Corrective RAG (CRAG), Self-Correct When Retrieval Fails
Adaptive RAG routes queries intelligently. But even with the right strategy, retrieval can fail, the documents retrieved might be irrelevant or insufficient. Corrective RAG adds a grading loop: after retrieval, an LLM grader assesses each document's relevance, and if too many are irrelevant, the system reformulates the query and retrieves again.
from pydantic import BaseModel, Field
from typing import Literal
# --- Relevance grader schema ---
class GradeDocuments(BaseModel):
"""Binary score for document relevance."""
binary_score: Literal["yes", "no"] = Field(
description="Is the document relevant to the question? 'yes' or 'no'."
)
structured_grader = llm.with_structured_output(GradeDocuments)
grade_prompt = ChatPromptTemplate.from_messages([
("system", "Grade whether the document is relevant to the question. Be strict, only 'yes' if the document directly helps answer the question."),
("human", "Question: {question}\n\nDocument:\n{document}"),
])
doc_grader = grade_prompt | structured_grader
# --- Grading node ---
def grade_documents(state: RAGState) -> RAGState:
"""Filter retrieved documents, flag if too few are relevant."""
question = state["question"]
documents = state["documents"]
relevant_docs = []
for doc in documents:
score = doc_grader.invoke({"question": question, "document": doc.page_content})
if score.binary_score == "yes":
relevant_docs.append(doc)
# If fewer than 2 relevant docs found, flag for query rewrite
needs_rewrite = len(relevant_docs) < 2
return {**state, "documents": relevant_docs, "route": "rewrite" if needs_rewrite else "generate"}
# --- Query rewriter node ---
rewrite_prompt = ChatPromptTemplate.from_messages([
("system", "Rewrite the question to improve retrieval. Make it more specific and use different terminology."),
("human", "Original question: {question}"),
])
rewriter = rewrite_prompt | ChatOpenAI(model="gpt-4o-mini")
def rewrite_query(state: RAGState) -> RAGState:
new_question = rewriter.invoke({"question": state["question"]}).content
# Increment counter to prevent infinite loops
return {**state, "question": new_question, "generation_count": state["generation_count"] + 1}
def should_rewrite(state: RAGState) -> str:
"""Conditional edge: rewrite if needed, but cap at 2 retries."""
if state["route"] == "rewrite" and state["generation_count"] < 2:
return "rewrite"
return "generate"
# Add CRAG nodes to the workflow
workflow.add_node("grade_documents", grade_documents)
workflow.add_node("rewrite_query", rewrite_query)
workflow.add_edge("retrieve", "grade_documents")
workflow.add_conditional_edges("grade_documents", should_rewrite, {
"rewrite": "rewrite_query",
"generate": "generate",
})
workflow.add_edge("rewrite_query", "retrieve") # Loop back to retrieve with new query
The generation_count guard is critical, without it, a bad query can loop indefinitely. In practice, two retries is usually sufficient; if two rewrites still yield irrelevant documents, the question probably can't be answered from your index and you should escalate to web search or return a graceful "I don't know."
Pattern 3: Modular RAG, Swappable Components
The final pattern isn't about runtime behaviour, it's about how you build for long-term maintainability. Modular RAG means designing your pipeline so you can swap any component (embedding model, vector store, LLM, retriever) without rewriting the whole graph.
LangGraph's node-based design makes this natural, each node is just a function. The trick is making sure your state schema and node interfaces stay stable:
from langchain_core.retrievers import BaseRetriever
from langchain_core.language_models import BaseChatModel
from langchain_core.embeddings import Embeddings
class RAGPipeline:
"""
Modular RAG pipeline, swap any component without changing the graph structure.
Usage:
# Production config
pipeline = RAGPipeline(
retriever=hybrid_retriever, # EnsembleRetriever from Part 2
llm=ChatOpenAI(model="gpt-4o"),
)
# Experimental config, test a new model without touching the graph
pipeline = RAGPipeline(
retriever=hybrid_retriever,
llm=ChatAnthropic(model="claude-3-5-sonnet-20241022"),
)
"""
def __init__(self, retriever: BaseRetriever, llm: BaseChatModel):
self.retriever = retriever
self.llm = llm
self.graph = self._build_graph()
def _retrieve(self, state: RAGState) -> RAGState:
docs = self.retriever.invoke(state["question"])
return {**state, "documents": docs}
def _generate(self, state: RAGState) -> RAGState:
context = "\n\n".join(d.page_content for d in state["documents"])
prompt = f"Answer based only on this context:\n\n{context}\n\nQuestion: {state['question']}"
answer = self.llm.invoke(prompt).content
return {**state, "answer": answer}
def _build_graph(self) -> any:
workflow = StateGraph(RAGState)
workflow.add_node("retrieve", self._retrieve)
workflow.add_node("generate", self._generate)
workflow.set_entry_point("retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)
return workflow.compile()
def invoke(self, question: str) -> str:
result = self.graph.invoke({"question": question, "generation_count": 0})
return result["answer"]
# Swap the LLM, graph structure unchanged
pipeline_v1 = RAGPipeline(retriever=hybrid_retriever, llm=ChatOpenAI(model="gpt-4o"))
pipeline_v2 = RAGPipeline(retriever=hybrid_retriever, llm=ChatOpenAI(model="gpt-4o-mini"))
# A/B test both against your golden dataset (from Part 1) to compare cost vs quality
This pattern pays dividends the moment you want to run an A/B test between two LLMs, migrate to a cheaper embedding model, or replace your vector store. Because nodes are functions and the graph is compiled from them, the topology stays stable across component swaps.
Which Pattern When
These three patterns aren't alternatives, they stack. A mature production system uses all three simultaneously: Adaptive RAG decides the route, CRAG catches retrieval failures, and Modular RAG ensures the whole thing stays maintainable as components evolve. The LangGraph primitives, state, nodes, conditional edges, compose cleanly because they operate on the same typed state dict throughout.
| Problem | Pattern | Key mechanism |
|---|---|---|
| Wasting compute on simple queries | Adaptive RAG | Router LLM + conditional edges |
| Retrieved docs are irrelevant | Corrective RAG (CRAG) | Grader + rewrite loop with cap |
| Hard to swap models/stores | Modular RAG | DI pattern + stable state schema |
Up Next: Agentic RAG
Parts 1-3 build a RAG system that is smart about ingestion, sophisticated about retrieval, and adaptive in its architecture. But it still responds to a single query in a single turn.
In Part 4, we cross into Agentic RAG, systems that decompose complex questions into sub-queries, iterate over multiple retrieval rounds, hand off between specialised agents, and use tools like SQL databases and MCP servers to reach answers that a passive pipeline simply cannot. The shift from retrieval to reasoning.
Enjoying the series? Share it with the person on your team who's still hardcoding their chunk size.