Agentic RAG — From Passive Retrieval to Active Reasoning with LangGraph

Part 4 of the RAG Series: The jump from passive retrieval to active reasoning. Build LangGraph agents that decompose questions, iterate over retrieval loops, use SQL tools and MCP servers, and hand off between specialised agents.

This is Part 4 of a 5-part series on building production-grade RAG systems in 2026. Previous parts: Part 1 (Knowledge Runtime), Part 2 (Retrieval Multiverse), Part 3 (Adaptive/Corrective/Modular RAG).


From Lookup to Reasoning

Every RAG system we've built so far operates in a single pass: one question, one retrieval, one answer. For the vast majority of queries, that's fine. But consider a question like: "Based on our Q3 sales data, which product categories underperformed against the targets in our strategic plan, and what do our customer support tickets say about why?"

This requires at minimum: a SQL query against structured sales data, retrieval from a strategic planning document, retrieval from customer support ticket summaries, and synthesis across all three. A single-pass pipeline cannot do this. It retrieves one set of context and generates one answer. The answer will be incomplete, or worse, it will hallucinate the parts it couldn't retrieve.

This is where Agentic RAG begins. The shift is from a system that looks things up to one that reasons through a problem — decomposing the question, retrieving iteratively, using specialised tools, and synthesising across multiple sources.


Pattern 1: Multi-Step Reasoning via Query Decomposition

The first capability is decomposing a complex question into sub-questions that can each be answered independently, then synthesising the results. LangGraph handles this naturally because its state machine can loop.

from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from typing import TypedDict, List, Annotated
import operator

class AgentState(TypedDict):
    original_question: str
    sub_questions: List[str]
    retrieved_answers: Annotated[List[str], operator.add]  # Accumulates across iterations
    final_answer: str
    current_step: int

llm = ChatOpenAI(model="gpt-4o", temperature=0)

# --- Decomposition node ---
decompose_prompt = ChatPromptTemplate.from_messages([
    ("system", """Break the question into 2-4 independent sub-questions that together answer the original.
    Return ONLY a JSON array of strings. Example: ["sub-q 1", "sub-q 2"]"""),
    ("human", "{question}"),
])

def decompose_question(state: AgentState) -> AgentState:
    import json
    result = llm.invoke(
        decompose_prompt.format_messages(question=state["original_question"])
    ).content
    # Strip markdown fences if present
    result = result.strip().strip("```json").strip("```").strip()
    sub_questions = json.loads(result)
    return {**state, "sub_questions": sub_questions, "current_step": 0}

# --- Iterative retrieval node ---
def retrieve_for_subquestion(state: AgentState) -> AgentState:
    """Retrieve and answer the current sub-question."""
    step = state["current_step"]
    if step >= len(state["sub_questions"]):
        return state

    sub_q = state["sub_questions"][step]

    # Use your hybrid retriever from Part 2
    docs = hybrid_retriever.invoke(sub_q)
    context = "\n\n".join(d.page_content for d in docs)

    answer = llm.invoke(
        f"Answer this specific question based only on the context provided.\n\nContext:\n{context}\n\nQuestion: {sub_q}"
    ).content

    return {
        **state,
        "retrieved_answers": [f"Q: {sub_q}\nA: {answer}"],
        "current_step": step + 1,
    }

def has_more_subquestions(state: AgentState) -> str:
    """Conditional edge: keep iterating or synthesise."""
    if state["current_step"] < len(state["sub_questions"]):
        return "retrieve"
    return "synthesise"

# --- Synthesis node ---
def synthesise(state: AgentState) -> AgentState:
    all_qa = "\n\n---\n\n".join(state["retrieved_answers"])
    answer = llm.invoke(
        f"Synthesise a comprehensive answer to the original question using the sub-answers below.\n\n"
        f"Original question: {state['original_question']}\n\n"
        f"Sub-answers:\n{all_qa}"
    ).content
    return {**state, "final_answer": answer}

# Build graph
workflow = StateGraph(AgentState)
workflow.add_node("decompose", decompose_question)
workflow.add_node("retrieve", retrieve_for_subquestion)
workflow.add_node("synthesise", synthesise)

workflow.set_entry_point("decompose")
workflow.add_edge("decompose", "retrieve")
workflow.add_conditional_edges("retrieve", has_more_subquestions, {
    "retrieve": "retrieve",   # Loop: keep processing sub-questions
    "synthesise": "synthesise",
})
workflow.add_edge("synthesise", END)

multi_step_rag = workflow.compile()

The Annotated[List[str], operator.add] type in the state is LangGraph's way of saying "accumulate, don't overwrite" — each iteration appends to retrieved_answers rather than replacing it. This is the key pattern for iterative state accumulation in LangGraph.


Pattern 2: Tool Use — SQL Agent for Structured Data

Not all answers live in documents. Sales figures, inventory counts, customer records, and operational metrics live in databases. The right tool for structured data is SQL, not vector search. LangChain's SQL agent wraps this capability in a form that LangGraph can use as a node.

from langchain_community.utilities import SQLDatabase
from langchain_community.agent_toolkits import create_sql_agent
from langchain_openai import ChatOpenAI

# Connect to your database
db = SQLDatabase.from_uri("postgresql://user:password@localhost/analytics_db")

# Create a SQL agent — it generates and executes SQL queries from natural language
sql_agent = create_sql_agent(
    llm=ChatOpenAI(model="gpt-4o", temperature=0),
    db=db,
    agent_type="openai-tools",
    verbose=True,      # Shows the generated SQL — always review in production
)

# Use as a LangGraph node
def query_database(state: AgentState) -> AgentState:
    """
    Node for structured data queries. Routes here when the sub-question
    requires data from a database rather than documents.
    """
    question = state["sub_questions"][state["current_step"]]

    try:
        result = sql_agent.invoke({"input": question})
        answer = result["output"]
    except Exception as e:
        # Graceful degradation — don't crash the whole pipeline on a bad query
        answer = f"Unable to retrieve structured data for this question: {str(e)}"

    return {
        **state,
        "retrieved_answers": [f"Q: {question}\nA (from database): {answer}"],
        "current_step": state["current_step"] + 1,
    }


# --- Route sub-questions to the right tool ---
class SubQuestionRoute(BaseModel):
    """Classify a sub-question to determine the right data source."""
    source: Literal["vectorstore", "database"] = Field(
        description="'database' for questions about numbers/metrics/records, 'vectorstore' for questions about documents/policies/text"
    )

sub_router = (
    ChatPromptTemplate.from_messages([
        ("system", "Classify where to find the answer: 'database' for quantitative data, 'vectorstore' for documents."),
        ("human", "{question}"),
    ])
    | ChatOpenAI(model="gpt-4o-mini").with_structured_output(SubQuestionRoute)
)

def route_sub_question(state: AgentState) -> str:
    step = state["current_step"]
    q = state["sub_questions"][step]
    result = sub_router.invoke({"question": q})
    return result.source  # "vectorstore" or "database"

# Add to workflow with routing
workflow.add_conditional_edges("decompose", lambda s: "route", {"route": "route_sub_q"})
workflow.add_node("route_sub_q", lambda s: s)  # Pass-through routing node
workflow.add_conditional_edges("route_sub_q", route_sub_question, {
    "vectorstore": "retrieve",
    "database": "query_database",
})

Pattern 3: Agent Handoff — The Librarian and the Strategist

For the most complex workflows, a single agent is insufficient. The pattern that's emerged in production 2026 systems is a supervisor-worker topology: a supervisor agent routes work to specialised worker agents, each of which is an expert in a narrow domain.

A useful concrete framing: the Librarian Agent is responsible for finding information — it knows your indices, your tools, your data sources. The Strategist Agent is responsible for reasoning over information once found — it synthesises, compares, and draws conclusions. The Librarian feeds context to the Strategist. Neither agent does both jobs well.

from langgraph.graph import StateGraph, END
from langchain_core.messages import HumanMessage, AIMessage
from typing import TypedDict, List

class SupervisorState(TypedDict):
    question: str
    retrieved_context: str    # Librarian's output
    final_answer: str         # Strategist's output
    handoff_complete: bool

# --- Librarian Agent: find and assemble context ---
def librarian_agent(state: SupervisorState) -> SupervisorState:
    """
    Specialised in retrieval. Uses hybrid search, SQL, and any available tools
    to assemble the richest possible context for the question.
    Does NOT attempt to answer — only gathers.
    """
    question = state["question"]

    # Multi-source context assembly
    doc_context = "\n\n".join(
        d.page_content for d in hybrid_retriever.invoke(question)
    )

    # Attempt structured data retrieval if question seems quantitative
    try:
        sql_result = sql_agent.invoke({"input": f"Find any relevant data for: {question}"})
        structured_context = sql_result["output"]
    except:
        structured_context = "No structured data available."

    assembled_context = (
        f"=== Document Context ===\n{doc_context}\n\n"
        f"=== Structured Data ===\n{structured_context}"
    )

    return {**state, "retrieved_context": assembled_context, "handoff_complete": True}


# --- Strategist Agent: reason over assembled context ---
def strategist_agent(state: SupervisorState) -> SupervisorState:
    """
    Specialised in reasoning and synthesis. Receives pre-assembled context
    from the Librarian. Produces the final answer.
    Does NOT retrieve — only reasons.
    """
    prompt = (
        f"You are a strategic analyst. Using ONLY the context provided, answer the question comprehensively.\n\n"
        f"Question: {state['question']}\n\n"
        f"Context:\n{state['retrieved_context']}\n\n"
        f"Provide a structured, evidence-based answer. Cite which context sections support each claim."
    )

    answer = ChatOpenAI(model="gpt-4o").invoke(prompt).content
    return {**state, "final_answer": answer}


# Build supervisor graph
supervisor = StateGraph(SupervisorState)
supervisor.add_node("librarian", librarian_agent)
supervisor.add_node("strategist", strategist_agent)

supervisor.set_entry_point("librarian")
supervisor.add_edge("librarian", "strategist")  # Handoff: Librarian → Strategist
supervisor.add_edge("strategist", END)

pipeline = supervisor.compile()
result = pipeline.invoke({
    "question": "Which product categories underperformed Q3 targets and why?",
    "retrieved_context": "",
    "final_answer": "",
    "handoff_complete": False,
})
print(result["final_answer"])

The separation of concerns here is the whole point. The Librarian is optimised for recall — it casts a wide net and doesn't worry about conciseness. The Strategist is optimised for reasoning — it has clean, assembled context and focuses entirely on producing a rigorous answer. Mixing these responsibilities produces agents that do both jobs poorly.


MCP Servers: The Universal Tool Interface

Beyond SQL, production agentic RAG systems in 2026 increasingly use Model Context Protocol (MCP) servers as a standardised way to expose tools — whether that's a CRM, a calendar API, a code execution environment, or a proprietary internal system. MCP decouples the agent from the tool implementation: you update the tool server without touching the agent code.

from langchain_mcp_adapters.client import MultiServerMCPClient
from langgraph.prebuilt import create_react_agent

# Connect to MCP servers — same tools work across any MCP-compatible framework
async def create_mcp_agent():
    async with MultiServerMCPClient(
        {
            "crm": {
                "url": "http://localhost:8001/mcp",   # CRM MCP server
                "transport": "streamable_http",
            },
            "analytics": {
                "url": "http://localhost:8002/mcp",   # Analytics MCP server
                "transport": "streamable_http",
            },
        }
    ) as client:
        tools = client.get_tools()   # All MCP tools as LangChain-compatible tool objects

        # Create a ReAct agent with the MCP tools
        agent = create_react_agent(
            model=ChatOpenAI(model="gpt-4o"),
            tools=tools,
        )

        result = await agent.ainvoke({
            "messages": [HumanMessage(content="What are the top 5 accounts by ARR that haven't been contacted in 30 days?")]
        })
        return result["messages"][-1].content

The power of MCP here is that the agent doesn't need to know how the CRM tool works internally — it just knows the tool exists and what it does. When the CRM migrates from Salesforce to HubSpot, you update the MCP server, not the agent. The agent code stays identical.


The Agentic Spectrum

Not every problem needs the full agentic treatment. Here's a practical guide to when to escalate:

Query ComplexityPatternLatency Cost
Single-source, single-hopStandard RAG (Part 1-2)Low
Requires routing or self-correctionAdaptive/CRAG (Part 3)Medium
Multi-source, multi-hopQuery decomposition (this post)Medium-High
Structured + unstructured dataSQL agent + RAG hybridHigh
Cross-system reasoningSupervisor + MCP toolsHighest

The temptation is to build the most sophisticated system immediately. Resist it. Start with the minimum pattern that handles your actual query distribution, measure with your golden dataset, and escalate only when the simpler system demonstrably fails.


Up Next: The Production Reality

We now have a complete agentic RAG stack — from clean ingestion to multi-agent reasoning. But none of it matters if you can't deploy it in a regulated environment, control who can access what, or prove to an auditor that the system behaved correctly.

In Part 5, we close the series with the unglamorous but essential work: governance, access control, EU AI Act compliance, and ROI frameworks. The stuff that keeps RAG systems running in production rather than being quietly decommissioned six months after launch.

Tag someone building their first agent — they're about to discover why the Librarian and the Strategist can't be the same node.

Subscribe to Vivek Wisdom

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe