DSPy: Stop Writing Prompts. Start Programming Them.

A practical, beginner-friendly guide to DSPy, the framework that makes your LLM apps easier to maintain. Build a working multi-agent system with LangGraph by the end.

DSPy: Stop Writing Prompts. Start Programming Them.
DSPy - Stop writing prompts, start programming them

A practical, beginner-friendly guide to the framework that makes your LLM apps easier to maintain. We will also build a small project using LangGraph at the end.

💡
TL;DR: DSPy lets you write LLM apps as Python code, not as fragile prompt strings. You declare what the model should do, and DSPy figures out how to ask. By the end of this post you will have built a working multi-agent system with LangGraph and DSPy in about 30 minutes.

The Problem Nobody Warns You About

You build an LLM app. It works. You feel great.

Then OpenAI ships a new model. Or you switch to Claude to cut costs. Or you bring in Llama for compliance reasons. Suddenly, your prompts behave differently. The polite "Please respond in JSON format with exactly these fields..." that worked on GPT-4 starts producing broken output on Claude. The few-shot examples you spent two weekends tuning are now overfit to the old model.

So you start tweaking again. Adding "IMPORTANT:" in caps. Adding more examples. Rewriting the system message. Twelve hours later, you have a prompt that works for this model, on this dataset, until the next thing changes.

This is the dirty secret of modern LLM work. Most "prompt engineering" is just messy string concatenation with extra stress.

DSPy exists to fix this.

What DSPy Actually Is (in plain English)

DSPy stands for Declarative Self-improving Python. It is an open-source framework from Stanford NLP. The idea is to treat prompts the way PyTorch treats neural networks. They become compiled artifacts you generate from higher-level code, not strings you write by hand.

The official tagline is "programming, not prompting" language models. That sounds buzzword-heavy, so let me translate.

In a normal LLM app, you write something like this:

prompt = f"""You are a helpful assistant. Given the following question, 
answer in 1-3 sentences. Be concise. Do not hallucinate.

Question: {question}
Answer:"""

In DSPy, you write this:

qa = dspy.ChainOfThought("question -> answer")
result = qa(question="What is the capital of France?")

That is it. No prompt template. You declared a signature (question -> answer) and picked a module (ChainOfThought, which means "think step by step before answering"). DSPy compiles that into an optimized prompt at runtime, fits it to whatever model you set up, and parses the output back into a typed Python object.

The really interesting part comes next. You can give DSPy a small dataset of examples and a metric, and it will automatically optimize the underlying prompt. It picks better instructions, chooses the most useful few-shot demonstrations, even rewrites the wording, all to maximize your metric. You do not write the optimized prompt. The compiler does.

🎯
Real-world impact: The official DSPy docs cite a case where GPT-4o-mini's accuracy on a labeling task jumped from 66 percent to 87 percent just by running an optimizer over the same code. No prompt rewriting needed.

People describe this shift as moving from assembly to C for LLM apps.

Why DSPy: The Honest Case

Skipping the marketing pitch, here is why this framework matters in practice.

1. Your code becomes portable across models. A signature like document -> summary is model-agnostic. Swap GPT-4o for Claude for a local Llama and your code does not change. DSPy regenerates the prompt for the new model. This alone is worth it if you are cost-conscious or worried about vendor lock-in.

2. Prompts become optimizable, not artisanal. DSPy ships with optimizers like BootstrapFewShot, MIPROv2, and the newer GEPA. Hand them training examples and a metric, like "did the answer match the gold label?", and they search the space of prompts and demonstrations for you. Quality gains in the 10 to 40 percent range over manual prompting are common.

3. Composition works the way it should. Multi-step pipelines like retrieval, then reasoning, then verification are just Python modules calling other Python modules. No fragile chains of string templates.

4. It plays nice with the rest of your stack. DSPy is not trying to replace LangGraph or LlamaIndex or your vector DB. It is the reasoning core. Use LangGraph for stateful orchestration, DSPy for the LM calls inside each node, your vector DB for retrieval. Everyone stays in their lane.

5. You can debug what actually happened. dspy.inspect_history() shows you the exact prompt that was sent and the exact response. No more guessing what your "chain" actually did.

Who Should Use DSPy

Be honest about where you are. DSPy is not the right tool for everyone.

You should use DSPy if:

  • You are building anything beyond a one-off ChatGPT-style demo. Classifiers, RAG systems, agents, multi-step pipelines.
  • You expect to switch or upgrade models at some point. Everyone does.
  • You have, or can collect, even 20 to 50 labeled examples of what good output looks like. That is enough to optimize.
  • You care about reproducibility. Being able to re-run the same logic six months from now and get sane behavior.
  • You are tired of prompts that are 600 words long and growing.
⚠️
Skip DSPy if: you are building a thin wrapper around a single chat call and "good enough" is fine, you have zero way to evaluate output quality (DSPy's superpower is optimization, which needs a metric), or you need to ship in two hours and have never seen the framework before.

The sweet spot is mid-complexity LLM apps where prompt quality actually matters and you will iterate on it. RAG systems are the classic example. Agent workflows are catching up fast.

The Mental Model: Three Concepts

If you internalize these three things, you can read any DSPy codebase.

1. Signatures: what the LM should do

A signature is a typed declaration of inputs and outputs. The shorthand form is a string:

"question -> answer"
"document -> summary"
"context, question -> answer"
"sentence -> sentiment: bool"

The class-based form lets you add descriptions and type hints:

class Classify(dspy.Signature):
    """Classify the sentiment of a movie review."""
    review: str = dspy.InputField()
    sentiment: Literal["positive", "negative", "neutral"] = dspy.OutputField()
    confidence: float = dspy.OutputField(desc="0.0 to 1.0")

Notice you are describing what, not how. No "please" and no "be careful." That is the compiler's job.

2. Modules: how the LM should do it

Modules are the strategies. Same signature, different module gives different reasoning behavior.

  • dspy.Predict: just answer.
  • dspy.ChainOfThought: think step by step before answering. Adds a reasoning field automatically.
  • dspy.ProgramOfThought: generate code, run it, use the result.
  • dspy.ReAct: an agent that can call tools.
  • dspy.Refine and dspy.BestOfN: run multiple times and pick the best. These replace the older Assert and Suggest API in DSPy 2.6+.

Swap modules freely. The signature stays the same.

3. Optimizers: making it better automatically

This is the magic part. You give the optimizer three things:

  • Your program, which is a module or a composition of modules.
  • A small training set of dspy.Example objects.
  • A metric function that returns a number. Higher is better.

It searches over instructions and few-shot demonstrations to maximize the metric. The common ones to know:

  • BootstrapFewShot: selects effective few-shot examples from your training data. Cheap and fast. Start here.
  • MIPROv2: jointly optimizes instructions and demonstrations. More expensive, often noticeably better.
  • GEPA: newer, uses reflective feedback loops. State of the art for some tasks as of 2026.

You compile once, save the optimized program, deploy. That is the workflow.

How DSPy Fits Into a Modern LLM Stack

Here is the honest layout of a production-ish LLM app in 2026:

LayerJobExamples
Model providerThe actual LMOpenAI, Anthropic, Google, Ollama, vLLM
Reasoning coreDefine and optimize how the LM thinksDSPy
OrchestrationState, branching, multi-step flowsLangGraph, LlamaIndex Workflows
RetrievalVector search, hybrid searchWeaviate, Pinecone, ColBERT, pgvector
ObservabilityTracing, evals, debuggingLangfuse, LangSmith, Arize
🔧
The mental shortcut: LangGraph orchestrates; DSPy reasons. LangGraph handles "after the classifier node, branch to either the math agent or the writer agent, with retry logic and memory." DSPy handles "what does the classifier node actually do internally, and how do we make it better over time?"

This is the combo we will build in the project below.

The Project: A Self-Improving Question Router with LangGraph and DSPy

Time to get hands dirty. We will build something small but real.

DSPy
The framework for programming, not prompting, language models. Official docs from Stanford NLP.

A multi-agent question-answering app that:

  1. Takes a user question.
  2. Uses a DSPy classifier to route it to one of three specialist agents: factual, creative, math.
  3. Uses LangGraph for routing and state management.
  4. Has each specialist agent powered by a DSPy module suited to its job.
  5. Can be optimized with a handful of labeled examples to improve routing accuracy.

By the end you will have something runnable, plus a feel for how the pieces fit together.

Prerequisites

  • Python 3.10 or higher
  • An API key for any supported LM (OpenAI, Anthropic, Google, or local via Ollama)
  • About 30 minutes

Step 1: Install

pip install dspy langgraph

Step 2: Configure your LM

DSPy uses LiteLLM under the hood, so the same syntax works for every provider:

import dspy
import os

# Pick one based on what you have a key for
lm = dspy.LM("openai/gpt-4o-mini", api_key=os.environ["OPENAI_API_KEY"])
# lm = dspy.LM("anthropic/claude-sonnet-4-5", api_key=os.environ["ANTHROPIC_API_KEY"])
# lm = dspy.LM("ollama/llama3.1")  # local

dspy.configure(lm=lm)

Step 3: Define your DSPy modules

from typing import Literal

class QuestionClassifier(dspy.Signature):
    """Classify a user question into one of three categories."""
    question: str = dspy.InputField()
    category: Literal["factual", "creative", "math"] = dspy.OutputField(
        desc="factual = needs a real-world fact; creative = needs imagination/writing; math = needs calculation"
    )

class FactualAnswer(dspy.Signature):
    """Answer a factual question concisely and accurately."""
    question: str = dspy.InputField()
    answer: str = dspy.OutputField(desc="1-3 sentences, factual, no fluff")

class CreativeAnswer(dspy.Signature):
    """Respond to a creative prompt with imagination and flair."""
    question: str = dspy.InputField()
    answer: str = dspy.OutputField(desc="Vivid, original, 2-4 sentences")

class MathAnswer(dspy.Signature):
    """Solve a math problem, showing your work."""
    question: str = dspy.InputField()
    answer: str = dspy.OutputField(desc="The final numeric answer, plus brief reasoning")

# Wire them up with appropriate modules
classifier = dspy.ChainOfThought(QuestionClassifier)
factual_agent = dspy.Predict(FactualAnswer)
creative_agent = dspy.Predict(CreativeAnswer)
math_agent = dspy.ChainOfThought(MathAnswer)  # CoT helps with math

Notice that each agent uses the module that fits its job. The math agent gets ChainOfThought because step-by-step helps. The factual agent uses plain Predict because we want a tight answer.

Step 4: Build the LangGraph

from langgraph.graph import StateGraph, END
from typing import TypedDict

class State(TypedDict):
    question: str
    category: str
    answer: str

def classify_node(state: State) -> State:
    result = classifier(question=state["question"])
    return {**state, "category": result.category}

def factual_node(state: State) -> State:
    result = factual_agent(question=state["question"])
    return {**state, "answer": result.answer}

def creative_node(state: State) -> State:
    result = creative_agent(question=state["question"])
    return {**state, "answer": result.answer}

def math_node(state: State) -> State:
    result = math_agent(question=state["question"])
    return {**state, "answer": result.answer}

def route(state: State) -> str:
    return state["category"]

# Assemble
graph = StateGraph(State)
graph.add_node("classify", classify_node)
graph.add_node("factual", factual_node)
graph.add_node("creative", creative_node)
graph.add_node("math", math_node)

graph.set_entry_point("classify")
graph.add_conditional_edges("classify", route, {
    "factual": "factual",
    "creative": "creative",
    "math": "math",
})
graph.add_edge("factual", END)
graph.add_edge("creative", END)
graph.add_edge("math", END)

app = graph.compile()

Step 5: Run it

result = app.invoke({"question": "What is 17 times 24?"})
print(f"Category: {result['category']}")
print(f"Answer: {result['answer']}")

result = app.invoke({"question": "Write a haiku about debugging."})
print(f"Category: {result['category']}")
print(f"Answer: {result['answer']}")

You now have a working multi-agent system in roughly 80 lines of code.

Step 6: Optimize the classifier

The raw classifier works, but it will misroute edge cases. Let us give it 8 labeled examples and let DSPy improve it:

trainset = [
    dspy.Example(question="What year did WW2 end?", category="factual").with_inputs("question"),
    dspy.Example(question="Who painted the Mona Lisa?", category="factual").with_inputs("question"),
    dspy.Example(question="Write a poem about rain.", category="creative").with_inputs("question"),
    dspy.Example(question="Invent a name for a coffee shop.", category="creative").with_inputs("question"),
    dspy.Example(question="What is 144 divided by 12?", category="math").with_inputs("question"),
    dspy.Example(question="If a train leaves at 3pm going 60mph...", category="math").with_inputs("question"),
    dspy.Example(question="Tell me a short story about a robot.", category="creative").with_inputs("question"),
    dspy.Example(question="What is the boiling point of water?", category="factual").with_inputs("question"),
]

def accuracy(example, pred, trace=None):
    return example.category == pred.category

from dspy.teleprompt import BootstrapFewShot
optimizer = BootstrapFewShot(metric=accuracy, max_bootstrapped_demos=4)
optimized_classifier = optimizer.compile(classifier, trainset=trainset)

# Save it
optimized_classifier.save("classifier_v1.json")
🚀
This is where DSPy earns its keep. Now swap classifier for optimized_classifier in your classify_node function. The classifier has been improved without you writing a single line of new prompt text. Try the same trick on every node in your graph and watch your end-to-end accuracy climb.

Step 7 (bonus): Inspect what DSPy actually did

dspy.inspect_history(n=1)

You will see the actual prompt that got sent to the LM. Instructions, few-shot demonstrations, the works. This is where you stop treating the framework as a black box.

Tips for Beginners

If I could go back to the day I started with DSPy, here is what I would tell myself.

1. Start with dspy.Predict. Do not reach for ReAct or ProgramOfThought on day one. Get a feel for signatures first. Add complexity only when you measure that you need it.

2. Write your metric before you write your optimizer call. A bad metric is worse than no metric. The optimizer will dutifully maximize the wrong thing. If you cannot write a metric, you cannot optimize. That is a feature, not a bug.

3. Twenty examples is enough to start. You do not need 10,000 labeled samples. DSPy is designed for small data. If you can get 20 good examples and 20 dev examples, you can run a real optimization.

4. Use inspect_history() constantly. It is the print-debugging of DSPy. When something behaves weird, look at the actual prompt that got generated. Nine times out of ten the bug is obvious once you see it.

5. Save optimized programs to disk. Re-running expensive optimizers every time you start your app is painful. Compile once, save the JSON, load it in production.

6. DSPy plus LangGraph means DSPy lives inside graph nodes. Do not try to make DSPy do the orchestration. Let LangGraph handle state, branching, and retries. Let DSPy handle what the LM should say at each step.

7. Observability matters early. Plug in Langfuse or LangSmith before your pipeline gets complicated. Tracing a 6-step DSPy plus LangGraph flow without observability is a special kind of pain.

8. Read the cheatsheet, not just the tutorials. The DSPy cheatsheet on the official docs is the single highest-value resource. Bookmark it.

Useful Resources

These are the links worth keeping open in a tab while you learn.

DSPy Cheatsheet
The single highest-density reference for DSPy. Every common pattern in one page: signatures, modules, optimizers, streaming, async, caching, and judges.
stanfordnlp/dspy on GitHub
The official DSPy repository. Issues are answered fast and the examples folder is gold.
LangGraph Documentation
Official docs for LangGraph, the stateful orchestration framework we used in this project.

Where to Go From Here

You now have:

  • A mental model for what DSPy is and why it exists.
  • A clear sense of whether it fits your use case.
  • A working multi-agent app combining LangGraph and DSPy.
  • An optimized classifier you trained yourself.

Some honest next steps, in order of difficulty:

  1. Extend the project. Add a fourth agent type. Add retrieval to the factual agent and turn it into a tiny RAG. Add a verification step that checks the answer before returning.
  2. Try a different optimizer. Swap BootstrapFewShot for MIPROv2. Compare the prompts each one produces.
  3. Swap models. Change one line, your dspy.LM(...), and run the whole thing on Claude or a local model. Notice that nothing else breaks.
  4. Read the GEPA paper and the MIPROv2 docs. Once you have built one app, the research papers actually make sense and they are surprisingly readable.
⭐
The shift in how you think about LLM apps takes about a week to fully click. You will write a few signatures, fight the framework once or twice, and then one afternoon you will find yourself reaching for a signature instead of a prompt template without thinking about it. That is when you have crossed over. Welcome to programming, not prompting.

Subscribe to Vivek Wisdom

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe