Harness Engineering: The Code Around Your LLM Calls That Keeps Production Alive
Part 1 of the AI Engineer Series. Prompt engineering gets the spotlight, but in production, most reliability comes from the system around the prompt. This is the harness layer, and most teams underbuild it.
The thing nobody tells you about LLM apps
You can build a working LLM demo in an afternoon. The API call is three lines of code. The prompt fits in a paragraph. The output looks great.
Then you put it in front of users and everything that can go wrong, goes wrong. The API times out. The model returns markdown when you asked for JSON. A retry storms the provider during an outage. A prompt change in your codebase silently breaks output quality for a week before anyone notices.
None of these problems are about the prompt. They are about the code around the prompt. That code is what we call the harness, and it is the difference between a demo and a system.
This is Part 1 of the AI Engineer Series. Before we talk about RAG, evals, fine-tuning, or any of the flashier topics, we have to talk about the boring layer that holds everything else up.
What a harness actually is
Treat every LLM call like an RPC to a non-deterministic service that can fail in creative ways. Because that is exactly what it is.
A harness is the wrapper code that turns that flaky probabilistic API into something your application can depend on. It handles four things:
- The contract between your code and the model
- What happens when the call fails
- How you version and ship prompt changes
- How you reproduce something that went wrong
Let us go through each one.
The contract layer
Most LLM bugs in production are not "the model gave a bad answer." They are "the model returned something that broke our parser, our typing, or our downstream call."
Define explicit contracts using Pydantic. Validate both inputs and outputs. The validation on inputs surprises people, but malformed inputs are a top cause of garbage outputs. A user pasting 500K tokens into a 32K-context field is a problem you want to catch before you spend money on the API call.
from pydantic import BaseModel, Field
from typing import Literal
class SummarizeRequest(BaseModel):
text: str = Field(min_length=10, max_length=50_000)
style: Literal["bullet", "prose", "headline"]
max_words: int = Field(gt=0, le=500)
class SummarizeResponse(BaseModel):
summary: str
confidence: float = Field(ge=0, le=1)
used_model: strNow your harness has type-checked rails on both sides. The function signature documents what the system promises. The model output is validated before it touches anything else. Pydantic v2 is the de facto standard here, and most modern LLM libraries including instructor are built on top of it.
Retries, timeouts, idempotency
LLM providers fail. Rate limits trip. Servers return 500s. Content filters block legitimate output. Streams hang. Your harness needs a plan for all of it.
The non-obvious rules:
Use exponential backoff with jitter. Never retry without jitter. During a provider outage, ten thousand clients retrying on the same backoff schedule create a thundering herd that prolongs the outage. AWS published the canonical write-up on this in 2015 and it still applies.
Set distinct timeouts. Connect timeout, time-to-first-token timeout, and total timeout are three different problems. A hung connection at 30 seconds is not the same problem as a slow stream taking 90 seconds. Treating them as one number means you fail badly on both.
Use idempotency keys where the provider supports them. Both Anthropic and OpenAI support an Idempotency-Key header. If your retry succeeds where the original silently did too, you do not want to double-charge or double-trigger downstream effects.
Cap retries globally. Per-request retry budgets are not enough. You also need a per-minute global cap. When a provider is degraded, you do not want every request in your queue triple-retrying at once.
The single rule that catches the most bugs: do not retry on 4xx errors. Those are your bugs. The input was malformed, the auth was wrong, the model name was misspelled. Retrying just wastes money and obscures the real problem. Only retry 429, 500, 502, 503, 504, and connection errors.
from tenacity import retry, stop_after_attempt, wait_exponential_jitter
from anthropic import Anthropic, APIStatusError, APITimeoutError, APIConnectionError
client = Anthropic()
RETRYABLE_STATUSES = {429, 500, 502, 503, 504}
def is_retryable(exc):
if isinstance(exc, APIStatusError):
return exc.status_code in RETRYABLE_STATUSES
return isinstance(exc, (APITimeoutError, APIConnectionError))
@retry(
stop=stop_after_attempt(4),
wait=wait_exponential_jitter(initial=1, max=30),
retry=is_retryable,
)
def call_llm(prompt: str, idempotency_key: str) -> str:
msg = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
messages=[{"role": "user", "content": prompt}],
timeout=30.0,
extra_headers={"Idempotency-Key": idempotency_key},
)
return msg.content[0].textThe tenacity library handles the retry mechanics well. Both the official Anthropic and OpenAI Python SDKs already have built-in retries with exponential backoff, so check the SDK docs before adding your own layer on top.
Prompts as code, not strings
Inline string literals for prompts are how teams accidentally ship quality regressions. Someone changes a system prompt to fix one edge case, breaks three others, and nobody notices for a week because there is no diff to review and no test to fail.
Prompts are configuration. Treat them like code:
- Store them in files, version them in git
- Use a templating system (Jinja2 is fine)
- Tag versions and pin them in deployments
- Review prompt changes in PRs like any other code
This unlocks two things. First, every prompt change becomes a diff that someone reviews. Second, you can roll back. A bad prompt in production is now just a revert away.
from pathlib import Path
import jinja2
class PromptLoader:
def __init__(self, prompts_dir: str):
self.env = jinja2.Environment(
loader=jinja2.FileSystemLoader(prompts_dir),
undefined=jinja2.StrictUndefined,
autoescape=False,
)
def render(self, name: str, version: str, **kwargs) -> str:
template = self.env.get_template(f"{name}/v{version}.j2")
return template.render(**kwargs)
prompts = PromptLoader("./prompts")
text = prompts.render("summarize", "3", style="bullet", doc="...")Strict undefined is the small detail that matters. If your template references a variable that was not passed in, you want it to fail loudly, not render an empty string and quietly degrade the output.
For teams that have outgrown raw Jinja, look at libraries like dspy or commercial prompt management in Langfuse and Braintrust. The principle is the same: separate the prompt from the code that uses it.
Deterministic replay
When something goes wrong in production, you need to reproduce it. Stack traces tell you almost nothing about LLM systems. The bug is in the interaction between the prompt, the variables, the model state, and a sampling step. You cannot debug it without the inputs.
Log every call. The full input. The prompt template name and version. The variables that filled it. The model name, temperature, top-p, max tokens, and any other sampling parameters. The full output. The latency. The cost.
Then build a replay tool. Given a request ID, it should reconstruct the exact original call and let you re-run it, optionally against a different prompt version. This is how you debug LLM systems. Not by reading logs, but by replaying interactions.
import sqlite3
import json
import uuid
from datetime import datetime, timezone
class CallLog:
def __init__(self, db_path: str = "harness.db"):
self.db = sqlite3.connect(db_path)
self.db.execute("""
CREATE TABLE IF NOT EXISTS calls (
id TEXT PRIMARY KEY,
ts TEXT,
prompt_name TEXT,
prompt_version TEXT,
variables TEXT,
model TEXT,
params TEXT,
output TEXT,
latency_ms INTEGER,
cost_usd REAL
)
""")
def record(self, prompt_name, prompt_version, variables,
model, params, output, latency_ms, cost_usd):
call_id = str(uuid.uuid4())
self.db.execute(
"INSERT INTO calls VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(call_id, datetime.now(timezone.utc).isoformat(),
prompt_name, prompt_version, json.dumps(variables),
model, json.dumps(params), output, latency_ms, cost_usd)
)
self.db.commit()
return call_idThe local SQLite version is fine for development. In production, you want this in a real warehouse with proper retention, or use a hosted tool like Langfuse or Phoenix that does this for you with a real UI on top.
Putting it together
A real harness call looks something like this:
def summarize(req: SummarizeRequest) -> SummarizeResponse:
idempotency_key = f"summarize-{hash(req.text)}-{req.style}"
prompt = prompts.render("summarize", "3", **req.model_dump())
raw = call_llm(prompt, idempotency_key)
result = SummarizeResponse.model_validate_json(raw)
log.record(
prompt_name="summarize",
prompt_version="3",
variables=req.model_dump(),
model="claude-sonnet-4-6",
params={"max_tokens": 1024, "temperature": 0.2},
output=raw,
latency_ms=latency,
cost_usd=cost,
)
return resultFive layers, each doing one job. The contracts catch malformed data. The retry logic handles transient failures. The prompt loader gives you versioning. The logger gives you replay. And the validation on output gives you confidence that what you return downstream is shaped the way your code expects.
None of this is glamorous. None of it shows up in a demo. But every team that ships LLM features into production ends up rebuilding some version of it, and the ones who build it first ship faster than the ones who learn the hard way.
What to build this week
If you take one thing from this post, build a small llm_harness library and use it as your default LLM interface for everything else in this series. It needs:
- A
call_llm()wrapper with retries, timeouts, and idempotency keys - Pydantic models for inputs and outputs
- A Jinja-based prompt loader reading from git-versioned files
- OpenTelemetry tracing for the LLM call (we will use this in Part 5)
- A replay log in SQLite
Two hundred lines of code. The compound benefit over the next ten posts in this series is enormous. Every concept we cover after this assumes you have a place to put it.
What is next
Next in the AI Engineer Series: context engineering and retrieval quality. Most "bad LLM output" is actually "bad context in the prompt." We will look at chunking strategies, hybrid search, reranking, and why retrieval is the part of the system that determines almost everything else.
References
- Your AI Product Needs Evals by Hamel Husain. The piece that shifted the broader conversation toward treating LLM systems like real software.
- Patterns for Building LLM-based Systems and Products by Eugene Yan. A reference catalogue of the architectural patterns this series builds on.
- Building Effective Agents from Anthropic. The clearest writing on when to add complexity to LLM systems and when not to.
- Timeouts, retries, and backoff with jitter from the AWS Builders' Library. The canonical reference on retry safety patterns.
- Pydantic documentation. The validation library that does most of the heavy lifting in the contract layer.
- Tenacity documentation. The Python retry library used in this post.