Evals: The Single Highest-Leverage Discipline in AI Engineering

Part 4 of the AI Engineer Series. Evals are the single highest-leverage discipline in AI engineering. Deterministic assertions, LLM-as-judge with calibration, pairwise versus pointwise scoring, and the discipline that separates teams who ship from teams who hope.

Without evals, you are guessing

Here is the test. Pick any LLM feature you ship today. Now answer this: did last week's prompt change make it better or worse? If you cannot answer with a number, you do not have evals, you have vibes. Vibes feel productive for two weeks, then collapse when you cannot tell whether your changes helped or hurt.

This is Part 4 of the AI Engineer Series, and it is the single highest-leverage discipline in everything we cover. Without measurement, harness engineering is theater, retrieval tuning is guesswork, and routing decisions are gut feel. With measurement, every other optimization becomes obvious.

The eval hierarchy

Not all evals are equal. Use them in order of cost and confidence.

Deterministic assertions come first. Did the output contain the expected substring? Is it valid JSON? Does it pass a regex? Does it cite at least one source? These are cheap, fast, and unambiguous. If you can express a quality check as a deterministic assertion, do it. They run in milliseconds and never have judge bias.

Reference-based metrics come next. You have a ground-truth answer and you compare against it. BLEU, ROUGE, exact match, semantic similarity. Useful for translation, summarization, extraction tasks where the right answer is known.

LLM-as-judge handles everything else. You give a stronger model a rubric and the output, and it grades. Cheap, scalable, and biased. We will get to mitigating those biases shortly.

Human evaluation is the gold standard. Slow, expensive, irreplaceable. You need it to calibrate your LLM judges, to catch failure modes humans notice but machines do not, and to validate your eval set itself.

Build a golden dataset, then freeze it

You need a stable set of examples to measure against. The rules:

  • 50 minimum, 200 ideal to start. You can always grow it.
  • Stratify the examples. Cover happy paths, common failure modes, adversarial cases, and historical bugs. A dataset of only easy cases lies to you.
  • Freeze it. Once your eval set is in use, do not let those examples leak into your training data or your few-shot examples. The cardinal sin of LLM development is eval contamination.
  • Version it. When you change the eval set (and you will, as you learn), bump the version. Scores across versions are not comparable.

The simplest version is a JSON file in git:

[
    {
        "id": "ext-001",
        "input": "Invoice #4421 dated March 3, total $1,247.50 for office supplies",
        "expected": {
            "invoice_number": "4421",
            "total": 1247.50,
            "category": "office_supplies"
        },
        "tags": ["happy_path", "extraction"]
    },
    {
        "id": "ext-002",
        "input": "Bill from last Tuesday, around $300, services rendered",
        "expected": {
            "invoice_number": None,
            "total": 300.00,
            "category": "services"
        },
        "tags": ["fuzzy_date", "incomplete_data"]
    }
]

LLM-as-judge, calibrated properly

LLM judges work, but only when calibrated. The common mistakes that ruin judge pipelines:

Vague rubrics. "Is the answer good?" returns noise. "Does the answer cite at least one source from the provided context AND avoid mentioning information not in the context?" returns signal. Write rubrics like you are writing assertions.

No human calibration. Run your judge on 30-50 examples that humans have also labeled. Compute agreement. If your judge and your humans agree less than 80% of the time, your rubric is broken. Fix the rubric, not the threshold.

Position bias. In pairwise comparisons, models systematically prefer the first option (or, on some models, the last). The 2024 LLM-as-judge paper showed GPT-4 being inconsistent across position swaps 40% of the time. The fix is to evaluate both orderings and only count consistent wins.

Verbosity bias. Longer answers seem better to judges. The inflation effect is around 15% in unmitigated pipelines. Mention conciseness explicitly in the rubric.

Self-preference. A model judging its own outputs is suspect. Use a different model family as judge. Claude judging Claude outputs is biased. Claude judging GPT outputs is more honest.

Here is what a calibrated judge looks like in code:

from pydantic import BaseModel, Field
from typing import Literal
import random

class JudgeResult(BaseModel):
    winner: Literal["A", "B", "tie"]
    reasoning: str
    confidence: Literal["high", "medium", "low"]

RUBRIC = """You are evaluating two answers to the same question.

Criteria, in order of importance:
1. Factual accuracy. Does the answer contain only claims supported by the provided context?
2. Completeness. Does it address all parts of the question?
3. Conciseness. Shorter answers are preferred when they convey the same information.

Do NOT prefer longer answers. Do NOT prefer answers that sound more confident.
Penalize answers that include information not in the context.

Output the winner (A, B, or tie), your reasoning, and your confidence."""

def judge_pairwise(question, context, answer_a, answer_b):
    # Randomize position to avoid position bias
    if random.random() < 0.5:
        first, second, swap = answer_a, answer_b, False
    else:
        first, second, swap = answer_b, answer_a, True
    
    result = client.messages.create(
        model="claude-opus-4-7",
        max_tokens=500,
        response_model=JudgeResult,
        messages=[{
            "role": "user",
            "content": f"{RUBRIC}\n\nQuestion: {question}\nContext: {context}\n\nA: {first}\n\nB: {second}"
        }],
    )
    
    # Unswap if needed
    if swap and result.winner != "tie":
        result.winner = "B" if result.winner == "A" else "A"
    return result

Pairwise beats pointwise

Asking a judge to rate an output 1-5 sounds easier than asking which of two outputs is better. It is not. Pointwise scores drift across runs because the model anchors them differently each time. Pairwise comparisons are more stable. Humans have the same issue, by the way.

Use pointwise for absolute quality gates (does this pass a fixed bar?). Use pairwise for ranking changes against a baseline (is the new prompt better than the old one?). The pairwise comparison is what tells you whether your work moved the needle.

The three eval loops

Evals run in three different places, each with different cost and latency budgets:

  • Dev loop. Subset of evals running locally on every prompt change, in under 30 seconds. If it takes longer than that, you will stop running it.
  • CI loop. Full eval suite on every PR. Block merges on regression. This is where the team-level discipline happens.
  • Production loop. Sample real traffic, run online evals against real outputs, track drift over time. The online judge catches the failure modes your offline eval set does not cover.

RAGAS is the standard library for RAG-specific evals. Promptfoo, Braintrust, and Langfuse all package the broader eval workflow. Pick one and use it. The tool matters less than the discipline.

The Hamel rule

Look at the data. Personally. For at least 50 examples. Every dashboard and metric in the world is no substitute for skimming actual model outputs with your own eyes. You will see failure modes that no automated metric catches, and you will catch your judge being wrong about cases your judge thinks are right. This single discipline separates teams who actually improve from teams who just collect numbers.

What to build this week

Build an eval suite for the structured-output task you set up in Part 3:

  • Create a 50-example golden dataset in JSON, in git, with stratified tags
  • Implement three checks: a deterministic assertion (schema validity), a reference-based metric (field-level accuracy), and an LLM-as-judge for the open-ended fields
  • Run your judge on 30 examples you have also labeled by hand. Measure agreement. If below 80%, refine the rubric and re-measure.
  • Run the full suite. Note the baseline numbers.
  • Make a deliberate prompt change you suspect helps. Re-run. Did the numbers move in the direction you expected?

You will find that the deterministic assertions catch most regressions, the reference metrics catch the structural ones, and the judge handles the rest. That layered defense is what makes evals work.

What is next

Part 5: LLM observability. Evals tell you whether changes are good. Observability tells you what is actually happening in production. We will cover trace structure, the OpenTelemetry GenAI conventions, sampling strategies, and how to catch regressions before users do.

Previous in the series

References

  1. AI Evals FAQ by Hamel Husain and Shreya Shankar. The most practical writing on the actual mechanics of building evals.
  2. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena by Zheng et al. The paper that documented position bias, verbosity bias, and self-preference bias.
  3. Who Validates the Validators? by Shankar et al. On the challenge of calibrating LLM judges against human preferences.
  4. LLM-Evaluators as Judges by Eugene Yan. Excellent overview of patterns and pitfalls.
  5. RAGAS documentation. The standard eval framework for RAG pipelines.
  6. Braintrust documentation. A mature eval platform that handles the dev, CI, and production loops in one place.

Subscribe to Vivek Wisdom

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe