The Operator's Library · No. 08
For operators who refuse to ship by vibe
Evaluating
AI Systems.

The unsexy, career-defining volume. How to test, score, and not deploy slop. Golden sets, eval types, regression on model swaps — and the discipline that turns AI work from "it seemed to work" into something you can stand behind.

Evaluating AI SystemsContents & Introduction
Contents & Introduction

"It worked when I tried it" is not evidence.


Most AI projects that fail in production passed every test the team ran — three vibe-checks on the happy path. Evaluation is the discipline that turns "seems good" into "we know." This is the volume that separates a Library reader from a Library operator.

Contents.

01Why evalsThe cost of not measuring. The price of measuring badly.p. 03
02Golden sets & the eval loopThe corpus, the runner, the scorer, the diff.p. 04
03Three eval typesExact match, rubric, LLM-as-judge. When each one fits.p. 05
04Regression on model swapsThe cheap habit that catches the expensive bug.p. 06
§When to retire an eval & colophonp. 07

Three principles.

  1. If you can't measure it, you can't ship it. Evals are the floor that lets you change the system confidently. Without them, every edit is a roll of the dice.
  2. Three cases is the minimum. Happy, edge, hostile. Below three, you don't have an eval — you have a vibe.
  3. Evals are code. Versioned, reviewed, runnable in CI. If yours live in a Notion page, they don't really live.

Whoever on your team owns the evals becomes the AI lead within six months. Always. It's the role that compounds.

02Contents
Ch. 01 · Why evalsEvaluating AI Systems
01
Chapter One

Why evals. The cost of not measuring.


Software has tests. ML has evals. The two are kin: a small set of known inputs with expected outputs, run automatically, on every change. The cost of skipping them is silent regression. The cost of doing them badly is false confidence. Both are addressable.

What you'd otherwise rely on.

  • "It worked last week." Different inputs. Different mood. No.
  • "The model didn't change." Your prompt did. Your context did. Your retrieval did.
  • "Users haven't complained." They will. The first complaint is the tip of weeks of quiet drift.
  • "I'll spot-check." You will, for two weeks, then you'll stop. The fix has to be automated.

Three concrete costs of no evals.

  1. You can't change the prompt. Any edit is a gamble; you back out everything because the cost of regression is unbounded.
  2. You can't swap models. Sonnet → Opus → next-gen. Without evals, every swap is a re-launch.
  3. You can't onboard the next person. They edit; nothing breaks visibly; quality drifts. Months later you realise.

What a healthy eval gets you.

  • Confidence to iterate the prompt daily.
  • The ability to swap models on a Tuesday, not a launch.
  • An onboarding artifact for new teammates ("read the evals to understand the system").
  • A regression suite to run before every release.
  • A conversation with your CFO about reliability that includes numbers.
The honest reframe

An eval suite is not bureaucracy. It is the asset that lets a small team move faster than a big one. Skipping evals is buying short-term speed with long-term paralysis.

An eval suite is your AI system's autonomic nervous system. Without it, every change is a risk; with it, every change is a measurement.

Today

For the highest-stakes prompt in your stack, write three test cases. Happy, edge, hostile. Save them next to the prompt. You now have an eval. The discipline is in the loop, not the perfection.

03Chapter 01
Ch. 02 · Golden sets & the eval loopEvaluating AI Systems
02
Chapter Two

Golden sets & the eval loop.


An eval is four small pieces: a golden set of inputs, expected outputs, a runner that executes the prompt against each, and a scorer that produces a pass/fail. Get the four pieces in place and you have a system. The pieces are smaller than you think.

The eval loop
GOLDEN SET RUNNER OUTPUT SCORER DIFF

Building the golden set.

Start with 10 cases. Five happy path, three edge, two hostile. Not 100; not 5. Ten is enough to catch most regressions and small enough to actually maintain.

Choose hostile cases on purpose. Missing fields. Ambiguous inputs. Inputs that look like prompt injection. The hostile cases are where most production bugs hide.

Pin expected outputs lightly. Don't pin the exact wording (it'll change with every model update). Pin the structural facts — "tone == friendly," "amount == 240," "action == draft."

Source from real history. Best golden cases come from things that actually happened (and broke). Each bug → one new eval case.

The runner.

A script that loads each case, fires the prompt with the case's input, captures the output. Should be runnable on demand and in CI. Doesn't need to be fancy — 50 lines is plenty.

The scorer.

This is where teams over-engineer. Don't. Match the scorer to the task:

  • Classification: exact match.
  • Extraction: field-by-field equality.
  • Drafting: rubric or LLM-as-judge (Ch. 03).
  • RAG: retrieval hit rate + answer faithfulness.

The diff.

The most important moment in the loop. When a case fails, you should be able to see, side-by-side: input, expected, actual, last-known-passing. A diff that takes 30 seconds to read is the difference between fixing the issue and giving up.

Watch out

Don't grow the golden set faster than you maintain it. 100 cases that no one reads is worse than 10 that everyone trusts. Quality over volume.

This week

Set up the four pieces for one prompt. Even at toy scale (10 cases, JSON expected output, equality scorer, console diff), the discipline is real. Add cases as bugs surface; not before.

04Chapter 02
Ch. 03 · Three eval typesEvaluating AI Systems
03
Chapter Three

Three eval types. Pick by task.


Most teams pick the wrong eval type for the task and then complain that evals are flaky. The trick is matching the type to the work. Three live patterns, each with a clear sweet spot.

TYPE 01 · EXACT MATCH

Output equals expected.

Output is structured (JSON, label, number, code). Compare with equality (or near-equality, with floating-point tolerance). Cheap. Deterministic. Doesn't lie.

Use for: classification, extraction, structured generation. The default when applicable.

TYPE 02 · RUBRIC

Hand-scored against criteria.

You define 3–6 binary criteria ("includes a CTA," "mentions price," "under 200 words"). Each case is scored against the rubric. A human or a script tallies.

Use for: drafting, summarising — anything where the output is text but quality is decomposable into checkable parts.

TYPE 03 · LLM-AS-JUDGE

Another model scores.

A separate prompt asks a model to compare actual vs. expected against the rubric and return a score. Cheap at scale; introduces its own bias; useful with care.

Use for: subjective quality at volume — but calibrate against human scores periodically, and never use the same model as judge and judged for sensitive tasks.

TYPE 04 · RETRIEVAL EVAL

Did we get the right chunks?

For RAG systems, the question before "is the answer good" is "did we retrieve the right material." Score against a labelled set of (query → expected chunk IDs).

Use for: any RAG system. Retrieval evals catch problems faster than end-to-end evals.

— DON'T DO —

Score "tone" with LLM-as-judge alone.

The judge model's idea of "tone" is its own. Calibrate against human scores on the same cases. If correlation is low, the eval isn't measuring what you think.

Mix: rubric checks + sampled human review. Cheap; correct.

— RULE OF THUMB —

Cheapest eval that catches the bug.

If exact-match works, don't reach for rubric. If rubric works, don't reach for a judge. Spend the complexity budget on cases, not scorers.

Default: exact-match for structure; rubric for prose; judge only when both fail.

Eval type is a design decision, not a default. Pick it consciously per task; mix where useful; document the choice.

05Eval types
Ch. 04 · Regression on model swapsEvaluating AI Systems
04
Chapter Four

Regression on model swaps.


A new model release is a software dependency upgrade. Treated like one, it's fine. Treated like "obviously better," it's a Friday-evening incident. Evals are the cheap habit that turns model migrations into a measurement, not a leap of faith.

The discipline, in five steps.

  1. Pin your current model in the prompt/config. Treat the model name as part of the system.
  2. Run the eval against the pinned model on every PR or weekly. Establish the baseline pass rate.
  3. When a new model is announced, branch the config. Run the same eval against the new model. Side-by-side diff.
  4. Investigate every regression. Even one case that flipped from pass to fail. Often it surfaces a real edge in your prompt that the old model masked.
  5. Decide. If the new model passes more cases (cleanly), migrate. If it passes the same with lower cost or latency, migrate. If it regresses on any hostile case, hold.

Two regressions you'd otherwise miss.

Format drift. New model produces the same content but with slightly different JSON shape. Downstream parser breaks. Eval catches it; production discovers it.

Tone shift. New model is "more helpful" — adds preamble you didn't ask for. Rubric eval catches it; users tolerate it then quietly switch off.

Three numbers to track per model.

  • Pass rate. % of golden cases passing. The headline number.
  • Cost per case. Average tokens × price. Move horizontally on this curve.
  • Latency per case. p50 and p95. Matters for any interactive surface.
The matrix you actually want

Rows: your prompts. Columns: candidate models. Cells: (pass rate, $/case, latency). Maintained quarterly. This is the single most actionable artifact a small AI team can produce.

A model is a swappable component. Evals are the contract that makes the swap honest.

When the next model drops

Resist the urge to "just try it." Run your eval suite first. The hour of waiting is the difference between deploying with data and deploying with optimism.

06Regression
Reference · When to retire an evalEvaluating AI Systems
Reference

When to retire an eval.


Evals decay. The system changed; the inputs changed; the bug you wrote that case for no longer exists. Hoarding evals is its own quality problem — false signal, wasted time, brittle suite. Retire on purpose.

Signals to retire a case.

  • It has passed for 6 months and the underlying scenario hasn't changed. Safe to archive.
  • It's been flaky for 6 weeks — passes and fails non-deterministically. Either fix the case or remove.
  • The behaviour it was testing for is now caught by a broader, newer case.
  • The expected output became wrong (the world moved). Update or remove.

Signals to add a case.

  • A real production bug. Always becomes a case. Always.
  • A new feature in the prompt. Add at least one happy + one hostile.
  • A new model under evaluation. Add cases that exercise the differences you care about.
  • An onboarding observation. New teammates spot weirdness; that's free signal.

Quarterly eval hygiene.

  1. Re-run the full suite against the current pinned model. Note the new baseline.
  2. Read every failing case. Decide: fix prompt, fix expected, retire case.
  3. Read three randomly-sampled passing cases. Sanity-check the expected output is still right.
  4. Promote any production bugs from the last quarter into permanent cases.
  5. Document the suite size, pass rate, and cost trend in your team's docs.
Colophon

Evaluating AI Systems · The Operator's Library · No. 08. Next: Vol. 09 — Cost & Performance Engineering — the discipline that pairs with this one for production work.

Eval suites are gardens. The work isn't planting once; it's tending the ones that bloom and pulling the ones that don't.

Ten cases. Three types. Pinned models. Quarterly tending. That's the discipline.

— END · OPERATOR'S LIBRARY NO. 08

07Hygiene