The unsexy, career-defining volume. How to test, score, and not deploy slop. Golden sets, eval types, regression on model swaps — and the discipline that turns AI work from "it seemed to work" into something you can stand behind.
Most AI projects that fail in production passed every test the team ran — three vibe-checks on the happy path. Evaluation is the discipline that turns "seems good" into "we know." This is the volume that separates a Library reader from a Library operator.
Whoever on your team owns the evals becomes the AI lead within six months. Always. It's the role that compounds.
Software has tests. ML has evals. The two are kin: a small set of known inputs with expected outputs, run automatically, on every change. The cost of skipping them is silent regression. The cost of doing them badly is false confidence. Both are addressable.
An eval suite is not bureaucracy. It is the asset that lets a small team move faster than a big one. Skipping evals is buying short-term speed with long-term paralysis.
An eval suite is your AI system's autonomic nervous system. Without it, every change is a risk; with it, every change is a measurement.
For the highest-stakes prompt in your stack, write three test cases. Happy, edge, hostile. Save them next to the prompt. You now have an eval. The discipline is in the loop, not the perfection.
An eval is four small pieces: a golden set of inputs, expected outputs, a runner that executes the prompt against each, and a scorer that produces a pass/fail. Get the four pieces in place and you have a system. The pieces are smaller than you think.
Start with 10 cases. Five happy path, three edge, two hostile. Not 100; not 5. Ten is enough to catch most regressions and small enough to actually maintain.
Choose hostile cases on purpose. Missing fields. Ambiguous inputs. Inputs that look like prompt injection. The hostile cases are where most production bugs hide.
Pin expected outputs lightly. Don't pin the exact wording (it'll change with every model update). Pin the structural facts — "tone == friendly," "amount == 240," "action == draft."
Source from real history. Best golden cases come from things that actually happened (and broke). Each bug → one new eval case.
A script that loads each case, fires the prompt with the case's input, captures the output. Should be runnable on demand and in CI. Doesn't need to be fancy — 50 lines is plenty.
This is where teams over-engineer. Don't. Match the scorer to the task:
The most important moment in the loop. When a case fails, you should be able to see, side-by-side: input, expected, actual, last-known-passing. A diff that takes 30 seconds to read is the difference between fixing the issue and giving up.
Don't grow the golden set faster than you maintain it. 100 cases that no one reads is worse than 10 that everyone trusts. Quality over volume.
Set up the four pieces for one prompt. Even at toy scale (10 cases, JSON expected output, equality scorer, console diff), the discipline is real. Add cases as bugs surface; not before.
Most teams pick the wrong eval type for the task and then complain that evals are flaky. The trick is matching the type to the work. Three live patterns, each with a clear sweet spot.
Output is structured (JSON, label, number, code). Compare with equality (or near-equality, with floating-point tolerance). Cheap. Deterministic. Doesn't lie.
Use for: classification, extraction, structured generation. The default when applicable.
You define 3–6 binary criteria ("includes a CTA," "mentions price," "under 200 words"). Each case is scored against the rubric. A human or a script tallies.
Use for: drafting, summarising — anything where the output is text but quality is decomposable into checkable parts.
A separate prompt asks a model to compare actual vs. expected against the rubric and return a score. Cheap at scale; introduces its own bias; useful with care.
Use for: subjective quality at volume — but calibrate against human scores periodically, and never use the same model as judge and judged for sensitive tasks.
For RAG systems, the question before "is the answer good" is "did we retrieve the right material." Score against a labelled set of (query → expected chunk IDs).
Use for: any RAG system. Retrieval evals catch problems faster than end-to-end evals.
The judge model's idea of "tone" is its own. Calibrate against human scores on the same cases. If correlation is low, the eval isn't measuring what you think.
Mix: rubric checks + sampled human review. Cheap; correct.
If exact-match works, don't reach for rubric. If rubric works, don't reach for a judge. Spend the complexity budget on cases, not scorers.
Default: exact-match for structure; rubric for prose; judge only when both fail.
Eval type is a design decision, not a default. Pick it consciously per task; mix where useful; document the choice.
A new model release is a software dependency upgrade. Treated like one, it's fine. Treated like "obviously better," it's a Friday-evening incident. Evals are the cheap habit that turns model migrations into a measurement, not a leap of faith.
Format drift. New model produces the same content but with slightly different JSON shape. Downstream parser breaks. Eval catches it; production discovers it.
Tone shift. New model is "more helpful" — adds preamble you didn't ask for. Rubric eval catches it; users tolerate it then quietly switch off.
Rows: your prompts. Columns: candidate models. Cells: (pass rate, $/case, latency). Maintained quarterly. This is the single most actionable artifact a small AI team can produce.
A model is a swappable component. Evals are the contract that makes the swap honest.
Resist the urge to "just try it." Run your eval suite first. The hour of waiting is the difference between deploying with data and deploying with optimism.
Evals decay. The system changed; the inputs changed; the bug you wrote that case for no longer exists. Hoarding evals is its own quality problem — false signal, wasted time, brittle suite. Retire on purpose.
Evaluating AI Systems · The Operator's Library · No. 08. Next: Vol. 09 — Cost & Performance Engineering — the discipline that pairs with this one for production work.
Eval suites are gardens. The work isn't planting once; it's tending the ones that bloom and pulling the ones that don't.
Ten cases. Three types. Pinned models. Quarterly tending. That's the discipline.
— END · OPERATOR'S LIBRARY NO. 08