The Operator's Library · No. 09
For operators whose bills have arrived
AI Cost &
Performance.

The discipline that separates a $40-a-month workflow from a $4,000-a-month one. Model selection, caching, batching, structured output. The boring optimisations that cut bills by 80% without touching quality.

AI Cost & Performance EngineeringContents & Introduction
Contents & Introduction

Cheap, fast, good — pick all three.


The promise of LLM cost engineering is unusual: you can frequently get cheaper, faster, and better — at the same time — by applying disciplines most teams skip. This handbook is the short list of those disciplines.

Contents.

01Anatomy of an LLM billWhere the money actually goes. Where you can cut.p. 03
02Model selectionSmall, medium, large. The right tool per task, not per team.p. 04
03Caching, batching, structured outputThe three highest-ROI optimisations.p. 05
04When to optimise (and when not to)The discipline of waiting until it matters.p. 06

Three principles.

  1. Measure before optimising. "It feels slow / expensive" is not data. Token logs are. Build them first.
  2. Right-size the model. A smaller model that passes your eval is a better production model than a larger model that does the same thing slower and at higher cost.
  3. Cache aggressively. Most LLM workloads have huge cache hit rates if you let them. Most teams don't even try.

The single biggest cost optimisation in AI is using the smallest model that passes the eval. Most teams reach for the biggest one out of caution. Reverse the default.

02Contents
Ch. 01 · Anatomy of an LLM billAI Cost & Performance Engineering
01
Chapter One

Anatomy of an LLM bill.


Almost every "AI is expensive" story is a context-cost story in disguise. Once you can see the bill broken down, the optimisations are obvious — and most of them don't require touching the prompt at all.

Where the money goes.

01 · Input tokens. Everything you send — system prompt, user message, attached files, retrieved chunks. Usually cheaper per token than output.

02 · Output tokens. What the model generates. Typically 3–5× the per-token cost of input. The shorter you can make outputs while keeping quality, the more you save.

03 · Multipliers. Tool calls, retries, multi-turn conversations. A "single" agent run can be 10–50 model calls under the hood.

04 · Hidden volume. The endpoint that's called 100× a day by an internal system you forgot existed. Almost always present; rarely audited.

The diagnostic.

  1. Pull a week of token logs. Per call: model, input tokens, output tokens, cost.
  2. Group by feature. Sort by cost descending.
  3. Look at the top 3 features. 80% of the bill lives there.
  4. Apply Ch. 02–03 to those three. Ignore the rest until they grow.

The bill table — example.

WEEKLY BILL · ACME CO · ILLUSTRATIVE
FeatureCallsAvg in/outWeekly $
Support draft2,4003.2k / 0.4k$320
Lead research18012k / 1.8k$140
Inbox triage9,0000.5k / 0.1k$95
Weekly brief1180k / 4k$8
Other$24

The story: support drafts dominate; lead research is expensive per call; inbox triage is high volume / tiny calls; weekly brief is irrelevant cost-wise. Each gets a different optimisation.

The unit cost question

For every feature: what's the cost per useful output? A $0.01 call is "cheap" — until you discover it runs 500,000× a month. Always compute per-call cost and per-feature volume.

Tomorrow morning

Build the bill table for your week. If your provider doesn't break it out, instrument it yourself — log model, tokens in, tokens out, feature, on every call. One day of plumbing, a year of clarity.

03Chapter 01
Ch. 02 · Model selectionAI Cost & Performance Engineering
02
Chapter Two

Model selection. Right-size, every time.


Default to the cheapest model that passes the eval. This is the single most actionable line in this handbook. Most teams default to the most capable model "for safety" — and pay 5–10× what they need to for tasks that don't require it.

TIER 01 · HAIKU-CLASS

Small, fast, cheap.

Classification, extraction, simple drafts, routing decisions. 80% of the calls in a mature stack. Fast enough for interactive use; cheap enough to run a lot.

Default for: front-of-chain steps, structured extraction, high-volume / low-stakes work.

TIER 02 · SONNET-CLASS

The workhorse.

Daily-driver. Drafts, summaries, research, agentic loops. Handles 80%+ of professional tasks well.

Default for: the actual product work. Unless you have data saying otherwise, start here.

TIER 03 · LARGE MODEL

The heavy lifter.

Hard reasoning, long-horizon agents, complex code. Slower; more expensive. Worth it when the failure mode of a cheaper model is real.

Default for: nothing. Promote to it only when an eval case fails on your mid-tier model and the failure is reasoning-bound.

— THE STRATEGY —

Cascade.

Run the cheap model first. If confidence is low or output fails validation, escalate to the next tier. Most calls never need to escalate; the few that do still get answered. Cost falls; quality stays.

Standard pattern: small-classify → medium-draft → large only for hard cases.

— DON'T —

Pick by the leaderboard.

Public benchmarks don't reflect your eval set. The model that wins on MMLU may flop on your support tickets. Benchmark on your own cases.

Re-run your eval (Vol. 08) against each candidate. Pick from your numbers.

— DO —

Pin the model version.

Record the exact provider and model version, not just the model family. Version drift can subtly change outputs. Treat model name as part of the system; rev it deliberately.

Update version → re-run eval (BPA Ch. 10) → migrate.

"Just use the best model" is an answer for prototypes. Production is a different problem, with a different answer.

This week

Pick one production prompt. Run the eval against two model tiers below current. Pick the cheapest that passes. Migrate. Track the bill change. Most teams find 40–70% cost reduction here.

04Selection
Ch. 03 · The three highest-ROI optimisationsAI Cost & Performance Engineering
03
Chapter Three

Caching, batching, structured output.


Three optimisations that pay back the day you ship them. Each is independent; each is easy to add to an existing system; together they routinely cut cost and latency by 50–80%.

Caching — prompt-level & semantic.

Prompt cache. Many requests share the same long system prompt, the same context, the same retrieved chunks. Cache those at the provider level (prompt caching) or your own (Redis with hash key). The first request pays full price; subsequent ones pay a fraction.

Semantic cache. Different questions that mean the same thing get the same cached answer. Embed the question; if a close enough match exists with a fresh answer, return it. Game-changer for high-volume Q&A.

Don't cache. Anything stateful, personal, or time-sensitive. The category boundary is the design.

Batching.

For non-interactive work (overnight runs, weekly briefs, bulk classification), batch into single requests where the API supports it, or fire in parallel with rate-limit awareness. Both reduce per-call overhead.

Batch APIs from major providers also typically discount per-token rates substantially. Read your provider's batch pricing; it's often 50%.

Structured output.

Force the model to emit JSON (or another schema) with provider-side structured output mode. Three benefits, immediately:

  • No retry from parse failures. Bad JSON → silent retry → 2× cost. Structured output prevents the bad shape.
  • Shorter outputs. JSON is more compact than prose for the same data. Fewer output tokens.
  • Cleaner chains. The next step can rely on the shape. (BPA Ch. 08.)
Watch out

Aggressive caching can serve stale data. Set a TTL per cache tier; invalidate explicitly on writes; have a way to bypass cache for debugging.

Three more, smaller, also worth it.

  • Shorter prompts. Audit your system prompt. Most are 2× longer than they need to be.
  • Trim retrievals. K=5 not K=20. Re-rank, don't stuff.
  • Streaming where it helps. Doesn't reduce cost, but reduces perceived latency dramatically for interactive use.

Most "we need a faster model" requests are actually "we need a smaller prompt." Audit before you upgrade.

One week's work

Day 1: turn on provider prompt caching. Day 2: convert your top-cost prompt to structured output. Day 3–5: instrument latency. Most teams cut bill by a third in week one.

05Optimisations
Ch. 04 · When to optimiseAI Cost & Performance Engineering
04
Chapter Four

When to optimise (and when not to).


Premature optimisation in AI is the same disease it is in software — wasted effort on things that don't matter. The fix is the same too: measure first, optimise the dominant line item, ship, repeat.

When to optimise.

  • Monthly bill exceeds $X you'd planned for.
  • One feature dominates the bill (Pareto-style).
  • Latency is causing user complaints or breaking SLAs.
  • You're moving from prototype to production.
  • You're about to add a model dependency that 10× your call volume.

When not to.

  • You're still validating product-market fit. Quality matters; cost barely.
  • The feature you'd optimise costs $20/month and works fine. Use the time better.
  • You don't have evals (Vol. 08). Optimising without evals breaks things you can't measure.
  • You're optimising for vanity — "we use the smallest model because it's clever" — instead of for an actual constraint.

The four-knob optimisation pass.

  1. Measure (Ch. 01). Per-feature bill, per-call cost.
  2. Right-size model (Ch. 02). Cascade where useful.
  3. Cache and structure (Ch. 03). Provider cache + structured output first; semantic cache if Q&A.
  4. Trim prompts and retrievals. Remove the dead weight; tune K.

Run the pass quarterly. Track the bill curve. A healthy AI-native system shows cost per useful output declining over time even as volume grows.

Colophon

AI Cost & Performance Engineering · The Operator's Library · No. 09. Next: Vol. 10 — AI Safety & Governance for Small Teams.

Measure, right-size, cache, structure, trim. In that order. Repeat quarterly.

Cheap, fast, good — pick all three by doing the four-knob pass.

— END · OPERATOR'S LIBRARY NO. 09

06When to optimise