The discipline that separates a $40-a-month workflow from a $4,000-a-month one. Model selection, caching, batching, structured output. The boring optimisations that cut bills by 80% without touching quality.
The promise of LLM cost engineering is unusual: you can frequently get cheaper, faster, and better — at the same time — by applying disciplines most teams skip. This handbook is the short list of those disciplines.
The single biggest cost optimisation in AI is using the smallest model that passes the eval. Most teams reach for the biggest one out of caution. Reverse the default.
Almost every "AI is expensive" story is a context-cost story in disguise. Once you can see the bill broken down, the optimisations are obvious — and most of them don't require touching the prompt at all.
01 · Input tokens. Everything you send — system prompt, user message, attached files, retrieved chunks. Usually cheaper per token than output.
02 · Output tokens. What the model generates. Typically 3–5× the per-token cost of input. The shorter you can make outputs while keeping quality, the more you save.
03 · Multipliers. Tool calls, retries, multi-turn conversations. A "single" agent run can be 10–50 model calls under the hood.
04 · Hidden volume. The endpoint that's called 100× a day by an internal system you forgot existed. Almost always present; rarely audited.
| Feature | Calls | Avg in/out | Weekly $ |
| Support draft | 2,400 | 3.2k / 0.4k | $320 |
| Lead research | 180 | 12k / 1.8k | $140 |
| Inbox triage | 9,000 | 0.5k / 0.1k | $95 |
| Weekly brief | 1 | 180k / 4k | $8 |
| Other | $24 | ||
The story: support drafts dominate; lead research is expensive per call; inbox triage is high volume / tiny calls; weekly brief is irrelevant cost-wise. Each gets a different optimisation.
For every feature: what's the cost per useful output? A $0.01 call is "cheap" — until you discover it runs 500,000× a month. Always compute per-call cost and per-feature volume.
Build the bill table for your week. If your provider doesn't break it out, instrument it yourself — log model, tokens in, tokens out, feature, on every call. One day of plumbing, a year of clarity.
Default to the cheapest model that passes the eval. This is the single most actionable line in this handbook. Most teams default to the most capable model "for safety" — and pay 5–10× what they need to for tasks that don't require it.
Classification, extraction, simple drafts, routing decisions. 80% of the calls in a mature stack. Fast enough for interactive use; cheap enough to run a lot.
Default for: front-of-chain steps, structured extraction, high-volume / low-stakes work.
Daily-driver. Drafts, summaries, research, agentic loops. Handles 80%+ of professional tasks well.
Default for: the actual product work. Unless you have data saying otherwise, start here.
Hard reasoning, long-horizon agents, complex code. Slower; more expensive. Worth it when the failure mode of a cheaper model is real.
Default for: nothing. Promote to it only when an eval case fails on your mid-tier model and the failure is reasoning-bound.
Run the cheap model first. If confidence is low or output fails validation, escalate to the next tier. Most calls never need to escalate; the few that do still get answered. Cost falls; quality stays.
Standard pattern: small-classify → medium-draft → large only for hard cases.
Public benchmarks don't reflect your eval set. The model that wins on MMLU may flop on your support tickets. Benchmark on your own cases.
Re-run your eval (Vol. 08) against each candidate. Pick from your numbers.
Record the exact provider and model version, not just the model family. Version drift can subtly change outputs. Treat model name as part of the system; rev it deliberately.
Update version → re-run eval (BPA Ch. 10) → migrate.
"Just use the best model" is an answer for prototypes. Production is a different problem, with a different answer.
Pick one production prompt. Run the eval against two model tiers below current. Pick the cheapest that passes. Migrate. Track the bill change. Most teams find 40–70% cost reduction here.
Three optimisations that pay back the day you ship them. Each is independent; each is easy to add to an existing system; together they routinely cut cost and latency by 50–80%.
Prompt cache. Many requests share the same long system prompt, the same context, the same retrieved chunks. Cache those at the provider level (prompt caching) or your own (Redis with hash key). The first request pays full price; subsequent ones pay a fraction.
Semantic cache. Different questions that mean the same thing get the same cached answer. Embed the question; if a close enough match exists with a fresh answer, return it. Game-changer for high-volume Q&A.
Don't cache. Anything stateful, personal, or time-sensitive. The category boundary is the design.
For non-interactive work (overnight runs, weekly briefs, bulk classification), batch into single requests where the API supports it, or fire in parallel with rate-limit awareness. Both reduce per-call overhead.
Batch APIs from major providers also typically discount per-token rates substantially. Read your provider's batch pricing; it's often 50%.
Force the model to emit JSON (or another schema) with provider-side structured output mode. Three benefits, immediately:
Aggressive caching can serve stale data. Set a TTL per cache tier; invalidate explicitly on writes; have a way to bypass cache for debugging.
Most "we need a faster model" requests are actually "we need a smaller prompt." Audit before you upgrade.
Day 1: turn on provider prompt caching. Day 2: convert your top-cost prompt to structured output. Day 3–5: instrument latency. Most teams cut bill by a third in week one.
Premature optimisation in AI is the same disease it is in software — wasted effort on things that don't matter. The fix is the same too: measure first, optimise the dominant line item, ship, repeat.
Run the pass quarterly. Track the bill curve. A healthy AI-native system shows cost per useful output declining over time even as volume grows.
AI Cost & Performance Engineering · The Operator's Library · No. 09. Next: Vol. 10 — AI Safety & Governance for Small Teams.
Measure, right-size, cache, structure, trim. In that order. Repeat quarterly.
Cheap, fast, good — pick all three by doing the four-knob pass.
— END · OPERATOR'S LIBRARY NO. 09