A working reference for what AI actually costs and what to do about it. Written for the engineer who has to make the change and the person who has to sign the invoice — which, increasingly, are the same person.
Why this is free, and why it is unfinished.
Cost work is mostly arithmetic and attention, and the arithmetic should not be behind a sales call. Chapters go up as they are written and get revised as the ground moves — this is a living document, not a launch. Where a number appears, it is either computed live from the price index or sourced and dated. Where something is a judgment call, it says so.
42 of 52 chapters published · about 301 minutes of reading so far · last revised 2026-08-11
You buy tokens from someone else's endpoint. Your levers are what you send, what you ask for back, and how often. Most teams live here and never leave, and that is often correct.
You rent accelerators by the hour instead of buying tokens. Cost per token becomes a function of how busy you keep the hardware. Relevant once volume is steady and large enough to fill a GPU.
Below the hourly rate is a physical machine with a fixed memory bandwidth and a fixed amount of VRAM. Those two numbers set the ceiling on everything above. Relevant when you are choosing accelerators or explaining why the cheap one was not cheaper.
The same GPU serves five requests a second or fifty depending entirely on how you schedule work onto it. This is where cost per token is actually decided, and where the largest wins live for anyone running their own inference.
Savings decay. This part is the practice that holds them: attribution, unit economics, guardrails, and the operating cadence. Applies at every layer above.
Why the unwritten chapters are listed.
Because the map is worth more than the territory covered so far. Publishing the whole outline says what this intends to be, lets you tell me which chapter you actually need next, and fixes each chapter’s address from the day it is planned — so a link written today still resolves when the chapter lands. Every infrastructure figure is priced on AWS, GCP, and Azure; a chapter citing one cloud is a draft, not a chapter. Tell me what is missing and it moves up the list.