What large language models actually are. What they're good at, what they're bad at, and how to design work around both — without mysticism and without dismissal.
Most disappointment with AI is the cost of a bad mental model. People treat LLMs like databases ("look up this fact") or like coworkers ("remember what we discussed") and are confused when they fail at both. This primer gives you a model that survives contact with the actual product.
The model is brilliant at language and frequently wrong about facts. Design work that benefits from the brilliance and is robust to the wrongness.
A large language model is, at its mechanical core, a function that takes a sequence of text and predicts the most likely next chunk of text. That's the whole engine. Everything else — chat, code, reasoning, agentic behaviour — is built on top of that one move, repeated.
01 · Trained, then frozen. The model learned patterns from a huge corpus of text. After training, the weights are fixed. It cannot learn from your conversation. Whatever it "remembers" is sitting in the prompt window with you.
02 · Next-token at a time. Output is generated one chunk at a time, each chunk influencing the next. There is no plan, no goal, no intent — just an extremely sophisticated guess at what word comes next given everything so far.
03 · The context window is the workspace. Everything the model "sees" while responding fits in a finite window. Things outside the window do not exist for that response. Long context helps; infinite context isn't a thing.
04 · Temperature is sampling, not creativity. "More creative" outputs are statistically less-likely tokens. Lower temperature is more conservative — and often more accurate for structured tasks.
Because the model predicts tokens, it is exceptionally good at the things language is shaped like — paraphrasing, translation, summarisation, drafting, transformation. It is much weaker at the things language merely references — facts, math, current events, anything that isn't already heavily encoded in the patterns it learned.
The trick is to lean on the strengths. Ask the model to shape information you provide. Don't ask it to retrieve information you assume it has.
The model's training data has a cutoff date. Anything after that date — events, prices, library versions, your customers — is invisible to it unless you put it in the prompt. Plan for this. (Ch. 06 of this series — RAG — is built around fixing it.)
For one day, every time you read an AI output, mentally ask: "would this answer be in the training data, or is it being recombined on the fly?" The shape of your trust will start to match the shape of the answer.
A hallucination is not a glitch. It is what a next-token predictor naturally produces when asked a question it doesn't have strong signal for. Understanding the mechanism is how you stop being surprised by it — and how you design work that catches it.
Imagine you're asked to finish a sentence about an obscure 1937 jazz musician you've never heard of. You'd pause, say "I don't know." The model can't pause. Its job is to produce the next likely token. So it produces something that sounds like the kind of sentence that would finish that question. Plausible-shaped, fact-shaped — but invented.
This is the same machine that gets right answers when the question has strong signal in the training data. Right answers and hallucinations come out of the same engine, in the same voice, with the same confidence.
01 · Ground the answer. Put the facts in the prompt. RAG, file attachments, copy-paste. If the answer is in the context, the model paraphrases instead of invents.
02 · Ask for sources. "Cite the line in the document where this claim is supported." Forces the model to point at evidence rather than narrate.
03 · Use tools for facts. Calculator for math; search for current events; database for your business data. The model orchestrates; the tools answer.
04 · Build verification into the workflow. For high-stakes outputs, a second model or rule check, then a human. A hallucination caught is just a draft revision.
Train yourself to ask "what is this answer based on?" after every consequential output. If the answer is "the model's training data, somewhere," that's the signal to verify. If the answer is "the document I pasted in," trust climbs.
A brief honest tour. Knowing the failure modes is more useful than knowing the capabilities — because the failure modes are how a careful operator scopes the work.
Models generate digit-shaped sequences, not computations. 7-digit multiplication, percentage calculations, statistical aggregates — wrong more often than right past trivial scale.
Fix: hand it to a calculator tool or a code interpreter. Never trust unverified math in an output.
"Give me exactly 7 examples." Often gets 6 or 8. Counting tokens, words, list lengths — surprisingly unreliable.
Fix: post-process. Or ask for many, then trim deterministically downstream.
News, pricing, library versions, who's CEO right now. The model is a snapshot. The world has moved.
Fix: provide current information in the prompt, or give the model a search tool. Both work.
Will invent plausible-looking paper titles, URLs, case-law citations. The shape is right; the URL is dead.
Fix: only cite from documents in context, or verify every URL programmatically.
"Are you sure?" rarely changes the model's mind in useful ways. It will flip an answer to please you, or double-down on a wrong one. Asking calibration questions of the model is a vibe, not a measurement.
Fix: external verification, not internal interrogation.
Coherent across a paragraph; drifts across a chapter; loses the thread across a project. The longer the chain, the more compounded the small errors.
Fix: decompose. Many short focused calls beat one long sprawling one. (Ch. 05 of this series — Designing Agents — turns this fix into a design pattern.)
Lean on what the model does well — language transformations. Outsource what it does badly — facts, math, structure — to other tools.
Before you ship any AI workflow, scan the six weaknesses. If your workflow depends on the model being strong at one of them, redesign before you ship — not after.
All of the above collapses to a single, useful, deliberately humble metaphor. Hold this metaphor and most of your decisions about AI get easier.
An extremely well-read intern with no memory, no internet, and an unshakeable inability to say "I don't know."
Run this metaphor through every decision:
Why pasting the brief works: the intern can read. Give it the context and it will use it.
Why "remember when I told you…" fails: the intern doesn't remember. Re-tell.
Why "what's the weather in Tokyo" needs a tool: the intern doesn't have internet. Give it one.
Why long autonomous loops drift: the intern has no manager checking in. Build the checkpoints in (Ch. 14 of the BPA volume — approval gates).
Why structured prompts beat vague ones: a clear brief gets clear work. Always has.
It refuses to flatter the model into mysticism, and it refuses to dismiss it as "just autocomplete." Both stances are wrong. It is a tool with specific shape; design for the shape, get the leverage.
Every other volume in this Library — Prompting, Working with Claude, BPA, Agents, RAG — is essentially an instruction manual for managing a brilliant intern at scale. Once the metaphor clicks, the rest is craft.
Whenever an AI output surprises you, ask: "would this surprise me if I'd handed this task to a very well-read intern with no memory?" The frame nearly always explains the surprise — and points at the fix.
A single page to keep next to your laptop. Read across: the task type, the failure shape, the right stance. Memorise the top three rows; the rest you can look up.
"Trust, verify, walk away" is not a vibe — it's a routing decision you make at the start of every task. Get faster at making it; everything downstream gets easier.
Pick three current AI tasks in your week. Place each one on the table. If any is in "trust" but should be in "verify" — or vice versa — adjust the workflow before next Tuesday.
Wrong question for an operator. Useful question: "is it useful, at what tasks, with what failure modes?" Skip the philosophy; build the workflow.
It will reduce them; it will not eliminate them. The architecture predicts text. Models will get more careful and more often grounded, but a system that needs zero hallucination still needs verification.
The dominant cost of getting value from AI is process, not capability. A team that builds the workflow now with a current model will out-execute a team waiting for the next one.
No. For most production tasks, a mid-tier model with a strong eval and good context beats a large model with weak both. (Ch. 09 — Cost & Performance.)
It can produce false statements, including ones contradicting available evidence. Whether that's "lying" is philosophy. What's relevant: it can happen, and design assumes it might.
It models language patterns produced by people who understand. The outputs are often indistinguishable from understanding for practical purposes. Be careful when "indistinguishable for practical purposes" stops being good enough — emotional contexts, especially.
Like any powerful tool, yes — in some uses. The risk profile is dominated by misuse (deepfakes, automated abuse) and by overreliance (uncritical adoption in consequential decisions). Both are addressable with governance. (Ch. 10.)
You stop being surprised. Surprise is the cost of a bad mental model; calm is the dividend of a good one. The rest of the Library shows you what to build with the calm.
The Operator's Mental Model of AI — A Primer. The Operator's Library · No. 01.
Set in Source Serif 4 and Outfit. Code in JetBrains Mono. Paper: warm cream. Accent: steel blue.
Read No. 02 next — Prompting Like a Pro — for the craft layer on top of this model.
The model is a tool. The work is yours. The leverage is real.
— END · OPERATOR'S LIBRARY NO. 01