RAG, memory, knowledge bases — the unglamorous infrastructure that turns a generic model into one that knows your business. The chapter most "AI doesn't get our company" complaints actually need.
A model that doesn't know your business will produce generic answers. A model with a well-built context layer will produce answers grounded in your actual documents, your real prices, last week's decisions. Most "AI doesn't get us" complaints are context-layer complaints in disguise.
RAG is not a new model. It is putting your documents in the model's hand right before it answers — and asking it to read.
A frontier model is enormously capable. It is also unaware of your customers, your products, your pricing, your decisions from last month. The gap between "smart in general" and "useful here" is filled by context. Cross that gap and the work changes.
Without context. "What's our refund policy?" → Generic answer. Possibly wrong. Possibly invented. Sounds confident either way.
With context. Same question. The model retrieves your actual refund policy page, paraphrases it, cites the section. Right answer, traceable.
For most production tasks, output quality is roughly: model strength × retrieval quality × prompt clarity.
Multiplicative. A great prompt with a bad retrieval pulls in garbage and outputs garbage. A great retrieval with a bad prompt gives the right facts the wrong way.
Most teams spend disproportionately on the first factor (chasing newer models) and underinvest in the other two. The order of returns is roughly reversed.
Eighty per cent of the gap between "generic" and "feels like it knows us" is closed by a working context layer. Not the next model. Not better prompts. Better retrieval.
If your AI feels generic, check your retrieval before you blame the model.
Ask your current AI tool five questions where the answer should depend on your business specifics. If three or more answers could have been written without ever seeing your company, you have a context-layer problem.
Retrieval-Augmented Generation. The acronym sounds technical; the idea is small. Before the model answers, fetch the relevant documents. Hand them to the model. Ask it to answer based on those. Done.
01 · Embed. Convert documents into vector representations — numerical fingerprints that capture meaning. Done once, off-line.
02 · Store. Put the vectors in a vector database (or use a managed index). Each vector points back to its original chunk.
03 · Embed the question. When a question arrives, convert it into the same vector space.
04 · Retrieve. Find the K (typically 3–10) chunks closest to the question.
05 · Augment. Insert those chunks into the prompt. Ask the model to answer using only them. Cite.
"Based on policy doc §3.2 and the FAQ updated last week, refunds are available within 14 days for unused subscriptions; refer to the linked page for edge cases."
Compare to the un-grounded version: "Most companies offer 14- or 30-day refunds." Same shape; entirely different value.
Pure vector search misses things keyword search catches (acronyms, exact codes, names). Pure keyword misses things vector catches (paraphrases, related concepts). Run both; merge the results. Cheap upgrade; big quality bump.
Top-K retrieval can quietly miss the right chunk if K is too small or the chunk is split across an embedding boundary. Tune K against an eval set (Vol. 08), not by vibe.
For a first RAG, pick 10–30 documents you know well. Get retrieval working there before you scale to thousands. You'll catch 80% of your design mistakes on the small set, cheaply.
How you cut up your documents shapes every answer the system gives. Most RAG quality problems are chunking problems. Get this right and a lot of other "tuning" stops being necessary.
01 · Fixed-size (250–800 tokens). Simplest. Works for prose. Fails on structured data, tables, code. Use overlap (50–100 tokens) to avoid splitting concepts.
02 · Semantic. Split at natural boundaries — paragraphs, sections, sentences. Better preservation of meaning; more setup.
03 · Document-aware. Use the structure of the source — H2s in markdown, slide titles in a deck, function bodies in code. The best results live here.
These three knobs interact. The right defaults to start with:
From there, measure on an eval set. Move one knob at a time. Half a day of tuning typically beats half a year of model upgrades.
If the question requires synthesising across many chunks (e.g., "compare our pricing across the last three years"), retrieval alone won't get you there. You may need structured retrieval (Ch. 04), summarisation upstream, or both.
Chunking is the place where most teams stop tuning and start blaming the model. Reverse that order; the model will surprise you.
Open three of your chunks at random. Read them in isolation. Could you answer a reasonable question using only that chunk? If two of three say no, your chunks are too small or too misaligned with the document's structure.
RAG is the right answer to "give the model the relevant documents." It is the wrong answer to "give the model the customer's full record" or "answer questions that span structured data." For those, you need other shapes of retrieval — and sometimes, no retrieval at all.
"Show me all overdue invoices for client X." There's a correct answer in a database. RAG would approximate it; SQL nails it.
Use tool-use (Vol. 05) to let the model write/issue queries against your structured data.
"Who has worked with both client A and client B?" Vector search struggles with traversal questions. Graphs are built for them.
Use when your domain is genuinely relational. Overkill for content-only domains.
The agent's own log of past runs, decisions, edits. Less "look it up" and more "remember what we did." Append-only storage; targeted retrieval.
Build alongside RAG, not inside it. Different access pattern, different operations.
For anything that changes hourly — prices, weather, news, status. No amount of pre-indexing will keep up.
Tool calls to a search API or your live system. Cache aggressively where staleness is tolerable.
A working production system often combines: RAG for content; structured queries for records; memory for run history; live fetch for current data. The "context layer" is plural.
Start with the one your top use case needs. Add the next when a real question demands it.
A perfect retrieval over stale data is wrong fast. The unsexy work is the ingestion pipeline — what gets re-indexed when, who flags stale docs, how deletes propagate.
Decide the freshness SLA per source. Daily, hourly, on-event. Build the pipeline to that SLA.
RAG is one tool. The context layer is the toolbox. Match shape to question; don't force every question through embeddings.
"Better context" is a programme, not a setting. Treat it like a database team would — pipelines, monitoring, evals, alerts.
The Context Layer · The Operator's Library · No. 06. Next: Vol. 07 — MCP & Integrations — the protocol that lets the model reach the rest of your stack.
Retrieve before you reason. Cite before you ship. Freshness before features.
— END · OPERATOR'S LIBRARY NO. 06