The Operator's Library · No. 06
For builders teaching the model your business
The Context
Layer.

RAG, memory, knowledge bases — the unglamorous infrastructure that turns a generic model into one that knows your business. The chapter most "AI doesn't get our company" complaints actually need.

The Context LayerContents & Introduction
Contents & Introduction

The infrastructure underneath the prompts.


A model that doesn't know your business will produce generic answers. A model with a well-built context layer will produce answers grounded in your actual documents, your real prices, last week's decisions. Most "AI doesn't get us" complaints are context-layer complaints in disguise.

Contents.

01Why context beats clevernessThe 80% problem most AI projects haven't seen.p. 03
02RAG, in plain EnglishEmbed, retrieve, augment. The pieces and how they fit.p. 04
03Chunking — the underrated decisionHow you cut up your documents shapes every answer.p. 05
04Beyond RAGMemory, structured retrieval, knowledge graphs.p. 06
§Failure modes & colophonp. 07

Three principles.

  1. Retrieval quality > model quality. A mid-tier model with great retrieval beats a top model with bad retrieval. Spend budget where it moves answers.
  2. Ground every consequential answer. Cite the source. If the answer can't point at a document, treat it as a guess.
  3. The knowledge base is a product. Stale → useless. Build the freshness pipeline before the retrieval pipeline.

RAG is not a new model. It is putting your documents in the model's hand right before it answers — and asking it to read.

02Contents
Ch. 01 · Why context beats clevernessThe Context Layer
01
Chapter One

Why context beats cleverness.


A frontier model is enormously capable. It is also unaware of your customers, your products, your pricing, your decisions from last month. The gap between "smart in general" and "useful here" is filled by context. Cross that gap and the work changes.

The before/after.

Without context. "What's our refund policy?" → Generic answer. Possibly wrong. Possibly invented. Sounds confident either way.

With context. Same question. The model retrieves your actual refund policy page, paraphrases it, cites the section. Right answer, traceable.

What "context" actually is.

  • Documents — policy, product, brand. The corpus.
  • Records — customers, deals, tickets. Structured data.
  • History — past conversations, decisions, runs. The institutional memory.
  • Live data — current prices, inventory, status. Things the snapshot lacks.

Three ways to deliver context.

  1. Paste it in. Works for one-offs. Doesn't scale.
  2. Attach it to a Project. Works up to a few dozen files.
  3. RAG. Works for thousands. The rest of this volume.

The leverage equation.

For most production tasks, output quality is roughly: model strength × retrieval quality × prompt clarity.

Multiplicative. A great prompt with a bad retrieval pulls in garbage and outputs garbage. A great retrieval with a bad prompt gives the right facts the wrong way.

Most teams spend disproportionately on the first factor (chasing newer models) and underinvest in the other two. The order of returns is roughly reversed.

The 80% problem

Eighty per cent of the gap between "generic" and "feels like it knows us" is closed by a working context layer. Not the next model. Not better prompts. Better retrieval.

If your AI feels generic, check your retrieval before you blame the model.

Diagnostic

Ask your current AI tool five questions where the answer should depend on your business specifics. If three or more answers could have been written without ever seeing your company, you have a context-layer problem.

03Chapter 01
Ch. 02 · RAG, in plain EnglishThe Context Layer
02
Chapter Two

RAG, in plain English.


Retrieval-Augmented Generation. The acronym sounds technical; the idea is small. Before the model answers, fetch the relevant documents. Hand them to the model. Ask it to answer based on those. Done.

The RAG loop
QUESTION EMBED RETRIEVE AUGMENT ANSWER user query → vector top-K chunks prompt + chunks grounded

The five moves.

01 · Embed. Convert documents into vector representations — numerical fingerprints that capture meaning. Done once, off-line.

02 · Store. Put the vectors in a vector database (or use a managed index). Each vector points back to its original chunk.

03 · Embed the question. When a question arrives, convert it into the same vector space.

04 · Retrieve. Find the K (typically 3–10) chunks closest to the question.

05 · Augment. Insert those chunks into the prompt. Ask the model to answer using only them. Cite.

What the user sees.

"Based on policy doc §3.2 and the FAQ updated last week, refunds are available within 14 days for unused subscriptions; refer to the linked page for edge cases."

Compare to the un-grounded version: "Most companies offer 14- or 30-day refunds." Same shape; entirely different value.

Hybrid retrieval — almost always.

Pure vector search misses things keyword search catches (acronyms, exact codes, names). Pure keyword misses things vector catches (paraphrases, related concepts). Run both; merge the results. Cheap upgrade; big quality bump.

Watch out

Top-K retrieval can quietly miss the right chunk if K is too small or the chunk is split across an embedding boundary. Tune K against an eval set (Vol. 08), not by vibe.

Start with a tiny corpus

For a first RAG, pick 10–30 documents you know well. Get retrieval working there before you scale to thousands. You'll catch 80% of your design mistakes on the small set, cheaply.

04Chapter 02
Ch. 03 · ChunkingThe Context Layer
03
Chapter Three

Chunking — the underrated decision.


How you cut up your documents shapes every answer the system gives. Most RAG quality problems are chunking problems. Get this right and a lot of other "tuning" stops being necessary.

Three chunking strategies.

01 · Fixed-size (250–800 tokens). Simplest. Works for prose. Fails on structured data, tables, code. Use overlap (50–100 tokens) to avoid splitting concepts.

02 · Semantic. Split at natural boundaries — paragraphs, sections, sentences. Better preservation of meaning; more setup.

03 · Document-aware. Use the structure of the source — H2s in markdown, slide titles in a deck, function bodies in code. The best results live here.

Three rules of thumb.

  • Chunks should answer questions. If your chunk doesn't contain enough context to answer a likely question, it's too small.
  • Chunks should be cite-able. Include a source pointer (doc, section, line) in metadata. Citations only work if you can point.
  • Tables and lists stay together. Splitting a table mid-row breaks retrieval. Treat tables as atomic.

Tune K, tune size, tune overlap.

These three knobs interact. The right defaults to start with:

  • K = 5 (retrieve top 5 chunks).
  • Chunk size = 400–600 tokens.
  • Overlap = 50–100 tokens.

From there, measure on an eval set. Move one knob at a time. Half a day of tuning typically beats half a year of model upgrades.

When chunking is the wrong fix

If the question requires synthesising across many chunks (e.g., "compare our pricing across the last three years"), retrieval alone won't get you there. You may need structured retrieval (Ch. 04), summarisation upstream, or both.

Chunking is the place where most teams stop tuning and start blaming the model. Reverse that order; the model will surprise you.

Tonight

Open three of your chunks at random. Read them in isolation. Could you answer a reasonable question using only that chunk? If two of three say no, your chunks are too small or too misaligned with the document's structure.

05Chapter 03
Ch. 04 · Beyond RAGThe Context Layer
04
Chapter Four

Beyond RAG.


RAG is the right answer to "give the model the relevant documents." It is the wrong answer to "give the model the customer's full record" or "answer questions that span structured data." For those, you need other shapes of retrieval — and sometimes, no retrieval at all.

SHAPE 01 · STRUCTURED

Database queries, not vector search.

"Show me all overdue invoices for client X." There's a correct answer in a database. RAG would approximate it; SQL nails it.

Use tool-use (Vol. 05) to let the model write/issue queries against your structured data.

SHAPE 02 · KNOWLEDGE GRAPH

Entities and relationships.

"Who has worked with both client A and client B?" Vector search struggles with traversal questions. Graphs are built for them.

Use when your domain is genuinely relational. Overkill for content-only domains.

SHAPE 03 · MEMORY

What happened earlier.

The agent's own log of past runs, decisions, edits. Less "look it up" and more "remember what we did." Append-only storage; targeted retrieval.

Build alongside RAG, not inside it. Different access pattern, different operations.

SHAPE 04 · LIVE FETCH

Web search and live APIs.

For anything that changes hourly — prices, weather, news, status. No amount of pre-indexing will keep up.

Tool calls to a search API or your live system. Cache aggressively where staleness is tolerable.

— THE STACK —

Real systems use all four.

A working production system often combines: RAG for content; structured queries for records; memory for run history; live fetch for current data. The "context layer" is plural.

Start with the one your top use case needs. Add the next when a real question demands it.

— FRESHNESS —

The pipeline matters more than the index.

A perfect retrieval over stale data is wrong fast. The unsexy work is the ingestion pipeline — what gets re-indexed when, who flags stale docs, how deletes propagate.

Decide the freshness SLA per source. Daily, hourly, on-event. Build the pipeline to that SLA.

RAG is one tool. The context layer is the toolbox. Match shape to question; don't force every question through embeddings.

06Beyond RAG
Reference · Failure modesThe Context Layer
Reference

Eight failure modes worth memorising.


F·01
Stale index. Docs updated; index didn't. The model "hallucinates" the old answer because that's what was indexed. Build the freshness pipeline.
F·02
Wrong chunk wins. A near-duplicate or out-of-date chunk gets retrieved over the right one. Add filters; de-dupe; date-aware ranking.
F·03
No citation. The model gives an answer but no source. Either you didn't ask, or the chunks lack stable IDs. Both fixable.
F·04
Lost in the middle. You stuffed 30 chunks into context; the model ignored the middle ones. Re-rank; trim; reorder.
F·05
Acronym blindness. Vector search misses "MRR" because no chunk had that exact form. Add keyword retrieval alongside.
F·06
Permission leak. The model surfaces a doc the user shouldn't see. Enforce access at retrieval, not at the prompt.
F·07
Spread question. Question needs synthesis across many chunks; RAG retrieves three; answer is partial. Move to map-reduce or structured.
F·08
Adversarial chunk. A document contains "ignore prior instructions"; the model obeys. Treat retrieved text as untrusted; wrap in safe templating.

"Better context" is a programme, not a setting. Treat it like a database team would — pipelines, monitoring, evals, alerts.

Colophon

The Context Layer · The Operator's Library · No. 06. Next: Vol. 07 — MCP & Integrations — the protocol that lets the model reach the rest of your stack.

Retrieve before you reason. Cite before you ship. Freshness before features.

— END · OPERATOR'S LIBRARY NO. 06

07Failure modes