The sequel to the BPA guide. What an agent actually is, the loop that makes one work, the tools that give it reach, the memory that lets it persist — and the failure modes you only learn by shipping one.
An agent is not a smarter chatbot. It is a system: a model in a loop, with tools, a goal, and a memory. Build it casually and you get a confidently-wrong autopilot. Build it deliberately and you get the most leveraged piece of software your business has ever had.
Anyone can spin up an agent in an afternoon. The work that separates a demo from a system is the boring part — tools, memory, observability, gates. Do that work.
An agent is a model placed inside a loop, given tools, a goal, and a way to remember. None of those pieces are exotic. Putting them together with discipline is what separates working agents from impressive demos.
01 · Goal. What the agent is trying to do, stated narrowly enough that "done" is recognisable. "Reply to inbound press requests with a draft" — not "handle press."
02 · Perception. How the agent receives inputs — a new email, a webhook, a file appearing. Sets the rhythm.
03 · The model. The reasoning core. Decides, in each turn of the loop, what to do next.
04 · Tools. The things the agent can actually do — read a doc, send an email, query a DB, run a calculation. The model decides; the tools act.
05 · Memory. What persists between turns and across runs. Short-term context, long-term knowledge, episodic logs.
A BPA-style chain (Vol. 04) is a fixed sequence of skills. The path is known up front; the agent's "choices" are predetermined.
An agent, by contrast, decides its own path turn by turn. Same model, same tools, different runs — different sequences of actions. That flexibility is the strength and the danger.
When the next step legitimately depends on what was learned in the previous step. If the path could be known up front, a chain is cheaper and more reliable. Use agents only where the dynamism is real.
Agents are not the universal upgrade to chains. They are the right tool for the smaller class of problems where the path itself is part of the work.
For your candidate agent: can you write out the exact sequence of steps it should take on a normal day? If yes, you want a chain (BPA Ch. 11), not an agent.
Every working agent runs the same loop: perceive, plan, act, observe. The names are simple. The discipline is making sure each move actually happens — many "agents" skip "observe" and then wonder why they drift.
01 · Perceive. Read the current state — the new input, the memory, the prior turn's result. Don't assume; explicitly retrieve.
02 · Plan. Decide the next single action. One thing. Resist the temptation to plan ten steps ahead — the world will change after step one anyway.
03 · Act. Call the tool. Send the request. Update the record. Generate the draft. Exactly one action.
04 · Observe. Read what came back. Did the action succeed? What did it return? Update memory. Then loop.
Observe. Beginners design agents that act, act, act, then check at the end. The agent has by then made six decisions on stale information. Real agents read after every action and pay the small cost to stay synchronised with reality.
An agent without an explicit step budget will, eventually, find a way to spend $40 of tokens on a $0.40 task. Always set the budget. Always alert when it's hit.
A loop without a stopping condition isn't an agent. It's a thermostat with a stuck contact. Build the off-switch first.
For any agent you build, write down the three stopping conditions before writing the loop. If you can't, you don't have a clear-enough goal yet — go back to Ch. 01.
An agent without tools is a brilliant intern locked in a closet. The tools are the agent's hands. Where most "agents don't work" stories really live: bad tools, bad schemas, bad seams between the model and the world.
01 · One job. A tool that does three things forces the agent to remember which mode it was in. Split.
02 · A clear name. send_invoice_email beats email_v2. The name is part of the prompt the model sees.
03 · A clean schema. Required vs. optional parameters explicit. Types declared. The model uses the schema to decide if a tool is right.
04 · A helpful description. One sentence on what the tool does, one on when to use it, one on when not to. The model reads these like docs.
05 · Structured outputs. Tools return JSON the model can parse, not prose it has to interpret.
06 · Idempotency where possible. Calling the same tool twice with the same input should be safe. Saves you when the agent retries.
Bad tools eat good agents alive. Spend more time on schemas than on prompts.
Write the schemas and descriptions for all your tools first. Show them to a colleague. If they can't tell from the descriptions when to use each, the model can't either.
An agent without memory rebuilds the world every turn. An agent with too much memory chokes on it. The art is keeping the right shape of memory at the right scale — three distinct kinds, each with its own job.
01 · Short-term (context). The conversation/loop so far. Lives in the context window. Cheap and current. Resets at the end of the run.
02 · Long-term (knowledge base). Facts that don't change — your brand voice, your product catalog, your customers. Stored externally; retrieved with RAG (Vol. 06). Doesn't fit in context; doesn't need to.
03 · Episodic (run logs). What this agent did, when, and why. Stored append-only. Used for debugging, learning, and audit. Often forgotten by first-time builders; non-negotiable in production.
Each memory needs four operations done well:
"Just feed the whole memory into context each turn" works for the first 50 runs. It fails badly after that — context gets long, costs balloon, the model loses signal in the middle of dense logs. Build retrieval from day one.
Memory design separates a one-shot agent from an agent that gets better at your business over time. Design it like a database, not like a diary.
For your agent, list ten things it "should remember." Now assign each to short-term, long-term, or episodic. The exercise itself usually surfaces a missing tool or a missing schema.
Multiple agents collaborating is the most impressive demo and the easiest place to over-engineer. Use these patterns only when the work genuinely needs perspectives a single agent in one pass can't provide. Otherwise: one strong agent beats five mediocre ones.
A top-level agent reads the situation, calls a specialist sub-agent for the specific kind of work, takes back the result, decides next. Clean separation; easy to reason about.
Use when: the work has clearly distinct sub-tasks. Common shape for research → write → review.
One advocates; one criticises; a third (or you) decides. Improves quality on judgement-heavy decisions. Expensive; slow; sometimes worth it.
Use when: decision is consequential and "one perspective" is the failure mode. Avoid for routine work.
Same agent run in parallel on independent items, then a synthesiser combines. Like map-reduce, for cognitive work.
Use when: items are independent and a structured merge is possible. Excellent for research across companies.
Specialised roles with clear handoffs. Writer drafts; editor improves; fact-checker verifies. Each handoff is a checkpoint.
Use when: the deliverable is high-stakes content where any single agent misses one dimension.
Spinning up five agents named "CEO," "CTO," "designer," etc., to "discuss" your problem. Token cost balloons; quality rarely beats a single careful prompt. The personas are theatre.
If you can't articulate exactly what each agent contributes that the others can't, you don't need them.
The strongest agent system you'll ever ship is "one well-designed agent with great tools and good memory." Multi-agent is a specialty case, not the default. Earn the second agent.
If your single agent works at 80%, fix the tools or context — not by adding agents.
Multi-agent is exciting in demos and frequently underperforms in production. Add agents the way you'd add staff — only when the work demands it.
The boring list. Each of these has cost a team a weekend somewhere. Reading the list is cheaper than discovering them.
Designing AI Agents · The Operator's Library · No. 05. Next: Vol. 06 — The Context Layer — for the RAG/memory infrastructure that makes agents grounded.
Bounded loop. Sharp tools. Living memory. Observability from day one.
— END · OPERATOR'S LIBRARY NO. 05