A plain-English, day-by-day field guide to turning the repetitive work in your business into reliable, auditable, self-improving systems. Built on skills, chains, and approval gates — not on hope.
Read Part I tonight. Run Parts II–V as your 30-day calendar — pick one process, build one skill, chain it, and ship it behind an approval gate. The plan (p. 20) is the map. The workflow catalogue (p. 21–22) is the menu. Educational, with opinions; not a script to copy without thought.
Most people who set out to "automate the business" stall in the same place. They pick a tool, open a blank prompt, and try to describe the whole business in one go. Six prompts later, the thing produces something that's nearly right and impossible to trust. They quit. They tell their friends AI isn't ready.
AI is ready. The mistake is upstream of the tool. Automation works when you start with one repeatable process you already do well, write it down the way you'd hand it to a new hire, and then ask Claude to follow that document — not to invent it.
A business automation is not "Claude doing your job." It is your process, written down well enough that Claude can execute it more reliably than the version that lived in your head.
Start with Chapter 01. The model is simple but most operators get it wrong the first time, and the rest of this guide depends on getting it right.
"Automation" gets used to mean anything that saves a click. That's why the conversation about it is so confused. Before we go further, name the specific shape we're building — a shape that's neither a prompt, nor an RPA script, nor a chatbot.
A business automation, mechanically, is:
01 · A trigger — a specific event the system watches for. A new email. A new row in a database. A schedule. A button click. Without a trigger you have a chat, not an automation.
02 · Context — the data, rules, and prior outputs the model needs to do the job. Templates, brand voice, the CRM, the last invoice. Bring this in deliberately; don't hope.
03 · An action — the structured work Claude actually performs. Read, classify, draft, decide, write back. Always more than one step.
04 · A structured outcome — a clean output the next step (or a human) can consume. A row, a draft, a flagged record. Not a wall of prose.
Strip any one of the four and the system disintegrates. A trigger with no context produces generic output. A great prompt with no trigger is a one-off favour. An action with no structured outcome is a black box no one can audit.
Classical robotic process automation records the exact clicks and keystrokes a human makes, then replays them. It's brittle: a button moves, a modal appears, a field is renamed — and the bot quietly produces garbage. You spend as much time maintaining the script as you saved.
Claude-based automation is different in two ways. It reads the screen, the document, or the API response contextually, the way a person would — labels, state, error messages. And it reasons through unexpected situations rather than failing silently.
RPA replays a recording. A skill executes a process. Recordings break when the world moves; processes don't.
"Reasons through unexpected situations" is a feature and a risk. Without guardrails, the model also reasons through things it shouldn't — like deciding to send the half-drafted email. Ch. 14 is the chapter that prevents that.
Pick one painful, repetitive process in your week. Write its four parts on a single line each: trigger / context / action / outcome. If you can't, the gap you find is the work for Day 1.
Most stalled automation projects skipped a rung. They tried to build an autonomous agent on top of a process that hadn't even been a good prompt yet. The fix is to climb the ladder one layer at a time — each layer earns the right to the next.
Layer 1 → 2: earn this with three or more chats that needed the same context. Stop pasting; promote the context into a Project. You now have a workspace, not a conversation.
Layer 2 → 3: earn this with three or more Project runs that followed the same steps. Stop describing the steps each time; write a skill.md. You now have a process document, not a procedure-in-your-head.
Layer 3 → 4: earn this with at least ten successful skill runs, an eval file that scores them, and a clear approval gate. Now — and only now — wire a trigger to it. You have an agent.
Never automate a process you haven't done as a skill at least ten times by hand. The first ten runs are where you discover the edges, not where you eliminate yourself from the loop.
Chat is for exploration: "is this even possible?" Don't ship anything from a chat.
Project is for repeated work that needs the same context — your brand voice, your customer list, your service catalogue. The single most important file in a Project is CLAUDE.md, a living briefing document.
Skill is the unit of automation. A markdown file Claude follows like a runbook. It has inputs, steps, an output format, and a learnings file that captures every edge case the run revealed.
Agent / Chain is two or more skills wired together with triggers and tool access. This is the layer most people aim at on Day 1 and shouldn't reach until Day 18.
Each layer earns the next. Skip the ladder and the system you build will collapse on its first weird input.
For the process you picked yesterday, name the layer it's currently at. Most readers are honest and say "Chat." That's fine. Today's work is moving it one rung — not four.
Most repetitive work in a business responds beautifully to a skill rebuild. Some doesn't — because the underlying problem isn't process. Here are the four honest situations where this guide alone won't fix what's wrong, and what to do instead.
Move file A to folder B every Tuesday at 9 a.m. No judgement, no variability, no exceptions. This is a cron job and a shell script. Don't waste a language model's reasoning on it.
Use: Zapier, Make, native cron, or a five-line script. Save Claude for the work that involves reading.
If you can't tell a junior hire what a successful run looks like, you can't tell Claude either. The model will produce something. You won't know whether to trust it.
Fix: spend a week scoring three runs by hand against criteria you write down. Only then automate.
Sending money. Posting publicly under the CEO's name. Legal filings. Patient triage. The model is good enough for the draft; it is not yet good enough to be the final sender on consequential actions without review.
Fix: keep the human in the loop forever on these. The approval gate is the product, not a constraint.
Pricing decisions, hiring decisions, which product to build, which client to fire. Claude can brief you brilliantly. It should not decide for you. Trying to make it decide is how good companies make boring choices fast.
Fix: automate the briefing. Reserve the decision.
The all-too-common pattern: pick the shiny new feature, then look around for a process to apply it to. Tail wagging dog.
Fix: pick the painful process first. Then ask whether automation is the right tool. It usually is — but only because you chose well.
You don't know how often this process runs. You don't know how long it takes. You don't know what a good outcome looks like. Six months from now you won't know whether the automation helped.
Fix: log three baselines this week — frequency, time per run, error rate. Then automate.
The honest test: would you be embarrassed to hand this exact process to a new hire? If yes, the problem is the process — and a model can't fix that.
Score the process you picked against the four situations above. If it falls in any of them, swap it out today — you'll save three weeks of frustrating week-four debugging.
Five questions. Score each from 0 (no) to 2 (definitely). Anything above 7/10 is a high-value candidate. Below 5/10, leave it alone — the time you'd spend automating won't come back.
01 · Frequency. Does this run at least weekly? Monthly tasks are rarely worth a skill; weekly tasks almost always are. Daily ones are urgent candidates.
02 · Stability. Have the steps stayed roughly the same for three months? Volatile processes are moving targets; nail them down first.
03 · Inputs. Are the inputs structured (forms, emails, files, rows) and reachable from Claude (Drive, Notion, Gmail, an API)? "It's in someone's head" is a no.
04 · Outputs. Can you describe a "good output" in one sentence and recognise it on sight? If not, write the rubric before the skill.
05 · Reversibility. If Claude does this wrong, can you fix it within an hour? If yes, ship behind an approval gate. If no, push the threshold from 7 to 9, and double the human review.
Weekly+ frequency, stable steps, structured inputs, clear "good," recoverable mistakes. Almost every business has at least three of these hiding in plain sight.
Run three candidate processes through the five-question test. Pick the highest scorer for the rest of the 30-day plan. Park the others — you'll get to them next month.
"Automation" sounds like one thing. It isn't. There are five distinct shapes the work usually takes, each with a different skill template, different failure modes, and different human-review costs. Knowing yours saves weeks.
A stream of variable inputs lands; the skill classifies, prioritises, and routes. Inbound leads. Support tickets. Press requests. Bug reports.
Output shape: a clean record with a category, a priority, an owner, and a recommended next action. Always structured.
Skill takes inputs and produces a draft — email, proposal, post, summary, contract redline. Human edits and ships. The highest-leverage type for most small teams.
Output shape: a near-final draft, plus a 3-line "what I changed and why" note for the editor.
Lead research, competitor monitoring, due diligence, pre-meeting prep. The skill gathers, filters, and synthesises into a one-page brief.
Output shape: bulleted brief with sources, opinion clearly separated from fact, and the two questions you'd ask next.
Meeting transcripts → action items. Long email threads → state. PDFs → key clauses. Weekly metrics → executive summary. Volume goes in, signal comes out.
Output shape: a fixed-template summary — never free-form. Same headings every time so the reader trains on it.
Update the CRM. Post to Slack. Move the Drive file. Schedule the calendar invite. Multi-tool work that used to be tab-juggling. Highest cost-of-being-wrong; require the strongest guardrails.
Output shape: a log of every action taken, every action skipped, and why — for review.
Most real automations are triage → draft → dispatch or research → summarise → draft. Chains are Ch. 11. Pick the dominant type now; chains come later.
Don't build a chain from scratch. Build each link as a standalone skill first.
Triage is cheapest to ship; dispatch is most valuable when it works. Most teams should start with draft — the risk-to-reward ratio is unbeaten.
Classify your chosen process. If it's two types stacked together, that's a chain — and the dominant type tells you which skill to write first. The temptation to start in the middle is strong; resist it.
The single biggest leverage point in this guide. If you can write a runbook a competent new hire could follow, you can write a skill Claude can follow. If you can't write the runbook, no model will rescue you. The map comes first.
Open a blank document. Set a 25-minute timer. Write the process as if you were instructing a new hire who has not seen it before. Cover, in this order:
If section 3 (rules) or 6 (edges) is empty, you don't have a process — you have a habit. That's fine. Spend the rest of the week running the process by hand and writing the rules as you go.
This runbook is a skill in waiting. Sections map directly onto the five sections in Ch. 07. The translation is mostly mechanical.
Write the runbook for your chosen process. 25 minutes, not 25 hours. The gaps you find are the work for the rest of the week, not a reason to give up. The skill you write next week will be at most as good as the runbook you write this week.
A skill is a markdown file Claude follows like a runbook. Five sections, in order, every time. Variation is fine inside the sections; skipping a section is how skills go feral.
01 · Purpose. One sentence. What this skill exists to do, for whom, and what "done" looks like. If you can't say it in one sentence, you have two skills, not one.
02 · Inputs. Explicit, named, with locations. "Read Notion DB 'Invoices' (id: …) where status = Unpaid." Not "look at my invoices."
03 · Steps. Numbered, deterministic. Each step has a single verb at the front: extract, classify, draft, write, log, ask.
04 · Output format. A schema. JSON, a markdown table, a Gmail draft template. Show one filled-in example. Output shape is the contract with the next step.
05 · Guardrails. What to never do. Approval gates. Edge cases and their handling. The "ask me first" list.
A great skill reads like a kitchen recipe written by an obsessive line cook — precise, opinionated, no fluff, with a clear "don't do this" at the end.
Headings, numbered steps, code-fenced examples. Claude treats markdown structure as scaffolding, not decoration — clean markdown makes clean execution.
Translate yesterday's runbook into the five-section skill format. Aim for 30–60 lines total. If it gets to 200, split it — you've found two skills, not written one well. (See Ch. 11.)
The most common failure in chained automation isn't reasoning. It's the seam between skills. Skill A emits a wall of prose; Skill B can't find the field it needs. The fix is to treat every skill's output as a contract.
{ lead: "Acme", country: "UK", staff: 50, priority: "mid", best_case_study: "Brigden Steel", next_action: "Follow-up Mon", confidence: 0.78, sources: [...] }
Use explicit file paths between steps. Skill A writes to outputs/leads/acme.json. Skill B reads it. No "remember the lead we discussed earlier" — that's a context bet, and it loses.
Validate at the boundary. Each skill begins by asserting its inputs exist and are well-formed. "Before proceeding to step 3, confirm that the file from step 2 exists and contains at least one row." One sentence in the skill prevents an hour of silent failure.
Fail loudly, never silently. If a required field is missing, the skill stops and asks. The worst outcome in automation is a confident wrong answer flowing into the next link.
Chains rarely break inside a skill. They break in the seams — at the handoff. That's where to put validation, not at the top of each skill.
Write the JSON schema for your skill's output before writing the steps. Five fields, not twenty. Then write one fully populated example below it. Both belong inside the skill file.
A skill at week one is mediocre. A skill at week six, with five to ten production runs and a maintained learnings file, is better than most of your team at the same task. The thing that compounds is not the model — it is the file alongside it.
Every skill folder has a companion learnings.md. After each run, three things happen:
A model can't remember the lesson from last Tuesday. The file can. The learnings file is the institutional memory the skill never had. Each promotion makes the skill smarter; each prune keeps it under 500 lines and readable.
Don't auto-promote. Models will dutifully turn every minor surprise into a rule, and your skill bloats. The promotion step belongs to a human. Quarterly is plenty.
Create learnings.md next to your skill file. Make it empty for now. Tomorrow, after the first run, write the first note. Don't wait for a perfect format — date, one sentence, that's the whole template.
"Did the skill work?" is the wrong question. The right one is "did it work on the inputs I expect to see?" Without an eval file, you'll discover the answer six weeks later, in production, when the bad output reaches a real customer.
A small JSON file living next to skill.md with three to ten test cases. Each case has an input, an expected output shape, and a pass/fail criterion. You run the eval whenever you change the skill — exactly like unit tests for code.
Three cases is the minimum: one happy path, one edge, one hostile. The hostile case — partial payment, missing field, ambiguous input — is the one that finds the real bugs.
Good: small, deterministic, structured input → structured expected output. Easy to diff. Easy to debug.
Bad: "the output should sound friendly." Sound is judgement, and judgement-eval bloats fast. Instead, assert tone == friendly as a field in the output, and let the skill self-declare. The skill is more honest than your gut.
A skill without an eval is a habit you got lucky with. An eval turns the skill into a thing you can ship and stand behind.
Evals are not gating bureaucracy. They are the floor that lets you change the skill confidently. Without one, every edit is a roll of the dice; with one, every edit is a small experiment with a known control.
Write three eval cases for your skill. Happy, edge, hostile. Five minutes each. Run them. Note what fails. The first failures are the most valuable data point of the whole 30 days.
A single great skill saves time. Three to seven, chained, build a system. Most production workflows live in this range — fewer than three usually means a skill is doing too much; more than seven means the workflow needs restructuring, not more skills.
Good — Lead to proposal: extract → research → match-case → draft → review. Clear boundary at each seam.
Good — Weekly review: pull metrics → summarise → flag anomalies → post draft to Slack for CEO comment.
Good — Inbox triage: classify → route → draft reply (for routable types) → save as draft.
Bad — Mega-skill: "Handle every new lead, top to bottom." One file. Untestable. Quietly breaks.
Bad — Implicit handoff: three skills that share a chat history instead of writing files. Reliable for two runs; fragile for the third.
Bad — Approval at every step: the human re-reads the same context six times. They stop reading by step 3. Now nothing is approved.
Sketch your chain on paper before you write the second skill. Boxes are skills; arrows are file handoffs; mark the one approval gate. If the picture has more than seven boxes, you're solving the wrong problem.
Once a chain is more than three skills, the shape of the chain starts to matter as much as the skills inside it. There are five recognised shapes. Knowing which one your work needs prevents months of fighting the wrong architecture.
The default. Each skill reads the previous skill's output and produces the next. Linear, easy to debug. Most weekly business processes fit here.
Use when: order matters and each step depends on the last. Avoid when: steps are independent and could run in parallel.
A top-level skill reads the situation, then decides which sub-skill to call. Good for triage problems where the next step depends on what kind of input came in.
Use when: inputs are heterogeneous. Avoid when: the routing logic is simple enough to be a switch statement.
One input becomes N parallel runs (one per row, one per source) which join into a single summary. Excellent for research across multiple companies, or summarising a week of meetings.
Use when: items are independent and the merge is structured. Avoid when: items have hidden dependencies.
Several agents with distinct roles (researcher, writer, editor) work together with shared memory. Powerful for content and long-form deliverables. Expensive in tokens; slow to debug.
Use when: the deliverable genuinely needs perspectives the same model in one pass can't provide. Avoid when: one skill with three steps does the job.
The whole chain runs on a cron or webhook trigger; results land in a queue for review later, or just commit if validation passes. The shape of a mature automation.
Use when: you have ten successful supervised runs and a passing eval. Avoid when: any of those preconditions is missing. Ch. 15.
(1) Are the inputs uniform or varied? (2) Are the steps order-dependent or independent? (3) Do I need a human at every stage, or can I batch reviews?
Answer those and the pattern picks itself. Don't pick by what sounds advanced.
Sequential is boring and works. Agent teams are exciting and frequently underperform a well-written sequential chain. Pick by the work, not by the demo you saw.
Pick the pattern that fits your chain. 80% of teams should pick sequential. Write the pattern name at the top of your chain's overview doc — it's the one design decision worth re-reading at the start of every weekly review.
A skill that can only read what you paste into it is a fancy autocomplete. A skill that can read Gmail, your CRM, your Drive, and your billing system is automation. The bridge between the model and your stack is the Model Context Protocol — MCP.
MCP is an open protocol that lets the model call tools that live outside it. Each tool is a small server that exposes named actions — gmail.read, notion.upsert, slack.post. The model asks; the server answers. Importantly, the model doesn't need to know how the tool works — only what it can do.
Most teams need five connectors, not fifteen. Adding more is rarely the bottleneck; making the five work reliably almost always is.
MCP is not magic. It doesn't grant the model new judgement. It grants the model new reach. The judgement still has to live in the skill. A perfectly connected agent with a sloppy skill will sloppily touch ten systems instead of one.
Every connector you enable expands what a model can read and write. Scope tokens to the minimum needed. Use a credential vault, not raw API keys in prompts. Audit logs are not optional. If your CFO can't see what the agent touched last month, you don't have automation — you have a liability.
This sequence is dull and correct. Skipping it is fast and risky. Pick.
Enable the three connectors your chosen workflow needs. Verify read access by asking Claude to summarise something specific that lives in each system. If any returns the wrong thing, fix the scope before you move on.
The single most important page in this guide. Every automation that sends, posts, pays, or commits to a customer needs a human in the loop until you have evidence it doesn't. Designing the loop well is the difference between "shipping" and "praying."
01 · Pre-action approval. The default. Skill drafts, human approves, action fires. Use for emails, posts, contracts, payments, anything externally visible.
02 · Spot-check sampling. The skill acts on every run; a human reviews 1 in N runs after the fact. Use for high-volume, low-risk work — internal classification, log triage.
03 · Exception-only review. The skill self-flags ambiguity, low confidence, or rule conflicts; only flagged runs reach a human. Use after enough supervised runs that you know what flagging looks like.
An approval gate is a "stop until I say go." A circuit breaker is a "stop automatically when something looks wrong." Three to build, no exceptions:
Whatever the gate, the human review surface needs three things visible at once: what the skill did (a one-line summary), what's about to happen (the draft), and why the skill chose it (the rule that fired). Approving blind is not approval.
Removing the human in the loop is a bet that the model never drifts. That is a bet. Make it deliberately, with evidence.
For your chain, pick the gate type. Implement one circuit breaker today — the rate limit is the easiest and most universally useful. The other two can wait until you've seen the first real run go wrong.
Headless is the goal — the skill runs on its trigger without a human in the room. It is also the most dangerous transition in the build. Fully autonomous operation means fully autonomous mistakes. Walk into it with eyes open.
"Headless" is not a synonym for "set and forget." It means "scheduled and observed." The on-call burden is real, especially in the first month after the switch.
Phase 1 — supervised manual. You trigger every run. You review every output. Days 11–17.
Phase 2 — supervised scheduled. The cron triggers; the skill drafts; you review and approve. Days 18–24.
Phase 3 — headless with exception review. The cron triggers; the skill acts; only flagged or out-of-bounds runs reach you. Day 25 onward, if Phase 2 was clean.
Headless without monitoring is not autonomy. It's negligence with extra steps.
Score your skill against the five preconditions. Honest grade. Fewer than four greens, you stay in Phase 2 another week. That's the right answer, not a failure.
Most automations that fail don't fall over loudly. They degrade quietly — outputs drift, edge cases pile up, costs creep, the team stops trusting the drafts. Six patterns to recognise early, and the fix for each.
Every edge case became a new rule. Now the file contradicts itself and Claude follows the wrong half. Symptom: outputs get worse the more rules you add.
Fix: prune. Split. Promote a "general principle" rule above five specific ones. Quarterly review of learnings.md.
Works in dev; fails in chains. Skill B can't find what Skill A "discussed earlier" because there was no earlier.
Fix: every handoff is a file with a known path and a known schema. Boring; reliable.
The reviewer rubber-stamps because there are 40 drafts a day. A bad one slips through. Now the system has the worst of both worlds.
Fix: cut to one gate per chain, near the end. Pair with exception-only review and aggressive flagging.
Same skill; different model; subtly different behaviour. Outputs change shape; downstream chain misreads them.
Fix: rerun the eval whenever the model changes. Treat model swaps the way you treat code releases.
Usually one of three: inputs got longer (someone dumped a 200-page PDF), a step started looping silently, or the skill got promoted to a larger model when the mid-tier model was fine.
Fix: log tokens per run. Set a budget alarm. Default to the cheapest model that passes the eval.
The loop only compounds if you close it. An unread learnings file is just a graveyard.
Fix: a 30-minute quarterly review on the calendar. Same date every quarter. Promote three, prune three.
A great automation needs maintenance the way a great garden does. Quiet, scheduled, unglamorous. Six hours a quarter is the price.
Run this page like a checklist. Most "the model got worse" reports are one of these six in disguise. Diagnose before you reach for a new model or a new tool.
One process. One skill becomes one chain. One approval gate. One running system at the end of the month. Don't skip days; don't double up. The constraint is the point.
Cut scope, not days. Drop the chain back to a single skill. Drop the cron back to manual. The plan beats one weekend of catching up.
A first automation realistically gives back 2–6 hours a week. The second, 4–10. The third, more — because each one re-uses skills you already wrote.
The real prize is not the hours. It's that the work that used to depend on someone remembering to do it now happens whether or not anyone remembered.
One running system in thirty days. Worth more than ten half-built ones in ninety.
Four automations that fit the framework. Each is one skill or a short chain. Estimated weekly time back is honest — measured by teams running these in production, with review time deducted. Read these as starting points, not templates to copy without thought.
New email lands. The skill extracts name, company, ask, budget signals; runs a quick web research pass; matches the best case study from the Notion library; drafts a proposal using your template. Draft only.
stack · Gmail · Notion (cases) · web search · Drive
Hourly: classify new emails as reply-needed, FYI, junk. For reply-needed, generate a draft. For FYI, summarise into a daily digest. For junk, archive (after 30 days of sampling proves it's safe).
stack · Gmail · Slack (digest)
30 minutes before any external meeting on the calendar, generate a 1-page brief: company facts, recent news, the attendee's role and prior interactions, three suggested opening questions.
stack · Calendar · CRM · web search · Slack DM
Whenever a blog post is published, generate platform-appropriate variants (LinkedIn long, X thread, newsletter teaser). Save as drafts in the relevant scheduling tool; do not post.
stack · CMS · Buffer/Hypefury · brand voice doc
Most "AI saves time" claims include the time it took to ship the skill. We don't. The numbers above are post-launch weekly returns, after review.
If your inbox volume is < 30 emails/day, skip workflow 02 — the ROI doesn't justify the build. Pick by volume, not by which one sounds coolest.
Four more automations, this time inside the back office. Higher stakes than the marketing set — money, payroll, employee data — so every one here ships with a pre-action gate by default. Loosen later, never sooner.
Mondays: scan the Invoices DB. For each unpaid, compute days overdue, pick tone (friendly → firmer → direct), draft a personalised follow-up email, save as Gmail draft. Send thank-yous for new "paid" rows.
stack · Notion (invoices) · Gmail · tone ladder doc
First of the month: pull revenue, AR aging, expense categories from accounting; pull traffic and ad spend; produce a 1-page narrative + 5 charts; flag anything > 15% off forecast. Draft for the owner to review and send.
stack · QuickBooks (read) · Analytics · ads APIs
New signed-MSA email arrives → extract counterparty, term, value; populate the playbook checklist; flag any non-standard clauses; place the draft signature packet in Drive for legal review. Above $X, route to human review only.
stack · Gmail · Drive · Docusign · clause library
Every new ticket: classify (billing / bug / how-to / urgent), set priority, route to the right queue, and for "how-to" only, draft an answer using the help-center as source. Human approves before send for medium/high priority.
stack · Helpdesk · Help-center MCP · Slack alerts
Workflow 05 is where most small businesses recover the cost of Claude in their first month. The ROI of one chased invoice exceeds the annual subscription many times over.
Patterns that quietly waste the next six months of your team's time. They look like work; they aren't. None of them are exotic — they're the same ten things every team gets wrong on the first build, including ours.
Most of these are unglamorous. So is most of the work that pays back. The flashy mistake to avoid is choosing the unflashy mistake to repeat.
Whenever you're about to start a new automation, scan the ten. If any apply, fix that first. The page costs five minutes to re-read; the mistakes cost weeks.
Default to a capable mid-tier model for skill execution; reach for a larger model when an eval case fails repeatedly and the failure is reasoning-bound, not prompt-bound. Smaller models are excellent for cheap, structured front-of-chain steps — classification, extraction.
For a single workflow run a few times a week, the token bill is usually modest. Check current pricing before forecasting. The dominant cost is large inputs (long PDFs, big email digests), not the number of runs.
Yes for most of Parts I–IV. Skills are markdown; Projects are no-code; MCP connectors are toggles. The places that benefit from a developer are cron triggers, custom MCP servers, and observability dashboards.
Scope every connector to least-privilege. Use credential vaults, not raw keys. Keep PII out of skill files. Maintain an audit log of every action. For regulated industries, the human-in-the-loop pattern is not optional, it's the compliance story.
Re-run the eval monthly. Spot-check five real outputs a week. Promote three notes and prune three from learnings.md quarterly. Sounds light; works.
Yes — for the deterministic glue (file moved, row added, webhook fired). The judgement steps belong to a skill. The two layers are complementary, not competitive. Use the right tool for each segment.
Three numbers. Time-back per week vs. baseline; flag-rate trend (should be flat or falling); cost per useful output (should fall after week 3). If any of the three drifts wrong, look at Ch. 16.
That's what gates and breakers are for (Ch. 14). The question to design around isn't "will it ever be wrong?" — it'll be wrong. The question is "when it is wrong, will we catch it before it leaves the building?"
Don't pitch the platform. Pitch the first weekly time-back number, measured. People who see four hours back per week become advocates faster than any deck.
Once you have three running workflows, the work shifts from "build the skill" to "design the system of skills." That's the subject of the next field guide. Same series. Same shelf.
Every honest answer here is shorter than the question. That's not laziness — it's the shape of work that actually compounds.
The starting point. Marta runs a 3-person brand studio. Inbound enquiries arrived faster than she could write proposals; she lost two leads in March because she didn't reply within a week. Baseline: ~6 hours/week on proposals, average 4 days response time.
Days 1–7. Picked "lead → proposal draft" as the process. Scored 8/10 on the candidate test (mostly the recoverability question — proposals are drafts, not sends). Archetype: draft. Wrote the runbook in 40 minutes; found two big gaps in her own process — no case-study tagging, no tone guide.
Days 8–14. Built the studio's first skill.md: extract → research → match → draft. Output schema: a Notion page using her existing proposal template. The eval revealed the skill was over-citing one case study (most-recent ≠ best-fit). Fixed.
Days 15–21. Wired Gmail (read) and Notion (read/write). Added a Slack DM as the approval surface — one-tap "approve & send" or "edit first." First three real leads went through end-to-end.
Days 22–28. Phase 2 supervised scheduled. Caught one wrong-tone draft (skill matched a casual case study to a formal RFP). Added a "formality detector" note to learnings.md; promoted to a rule on day 27.
Outcome. Proposal time: 6 hrs/week → 1.8 hrs/week. Average response: 4 days → 8 hours. Won the next two enquiries that would have ghosted. Cost: $34/month in tokens.
The starting point. Every Monday Dan compiled a 9-tab spreadsheet, wrote a narrative summary, and shipped it to the leadership team. Nobody read past the first paragraph. Baseline: 5 hours of his own time, plus 30 minutes from three other people pulling data.
Days 1–7. Picked "weekly leadership brief" — high frequency, painfully stable, structured inputs, fuzzy "good output." Spent days 4–7 actually defining what "good" meant: a 1-page brief, 5 charts, anomalies flagged with a hypothesised cause. Wrote the rubric.
Days 8–14. Built three skills: pull-metrics, summarise-narratively, flag-anomalies. Defined the schema for the handoffs (rows with fields, not paragraphs). The eval caught a recurring failure — the skill flagged seasonal patterns as anomalies. Added a "compare-to-prior-year" rule.
Days 15–21. Connected the analytics, billing, and product APIs. Read-only. The first end-to-end run produced a brief Dan would have written himself — at minute 14 of a 14-minute job, not hour 5.
Days 22–28. Approval gate: the brief lands as a Slack thread for leadership Monday at 8 a.m.; Dan approves to publish to the company wiki. After two weeks of clean runs, Dan went exception-only — only anomaly-flag runs require his review.
Outcome. Total time: 6.5 hrs/week → 25 minutes/week. Read-through rate (measured by reactions): 1 of 8 leaders → 7 of 8. The brief stopped being a deliverable and started being a meeting input.
The unexpected win wasn't the time back. It was that the work got more important — because more people read it.
Once your first chain has run cleanly for a month, the second is roughly twice as fast to build. The skills you wrote re-use. The instinct sharpens. The catalogue (p. 21–22) is a starter; the longer list belongs to the next field guide in this series.
A business with three running automations runs differently from one with zero. The work doesn't vanish; it climbs up a layer. The team's days fill with the things they're actually good at.
Business Process Automation with Claude — A Field Guide.
Edition 2026 · 01. Volume 15 of the Field Guide series.
Set in Source Serif 4 (display) and Outfit (text). Monospace specimens in JetBrains Mono. Paper tone: warm cream. Accent: steel blue.
Built for founders and operators who already know the work — and want the boring middle of it to run while they sleep.
Pick one process. Write the runbook. Ship the skill. Watch the work change shape.
— END · FIELD GUIDE NO. 15