A field guide · October 2026

The Operator’s Handbook

Running a one-person machine with AI agents, day by day
by Mat Siems
Part I

The Operator's Chair

What the job is when agents do the typing.

Chapter 1 · Part I

One Person, Many Hands

This is a handbook for a particular kind of worker: one person, no staff, and a fleet of AI agents that will cheerfully do almost anything you ask of them, at any hour, in parallel. You might be a freelancer, a founder with no co-founder, a consultant, a writer with a codebase, or simply someone who noticed that the number of things they can get done in a day has stopped being limited by how fast they type. You are, whether you like the word or not, an operator.

The word matters because it names the job honestly. You are not a manager, because nobody reports to you who can take responsibility for anything. You are not quite a maker any more, because most of the making is done by something else. You run a small machine. The machine is made of agents, tools, notes, a list of projects and, most importantly, your own working day. When the machine runs well, a single person ships the output of a small team. When it runs badly, a single person spends the whole day supervising chaos and goes to bed with nothing finished.

What separates the two is rarely the tools. Everyone has roughly the same agents now. The difference is the operating system: the habits that decide what gets started, how it is described, how it is checked, when it ships and what gets written down afterwards. Those habits are what this book is about. It is not a business book. It will not tell you what to sell or how to price it. It is about the inside of one person's working day, and how to make that day produce finished things rather than busy feelings.

The agents supply the hands. You supply the order in which they move.

The book runs in ten parts. It starts with the job itself, then walks through the operating day, the writing of briefs, memory and notes, the project register and the delegation of threads, the reviewing of work, release discipline, the weekly and quarterly loops, the upkeep of tools and the handling of mistakes. It ends where everything in it points: the operator's real job is deciding, not doing. Each chapter teaches one thing you can try this week. None of them needs new software. Most need a text file and a little nerve.

A word on tone. This is written for people who are already a bit tired of being told that everything has changed. Plenty has. Plenty has not. People still need to know what they are trying to achieve, still need to check whether they achieved it, and still need to stop working at some point and eat dinner. Agents make the first two more important, not less, and they make the third harder to remember.

So start here, with a small exercise. Write down, in one sentence each, the three things your agents did for you last week that you would not have done yourself. Then write down the one thing you did that no agent could have done. That second list is shorter. It is also your job.

The operator runs the order, not the hands Start what gets started Describe how it is briefed Check how it is reviewed Ship when it goes out Record what is written down The operator one person sets the order the operating system The hands agents tools notes projects Everyone has the same agents. The habits are the difference.
Fig 1 · One Person, Many Hands. Five habits ring the operator, who sets their order while agents and tools do the work.
Chapter 2 · Part I

The Day Is the Machine

Most people who run agents think of their system as a set of tools: this assistant, that terminal, these connectors, those scripts. The tools matter, but they are not the machine. The machine is your working day. It is the order in which things happen between waking and stopping, and it determines what the tools are pointed at. Two operators with identical tools and different days will produce wildly different weeks.

Think of the day as having three parts. In the morning you review: you read what happened while you were away, decide what matters now, and choose a small number of outcomes for the day. In the middle you run sessions: bounded stretches of time in which you brief agents, watch their work, review what comes back, and either ship it or send it round again. At the end you close: you record what finished, what did not and what tomorrow should start with, and then you stop. Review, sessions, close. Everything else in the operating day is a variation on that arc.

This sounds obvious, and it is, which is why so few people actually do it. The default day for an operator with agents is reactive. You open the laptop, notice that a background job finished, read it, fix something small, start a new thread because an idea occurred to you, answer a message, check another thread, and look up to find it is four in the afternoon and you have touched eleven things and finished none. Agents make this worse, because each one generates more output to look at. The machine without a shape becomes a machine for generating notifications.

A day without a shape is shaped by whatever shouts loudest.

The fix is to treat the shape of the day as a design decision rather than an accident. Decide when the review happens and how long it takes. Decide how many sessions the day can hold and roughly what size they are. Decide when the close happens and what it produces. Write these down somewhere you will see them. They will be wrong at first, and you will adjust them, but a wrong plan you can adjust beats no plan at all, because without one there is nothing to adjust.

It helps to notice that agents do not get tired and you do. The machine can run all night; the operator cannot. So the day is the part of the system that has hard limits, and hard limits are where design pays off. A database designer worries about the slowest query. An operator should worry about the scarcest hour, which is usually the first good one of the morning, and should spend it on deciding rather than on fiddling.

Try it this week. For five working days, write three times on a card: when review ended, when the last session ended, and when you closed. Do not try to improve anything. Just look. Most people discover that the close never happens at all, and that the review quietly merges into the first session. That is not a moral failing. It is a machine with a missing part. Fit the part.

Same tools, two different days Designed has a shape Review decide Session one outcome Session one outcome Session one outcome Close then stop Reactive no shape no close 11 things touched · 0 finished · notifications set the order 08:00 12:00 16:00 18:00 A day without a shape is shaped by whatever shouts loudest. Five-day card review ended · last session ended · closed
Fig 2 · The Day Is the Machine. A designed day of review, sessions and close beside a reactive day that finishes nothing.
Chapter 3 · Part I

Doing Versus Deciding

For most of working history, doing and deciding came bundled. If you wanted a report written, you decided what it should say and then you wrote it, and the writing took so long that it felt like the real work. The deciding hid inside the doing. You chose the structure while drafting the first paragraph, changed your mind about the conclusion halfway through, and the finished report was the record of all those small decisions.

Agents unbundle the two. The doing, which used to take hours, now takes minutes and can happen without you. The deciding does not shrink at all. If anything it grows, because the agent will do exactly what you decided, and if you decided vaguely you get a vague result very quickly. The operator's day therefore tilts. Less of it is spent making things. More of it is spent choosing what to make, describing it precisely, judging what came back, and choosing again.

This is uncomfortable for anyone whose identity was built on doing. Making things is satisfying in a way that deciding rarely is. You can point to a page you wrote or a feature you built. It is harder to point to a good decision, because good decisions mostly look like nothing happening: a project not started, a feature not added, a bad pull request not merged. The temptation, when agents take away the doing, is to grab some of it back for the comfort, to rewrite the agent's draft in your own words when the draft was fine, or to fix the bug yourself because waiting felt idle.

Resist the temptation, but not dogmatically. There is a simple test for any piece of work in front of you: does this need to be done by me specifically? Some things do. A conversation with someone who trusts you. A judgement about quality only you can make. A piece of writing whose whole value is that it is yours. Everything else is a candidate for delegation, and the question becomes how to brief it well rather than whether to do it yourself.

Doing feels like progress. Deciding is progress.

Notice also that deciding is not the same as thinking about things. Operators can lose whole mornings to deliberation that never lands. A decision is a sentence with a verb in it: we will ship the small version on Thursday; we will drop the second feature; we will rewrite the brief and run it again. If your thinking does not end in a sentence like that, it was not deciding. It was weather.

So this week, keep a tally. Each time you sit down to do a piece of work yourself, ask the test question first, and mark which way it went. Do not change your behaviour; just count. By Friday you will know what fraction of your day is doing that only you can do, and what fraction is doing you have not yet learned to hand over. The second number is the size of your opportunity. It is usually larger than you hoped, and also larger than you feared.

The test question A piece of work in front of you Needs me, specifically? yes no Only you a conversation built on trust a judgement only you make writing whose value is yours Hand it to an agent the question becomes how to brief it well, not whether to do it yourself Decide a sentence with a verb we will ship the small version on Thursday Doing feels like progress. Deciding is progress.
Fig 3 · Doing Versus Deciding. One test question routes each piece of work to you or to an agent, ending in a decision.
Chapter 4 · Part I

The Three Operator Questions

Every piece of work an operator starts should survive three questions. What do we want? Why now? How will I know? They take about a minute to answer, and the minute is the best-spent minute of the day, because work that fails them tends to consume hours before anyone notices it was never worth doing.

The first question, what do we want, sounds trivial and is not. Most failed agent runs fail here. The operator had a feeling, typed it into a prompt, and received something that matched the words but not the feeling. The fix is to answer in terms of a result rather than an activity. Not look at the onboarding flow but a new user can sign up and reach their first saved item without help. Activities can be performed indefinitely. Results either exist or they do not.

The second question, why now, is the operator's defence against the infinite. Agents make it cheap to start things, which means that the backlog of plausible ideas grows faster than any person can review the output. Why now forces you to rank. Perhaps something is broken and costing you every day. Perhaps a door will close next week. Perhaps it is simply the most useful thing on the list. All of those are fine answers. Because I thought of it this morning is not, although it is the most common one.

The third question, how will I know, is the one that makes delegation possible. If you cannot say how you would recognise success, you cannot ask an agent to achieve it, and you certainly cannot review whether it did. A good answer names a check: the tests pass and the new one fails without the change; the page loads on a phone; the summary fits on one screen and mentions all four suppliers. A bad answer is I'll know it when I see it, which means you will see something and then argue with yourself about it.

If you cannot say how you would know, you are not ready to ask.

The questions narrow as they go, which is why they are worth asking in order. The first one is broad and generous; it lets in anything you might want. The second filters by timing and cost. The third filters down to the work you can actually verify. What emerges at the bottom is small, specific and checkable, which is exactly the shape that agents handle best and operators review fastest.

You do not need a form for this. A line at the top of each brief will do. Some operators write the three answers as the first three sentences of every task, before any instruction. It looks a little ceremonial at first. After a fortnight you stop noticing it, and you start noticing instead how much less work you throw away. The questions do not make you cleverer. They simply stop you from being busy on the wrong thing, which, in a day full of eager agents, is most of the battle.

Three questions, asked in order What do we want? a result, not an activity Why now? rank it against the list How will I know? name the check Ready to brief small · specific · checkable FAILS WHEN THE ANSWER IS an activity I thought of it today I will know it when I see it look at onboarding novelty, not need an argument later If you cannot say how you would know, you are not ready to ask.
Fig 4 · The Three Operator Questions. Three questions narrow work down to something small, specific and checkable.
Chapter 5 · Part I

Throughput Is Not the Point

The first thing agents give you is volume. Ask for ten variations and you get ten. Ask for a feature and you get the feature, its tests, a migration and a tidy summary. Start five threads before lunch and you will have five sets of results by teatime. It is intoxicating, and like most intoxicating things it is a poor guide to whether you are getting anywhere.

Output is what the machine produces. Outcome is what changes in the world because of it. An operator can generate enormous output with no outcome at all: drafts that are never published, branches that are never merged, prototypes that are never shown to anyone. The dashboard looks busy and the work looks impressive and nothing has moved. This is the busy trap, and agents dig it deeper, because the cost of producing one more thing has fallen almost to zero while the cost of finishing one more thing has not fallen at all.

Finishing is expensive because it involves you. Someone has to review the work, decide it is good enough, push it into the world, and deal with what happens next. Each of those steps draws on your attention, which is fixed. So an operator who starts more than they can finish is not increasing their throughput. They are building a queue, and queues have a habit of turning into guilt.

Started is a cost. Finished is a result.

It helps to picture your work on two axes. One axis is output: how much was produced. The other is outcome: how much of it reached someone and made a difference. Most operators, when they first get agents, slide along the output axis and stay low on the outcome one. The quadrant you want is the one where output is modest and outcome is high: fewer things, finished. It feels less productive. It is enormously more productive, as anyone who has ever shipped one thing instead of starting four can tell you.

The practical habit is to measure finishes rather than starts. At the end of each day, count only what crossed the line: merged, published, sent, delivered, decided. Do not count threads opened or drafts created. If the count is zero, that is information, not a scolding. Look at why. Usually the answer is that too many things were started and none was given the review attention it needed to finish.

There is a secondary benefit. When you count finishes, you begin to brief differently. You stop asking agents for exploratory sprawl and start asking for the smallest thing that could be finished today. You split large jobs into pieces that can each ship on their own. You notice that a half-finished large thing is worth less than a finished small thing, because the small thing can be used and the large thing can only be admired.

Try it for one week: a single number on a card, finishes per day. Ignore everything else the machine tells you about how busy it has been. The machine is always busy. That is its nature. Your nature is to decide which part of that busyness becomes real.

Output is not outcome Outcome reached someone low Output: how much was produced low Few, finished the quadrant you want merged, published, sent feels less productive is far more productive Many, finished expensive every finish needs your review your attention is fixed rare for one person Quiet a rest day, or a stall nothing started nothing moved The busy trap where agents push you drafts never published branches never merged prototypes never shown count finishes per day, not threads opened
Fig 5 · Throughput Is Not the Point. Output against outcome: aim for few things finished, not the busy trap of many started.
Chapter 6 · Part I

Direct, No Preamble

Agents do not need to be warmed up. They do not need to be told that you hope they are well, that this is a bit of a strange request, or that you have been thinking about something for a while. They need to know what you want, what they should know to do it, and how you will judge the result. Everything else is preamble, and preamble costs more than it looks.

It costs, first, in clarity. A brief that opens with three sentences of context-setting makes the actual instruction harder to find, for the agent and for you when you reread it later. Agents give weight to everything in front of them, so a meandering opening can tilt the work in directions you did not intend. If you mention in passing that you are worried about performance, do not be surprised when a simple text change arrives wrapped in a caching layer.

It costs, second, in your own thinking. Preamble is often what we write while we are still working out what we want. That is a fine thing to do, but do it in a notebook, not in the brief. Draft freely, then delete everything above the first sentence that contains a verb and a result. What remains is usually the brief you meant to write.

Say the thing. Then say how you will know it is done. Then stop.

The same applies in the other direction. Ask agents to be direct with you. A good result leads with the outcome and the evidence, not a narrative of the journey. Done: the importer now handles empty rows; the new test fails on the old code and passes on the new; full suite green is worth more than three paragraphs describing the investigation. You can always ask for the story. You should not have to dig for the verdict.

This is not about being curt. It is about respecting the scarcest thing in the system, which is your attention. Every unnecessary sentence an agent writes is a sentence you have to read before you can decide what to do next. Multiply that by twenty threads a day and preamble becomes a tax on your whole operation. Many operators put a standing instruction in their project memory to this effect: lead with the result, then the evidence, then any open questions, and keep the summary short. It is one of the highest-return lines you can write.

Directness also makes disagreement cheaper. If an agent thinks your approach is wrong, you want it to say so in the first line, not to bury the concern in the fourth paragraph after doing the work anyway. Ask for that explicitly. Plain speech in both directions turns the collaboration from a polite exchange of documents into something closer to a working conversation.

So, this week, look at your last ten briefs and cut the first sentence of each. Then see if the brief still makes sense. In most cases it will make more sense. The preamble was for you. The agent never needed it, and, if you are honest, neither did you.

Cut the warm-up, in both directions With preamble Direct Brief Report Hope you are well! This is a bit of a strange one I have been thinking for a while also worried about speed… Please fix the label text. Goal: fix the label text Done when: it reads Export Do not: touch the layout then stop First I looked at the… Then I explored how… Interestingly, the… After some time… …so it should work now. Done: empty rows handled Evidence: test fails, passes Full suite: green Open questions: none Say the thing. Then say how you will know it is done. Then stop.
Fig 6 · Direct, No Preamble. Preamble-heavy briefs and reports compared with direct ones that lead with the result.
Chapter 7 · Part I

Bias to Action, Bound by Evidence

Operators who do well with agents tend to share a temperament: they would rather try something than discuss it. When a question comes up that an experiment could answer, they run the experiment. When a draft is needed, they get one made and then argue with the draft rather than with an empty page. Agents reward this temperament handsomely, because experiments that used to take a day now take ten minutes, and there is no longer much excuse for long deliberation about things that could simply be tested.

But bias to action has a failure mode, and agents reward that too. The operator who acts fast without checking ends up with a lot of things that look done and are not. A migration that ran but quietly skipped records. A page that renders beautifully on a laptop and falls apart on a phone. A summary that is confident and wrong. Each one is cheap to produce and expensive to discover later, usually at the worst possible moment and often by someone else.

The answer is not to slow down. It is to bind your speed to evidence. Act as fast as you like, provided that each action produces something you can check, and that you actually check it before building on it. That is the overlap that matters: not caution, not haste, but quick moves that each leave behind a piece of proof. The operator lives in that overlap.

Move as fast as your evidence can keep up with.

In practice this means building the check into the action. When you brief an agent, include the verification step: run the tests, load the page, count the rows before and after, compare the summary to the source. When you take an action yourself, decide in advance what you will look at afterwards. If there is nothing to look at, either find something or acknowledge that you are guessing and treat the result as provisional.

It also means being honest about which actions are reversible. Trying a new layout on a branch is cheap to undo; act freely. Sending an email to everyone on a list is not; slow down and look twice. A useful habit is to label each action, silently, as either a sketch or a commitment. Sketches can be fast and sloppy. Commitments need evidence. Most of the trouble operators get into comes from treating a commitment as if it were a sketch because the agent made it feel so easy.

There is a pleasure in this way of working that is easy to miss. When every move leaves evidence, you stop carrying anxiety about whether things are really done. You know, because you looked. That frees attention for the next decision, which in turn makes you faster. Evidence is not a brake on bias to action. It is the thing that lets you keep your foot down.

This week, pick one task you would normally deliberate over and run it as an experiment instead. Before you start, write the one check that would tell you it worked. Then do it, check it, and see how long the whole thing took. It will be shorter than the deliberation would have been.

Move as fast as your evidence keeps up Act fast no checking skipped records breaks on a phone confident and wrong Check no acting long deliberation nothing tested slow to learn Operator quick moves with proof Sketch: reversible, act freely Commitment: needs evidence
Fig 7 · Bias to Action, Bound by Evidence. The operator works where fast action overlaps with checking the evidence.
Chapter 8 · Part I

Artifact First

There is a simple rule that saves operators a remarkable amount of time: when in doubt, make the thing. Do not describe the dashboard; get a rough dashboard built. Do not debate the structure of the report; have a draft written in two structures and read both. Do not imagine how the onboarding email will feel; produce it and read it as if you had received it. An artifact on the table ends more arguments than any amount of discussion about one.

This used to be expensive advice. Making a prototype took days, so it was sensible to talk things through first and build only once you were fairly sure. Agents have inverted the economics. A rough version of almost anything now costs less than the meeting you would have had about it, even if the meeting is only with yourself. So the order of operations changes. Instead of think, decide, build, it becomes build roughly, look, decide, build properly.

The reason this works is that people, including you, are much better at reacting than at imagining. Shown a page, you know within seconds that the headline is too long and the button is in the wrong place. Asked to imagine the page, you can spend an hour and still miss both. An artifact converts vague preferences into specific objections, and specific objections are things an agent can act on.

You cannot review an intention. You can review an artifact.

There are two disciplines that keep artifact-first from turning into artifact sprawl. The first is to label the artifact honestly. A sketch is a sketch. Tell the agent it is a sketch, so it does not spend effort on polish, and tell yourself it is a sketch, so you do not fall in love with it. Many operators keep a separate folder or branch for throwaway artifacts precisely so that nothing in it can be mistaken for real work.

The second discipline is to end each artifact with a decision. The point of making the thing was to learn something. Once you have looked at it, write down what you learned and what you will do: keep this direction, abandon it, or change one specific thing and look again. An artifact that does not lead to a decision is just output, and the previous chapters have been clear about what output alone is worth.

Artifact first also improves your briefs. When you have a rough version in front of you, the next brief can point at it: keep the layout, change the tone of the copy, make the table sortable. Pointing is far more precise than describing. The first artifact is often less valuable for itself than for the vocabulary it gives you to ask for the second.

This week, find a decision you have been putting off because you could not quite picture the options. Ask an agent for two rough versions, side by side, in whatever form the thing will finally take. Give yourself ten minutes to look at them. Notice how quickly the decision makes itself once there is something to look at. It usually just needed a face.

Build roughly, then decide BEFORE AGENTS: THINK FIRST Think Decide Build takes days prototype is costly NOW: MAKE THE THING FIRST Build roughly minutes Look react to it Decide one clear call Build properly point at it or change one thing and look again Label it a sketch throwaway folder or branch; no polish End it with a decision keep, abandon, or change one thing You cannot review an intention. You can review an artifact.
Fig 8 · Artifact First. The old order of think then build against building roughly, looking and deciding.
Chapter 9 · Part I

The Small Machine Rule

Every operator eventually builds too much machine. It starts innocently. You write a helpful script, then a template, then a set of instructions for a recurring task, then a dashboard to track the recurring tasks, then an agent to update the dashboard. Each piece makes sense on its own. Together they become a system that requires its own maintenance, its own documentation and, before long, its own operator. You set out to run a small business and discovered you were running a small software company whose only customer is you.

The small machine rule is a defence against this. It says: keep the whole system small enough that you can hold it in your head on a bad day. Not on a good day, when you are rested and curious and enjoy tinkering, but on a tired Thursday afternoon when something has broken and you need to know where to look. If you cannot sketch your system from memory on the back of an envelope, it is too big.

What does a small machine contain? Usually far less than people expect. One register of projects, so you know what is in flight. A handful of reusable recipes for the work you do repeatedly. A small, stable set of tools you know well. And, at the bottom of everything, one daily rhythm that tells you when to review, when to work and when to stop. That rhythm is the foundation; the rest sits on top of it. An operator with a good rhythm and three tools will outperform one with a bad rhythm and thirty.

If you need a manual to run your manual, stop building.

The rule cuts against a natural instinct. Agents make building tools so easy that it can feel irresponsible not to automate every repeated step. But automation has a carrying cost. Each script must be kept working as the world around it changes. Each template drifts out of date. Each connected service needs its permissions reviewed and its failures noticed. The cost is small for each and large in aggregate, and it is paid in exactly the currency you are shortest of, which is attention.

So before adding a new piece to the machine, ask two questions. Will I use this at least weekly? And will I notice when it breaks? If the answer to either is no, do the task by hand or by a one-off brief, and wait until the pattern proves itself. A recipe that you have run manually five times is ready to be written down. A recipe you have run once is a guess.

It is equally useful to subtract. Once a quarter, list every part of your system and ask which ones you have not used in a month. Remove them, or at least move them somewhere out of sight. Operators rarely regret a removal. They frequently regret the hour spent debugging a clever automation that was saving them four minutes a week.

This week, draw your machine. One page, from memory. Whatever you forget to include is a candidate for deletion. Whatever you cannot explain in a sentence is a candidate for simplification. The machine you can draw is the machine you can run.

A machine you can draw from memory One register what is in flight A few recipes for work you repeat A handful of tools small, stable, known well One daily rhythm review · work · stop: the foundation top base BEFORE ADDING A PIECE Used at least weekly? Would I notice it break? Add it both yes By hand either no run it by hand 5 times before writing it down Each quarter: remove anything unused for a month If you need a manual to run your manual, stop building.
Fig 9 · The Small Machine Rule. A small machine rests on one daily rhythm, with two tests before adding any piece.
Chapter 10 · Part I

The Name on the Door

There is one fact about the operator's job that does not change no matter how capable the agents become: your name is on the door. When something ships, it ships as yours. When it breaks, it breaks as yours. A client, a reader or a user will never ask which agent wrote the paragraph or which thread produced the bug. They will ask you, and they will be right to.

This is not a burden to resent. It is the thing that makes the job a job. If the agents could own the outcomes, they would not need an operator, and the whole arrangement would be simpler and considerably less interesting for you. Ownership is what turns a pile of capable tools into a working operation. Somebody has to decide what good enough means, and somebody has to answer for it afterwards.

What follows from ownership is mostly a matter of attention. Agents do the work, broadly and quickly. You review the work, narrowly and carefully, because reviewing is where your ownership becomes real. And you own the result, which means that the review must be good enough that you would be comfortable defending the result to someone who matters. That progression narrows as it goes: lots of work, less review, one owner. The narrowing is the point.

You can delegate the effort. You cannot delegate the apology.

Ownership also shapes the briefs you write. An operator who knows they will answer for the outcome writes clearer constraints, asks for stronger evidence and is less tempted to wave things through on a Friday afternoon. It is a useful thought experiment, before approving anything, to imagine explaining it to the person most affected by it. If the explanation would start with well, the agent decided, you have not finished reviewing.

None of this means doing everything yourself out of anxiety. That is the opposite failure, and it is just as common among conscientious people. Ownership is compatible with heavy delegation, provided that what you delegate comes back through a checkpoint you control. A good editor owns a magazine without writing every article. A good captain owns the ship without turning every bolt. The skill is in placing the checkpoints where they will catch what matters, which is what most of the rest of this book is about.

There is a quieter benefit, too. Owning the result gives the work a centre of gravity. When you are juggling a dozen threads and a fleet of agents, it is easy to feel that the work is happening to you rather than through you. Remembering that you will sign it pulls you back into the operator's chair. You stop watching the machine and start running it.

So this week, before you approve anything that will be seen by another person, pause for five seconds and ask: would I sign this? Not is it probably fine, but would I put my name on it. Most of the time the answer will be yes. The few times it is not will be the most valuable five seconds of your week.

Work narrows to one name Agents do the work broadly, quickly many threads You review it narrowly, carefully checkpoints you control You own it one name on the door Before approving: would I sign this? not "the agent decided" You can delegate the effort. You cannot delegate the apology.
Fig 10 · The Name on the Door. Work narrows from many agent threads to your review and finally to your signature.
Part II

The Operating Day

Morning review, sessions and the close.

Chapter 11 · Part II

The Shape of a Good Day

A good operating day has a shape you could draw with three strokes. A short stretch of review at the start, a long middle of sessions, and a short close at the end. The details vary by person, by season and by how much coffee is in the house, but the shape holds. When operators describe a day that went well, they almost always describe this shape. When they describe a day that went badly, they usually describe its absence.

The review comes first because it is where the day is decided. You read what happened since you last looked, separate what needs your judgement from what does not, and pick a small number of outcomes for today. It is the part of the day with the highest leverage and the lowest drama. Nothing gets built during the review. Everything that gets built afterwards depends on it.

The sessions fill the middle. A session is a bounded stretch of work aimed at one outcome: you brief, the agents work, you review, you ship or send back. The day might hold two sessions or six, depending on their size and on you. What matters is that each one has a beginning and an end, so that the middle of the day is a sequence of finished attempts rather than one long smear of partial attention.

The close comes last because it is what makes tomorrow's review possible. You record what shipped, what is still open and what the first move tomorrow should be. Then you stop, in the full sense of the word: threads paused or left running on purpose, laptop closed, attention returned to whatever the rest of your life contains. The close is the part most operators skip, and its absence is why so many mornings start with twenty minutes of trying to remember where things were.

Begin on purpose, end on purpose, and the middle mostly takes care of itself.

Why does such a simple shape work so well? Because it separates two modes of thinking that interfere with each other. Review and close are about deciding: what matters, what is finished, what comes next. Sessions are about directing and judging specific work. When the modes blur, deciding gets done badly in the cracks between tasks, and tasks get interrupted by half-made decisions. Giving each mode its own time protects both.

The shape also gives you somewhere to put the things that do not fit. A new idea in the middle of a session goes onto a list for tomorrow's review, rather than becoming a new thread. A worry at the end of the day goes into the close notes rather than into your evening. The shape is a set of containers, and containers are what keep a busy system from spilling.

This week, do not try to perfect the day. Just put the three strokes in your calendar: a review block, a close block, and everything between them labelled sessions. Keep the review and close short, twenty minutes each is plenty to start with. See what happens to the middle when its edges are firm. Most people find it behaves better. Middles usually do, once someone has drawn the edges.

Three strokes: firm edges, a free middle deciding directing and judging deciding Review 20 min choose SESSIONS: 2 TO 6, EACH WITH AN END Session Session Session Session Close 20 min record, stop start of day end of day THE SHAPE IS A SET OF CONTAINERS New idea mid-session List for tomorrow's review not a new thread Worry at the end of day The close notes not your evening Begin on purpose, end on purpose; the middle mostly takes care of itself.
Fig 11 · The Shape of a Good Day. A firm review and close bracket a middle of bounded sessions, with containers for strays.
Chapter 12 · Part II

The Morning Review

The morning review is the twenty minutes that decide whether the rest of the day goes anywhere. It is not catching up. It is not checking messages. It is a small, structured act of triage that takes everything that has accumulated since yesterday's close and narrows it down to the handful of outcomes you will pursue today.

Start wide. Look at everything that ran or arrived since you last looked: background threads that finished, pull requests waiting, notes from yesterday's close, messages that need a reply, anything an automated job has flagged. Do not act on any of it yet. The temptation is to fix the first small thing you see, because it is satisfying and quick. Do not. The first small thing is rarely the most important thing, and once you start fixing, the review is over and the day has been chosen by accident.

Then narrow. For each item, ask whether it needs a decision from you. Many do not. A thread that finished and passed its checks may only need merging, which is a two-second action you can batch. A notification that something ran successfully needs nothing at all. What remains, the items that genuinely need your judgement, is your decision list. It is usually shorter than the inbox suggested.

Finally, choose. From the decision list and from your register of projects, pick the outcomes for today. The next chapter suggests three, but the number matters less than the act of choosing. Write them down where you will see them all day. They are the day's contract with itself.

The review is where the day is won, quietly, before anything has happened.

A few practical notes. Do the review before opening anything that can pull you into a conversation, because conversations are designed to be continued and the review is designed to be finished. Keep it time-boxed; if it takes forty minutes every morning, something upstream is producing too much noise, and that is worth fixing in its own right. And do it in the same place each day, with the same sequence, so that it becomes a habit rather than a decision. Decisions are expensive first thing. Habits are free.

Some operators ask an agent to prepare the review for them: a short digest of what ran, what passed, what failed and what is waiting. This is a good use of an agent, provided that the digest is a starting point and not a substitute. Read the digest, then glance at the underlying items for anything that smells wrong. The agent can summarise. It cannot yet tell you which of the summarised things you will care about most at eleven o'clock.

Try this tomorrow. Before you open anything, write today's date and the words today's outcomes on a blank page. Then do the review, wide, narrow, choose, and fill in the page. Close the review deliberately, perhaps by closing the tab you did it in. Then start the first session. Notice how different the first hour feels when it begins with a decision rather than a scroll.

Wide, narrow, choose 1 · WIDE: LOOK, DO NOT ACT 2 · NARROW 3 · CHOOSE Everything that ran since the last close background threads pull requests waiting yesterday's close notes messages to answer automated flags Needs a decision? the decision list keep only what needs your judgement shorter than the inbox Today's outcomes three, finishable written down the day's contract No: batch it or ignore merge, or nothing 20 minutes · before any conversation · same place, same order The review is where the day is won, quietly, before anything has happened.
Fig 12 · The Morning Review. The morning review looks wide, narrows to decisions and chooses today's outcomes.
Chapter 13 · Part II

Reading What Ran Overnight

One of the great pleasures of running agents is waking up to finished work. You briefed a thread before bed, it ran while you slept, and now there is a pull request, a draft, a report or a dataset waiting. It feels like having staff. It is also one of the easiest places to make a quiet, expensive mistake, because work done while you were not watching is work you know least about.

The first rule of reading overnight work is to read the result before the summary. Agents write good summaries, and good summaries are persuasive. They tell you what was done, why, and that everything passed. They are not lies, but they are written by the party that did the work, and they naturally emphasise what went well. Open the actual output first: the diff, the document, the data. Form your own impression. Then read the summary and see whether it agrees with you.

The second rule is to check the evidence, not just its existence. A thread that says all tests pass has told you something, but you want to know which tests, and whether any of them actually exercise the new work. A report that cites sources should have sources you can open. A dataset that claims a hundred rows should have a hundred rows. These checks take a minute each. They are the minute that separates trusting the machine from hoping the machine was right.

Unwatched work deserves a slower first read than watched work, not a faster one.

The third rule is to decide, cleanly, one of three things: ship it, send it back with specific notes, or park it with a reason. Do not leave overnight work in an ambiguous state, opened and half-read, because that is how it becomes the fourth item on tomorrow's review as well. If it is good, merge or publish it now. If it needs changes, write the changes as a short brief and send the thread round again. If it raised a question you cannot answer yet, write the question down in the register next to the project and close the tab.

There is a pattern worth noticing over time. Some kinds of overnight work come back reliably good, and some come back reliably muddled. The first kind is usually narrow, well-specified and easy to verify. The second is usually broad, exploratory or dependent on judgement calls the agent had to make without you. This is not a reason to stop running broad work overnight. It is a reason to brief it differently: ask for options rather than a decision, and expect to spend more review time on it in the morning.

This week, keep a short note of every overnight result you read: what it was, whether you shipped, resent or parked it, and how long the read took. By Friday you will know which kinds of work are safe to run while you sleep and which kinds really need you awake. That knowledge is worth more than any amount of agent capability, because it tells you where to point the capability you already have.

Read the work before the summary Overnight result PR, draft, data 1 · Read output diff, doc, data 2 · Read summary does it agree? 3 · Evidence which tests? decide cleanly, one of three Ship it merge or publish now Send it back notes as a short brief Park it question in the register never: opened, half-read, left for tomorrow OVER TIME, NOTICE Comes back good narrow, specified, easy to verify Comes back muddled broad: ask for options instead
Fig 13 · Reading What Ran Overnight. Read overnight output, then the summary, then the evidence, and decide one of three ways.
Chapter 14 · Part II

Choosing Today's Three

Pick three outcomes for the day. Not three tasks, not thirty, and not one. Three outcomes, each something that will be visibly finished by the close: shipped, sent, decided or published. This is the single most useful constraint an operator can adopt, and it is useful precisely because it feels too small.

Why three? Partly because it is roughly what one person can review properly in a day when agents are doing most of the work. Each outcome will need at least one careful review, often two or three, and each review needs genuine attention. Partly because three is enough to absorb a disappointment: if one outcome stalls, the day still produces two. And partly because three is a number you can hold in your head without writing it down, although you should write it down anyway.

The choosing is where the skill lies. A useful way to think about it is to place candidate outcomes on two axes, impact and effort. Impact is how much finishing it would change: for a client, for users, for the health of your projects. Effort is not the agent's effort, which is cheap, but yours: how much briefing, reviewing and deciding it will require. The outcomes you want are high in impact and moderate in your effort. Those are today's three.

A list of thirty is a wish. A list of three is a plan.

Be wary of two common distortions. The first is choosing only easy outcomes, because finishing feels good and easy things finish. A day of three trivial wins is pleasant and does not move anything that matters. Make sure at least one of the three is something you would have been slightly reluctant to start. The second is choosing only huge outcomes, because they feel important. A huge outcome rarely finishes in a day. If something matters but is large, the outcome for today should be a slice of it that can finish: the first section shipped, the design decided, the data cleaned.

The three are not a prison. Things will come up. If something urgent arrives, you can swap it in, but do it explicitly: cross one off, write the new one in, and notice that you made a trade. What you should not do is add a fourth and fifth silently. That is how three becomes eight, and eight becomes a day where nothing quite finishes.

At the close, look at the three. Mark each as done, partly done or not done, and write a sentence about why for the ones that slipped. Over a few weeks a pattern will appear. Perhaps you consistently underestimate review time. Perhaps afternoons are where outcomes go to die. Perhaps you choose well on Mondays and badly on Thursdays. Each pattern is a lever.

Start tomorrow. In the morning review, write three outcomes at the top of the page. Make each one finishable, specific and worth doing. Then spend the day on them, and only them, unless you consciously trade. It is a modest discipline. Its effect on a week is not modest at all.

Choosing today's three Impact low low Your effort: briefing, review Quick wins rare; take them Huge will not finish today take a slice that can Trivial wins pleasant; moves nothing Avoid costly and minor Today's three high impact moderate effort at close: done · partly · not swap one out, never add a 4th
Fig 14 · Choosing Today's Three. Today's three sit at high impact and moderate effort, away from trivial or huge work.
Chapter 15 · Part II

Time-Blocked Sessions

A session is a block of time with one purpose. It begins with a brief, ends with a decision, and contains nothing else. Operators who work in sessions get more done than operators who work in streams, for the same reason that a kitchen with timers produces more meals than one where everything is left to simmer until someone remembers it.

The structure inside a session is simple. You write or load the brief, ideally in the first ten minutes. You let the agents run, watching enough to catch a wrong turn early but not so closely that you are effectively doing the work through them. You review what comes back, carefully. And you decide: ship it, send it round again within the session, or park it with a note. The review and decision step is the heart of the session, and it is the part most likely to be squeezed if the session has no fixed end.

How long should a session be? Long enough to finish something, short enough that you stay sharp through the review. For many people that is somewhere between forty minutes and two hours. The exact length matters less than the fact that you decided it in advance. A session with a fixed end changes your behaviour at the start: you brief more tightly, because you know there is limited time to correct a vague brief later.

Give the work a container and it will mostly stay inside it.

Sessions also make parallel work manageable. It is tempting, with agents, to start many threads and drift between them. That can work, but it works far better when each thread belongs to a session with an owner and an end. Within a session you might run three agents on three slices of the same outcome. What you should avoid is five threads on five unrelated outcomes, all half-watched, none reviewed properly. That is not parallel work. It is parallel neglect.

Put sessions in the calendar, even if the calendar is only yours. Label each with its outcome, not its activity: ship importer fix, not work on importer. When the session ends, stop, even if the work is not quite done. Write a short note of where it stands, and either schedule another session or move the outcome back to the register. The discipline of stopping is uncomfortable at first. It is also what teaches you, over a few weeks, how big a session-sized piece of work really is.

Between sessions, take a real break. Not a break spent checking other threads, which is just another session in disguise, but a break in which nothing is being reviewed. The agents can keep running. You do not have to. The quality of your next review depends far more on that break than on anything you could have read during it.

This week, try running every piece of work through a session: a block in the calendar, an outcome in its title, a decision at its end. Count how many sessions you complete each day. That number, more than hours worked, is a good measure of an operator's real capacity. Most people find it is smaller than they assumed, and that working to it rather than against it is a great relief.

Inside one session Brief first 10 min Run and watch catch wrong turns Review and decide the heart of it Ship or park with a note 0:00 40 min to 2 h: a fixed end Parallel work three agents, three slices Agent Agent Agent One outcome Parallel neglect five threads, five outcomes T1 ? T2 ? T3 ? T4 ? T5 ? Give the work a container and it will mostly stay inside it.
Fig 15 · Time-Blocked Sessions. A session runs brief to decision; parallel slices of one outcome beat parallel neglect.
Chapter 16 · Part II

Starting a Session Cleanly

The first five minutes of a session decide most of what happens in the next hour. A session that starts cleanly, with the right context loaded and a clear brief, tends to run smoothly. A session that starts with right, where was I tends to spend its first half finding out, and its second half recovering from the guesses it made while it was finding out.

The clean start has three moves. First, read the handoff. If this outcome was touched before, there should be a note from the last session or from yesterday's close saying where it stands: what was done, what is open, what the next move was meant to be. Read it, and read any relevant entries in the project's memory. This takes two minutes and saves twenty. If there is no handoff, that tells you something about how the last session ended, which you can fix later.

Second, write the brief. Even if the session is a continuation, write down in a few sentences what this session is for, what done looks like and what the agent should know. This is partly for the agent and partly for you. Writing the brief forces you to decide what you actually want from the next hour, and the act of deciding is half the value of the session.

Third, confirm the plan before work begins. For anything bigger than a trivial change, ask the agent to say how it intends to proceed before it does. Many agent tools have a planning mode for exactly this reason. Read the plan. If it is wrong, correcting it now costs a sentence. Correcting it after forty minutes of work costs the forty minutes and your patience.

A wrong plan is cheap. A wrong plan carried out is not.

These moves sound fussy, and for small tasks they can be compressed into a single line of brief. But the habit matters more than the ceremony. Operators who skip the clean start often blame the agent when a session goes sideways: it misunderstood, it went off in a strange direction, it did not know about the constraint. In most of those cases, the agent was working from exactly what it was given, which was not much.

A clean start also includes a clean environment. Close threads that are not part of this session. Clear away the tabs from the last one. If the agent works in a repository, make sure you are on the right branch and that nothing half-finished is lying around from before. These are small acts of tidiness, and they prevent a whole class of confusing errors in which the agent is working on one thing while some leftover from another quietly interferes.

Try timing it. For the next five sessions, note how long the clean start takes, and then rate the session afterwards on a simple scale: smooth, bumpy or wasted. Most operators find that the sessions with a proper start are rated smooth far more often, and that the start rarely took longer than five minutes. Five minutes is a small price for a smooth hour. It is also, conveniently, about how long it takes to make a cup of tea, and the two pair well.

The first five minutes 0:00 Clear the bench close strays, right branch stops leftovers interfering 0:01 Read the handoff last note + project memory two minutes save twenty 0:03 Write the brief what for, done, what to know decides what the hour is for 0:04 Confirm the plan ask first, then read it a wrong plan costs a sentence 0:05 Work begins smooth, bumpy or wasted? rate it afterwards A wrong plan is cheap. A wrong plan carried out is not.
Fig 16 · Starting a Session Cleanly. Five minutes of clean start: clear the bench, read the handoff, brief, confirm the plan.
Chapter 17 · Part II

The Middle of the Day

Once the agents are running, the operator faces a curious problem: what to do with themselves. The work is happening. It does not need your hands. It does, occasionally, need your judgement, and you do not know exactly when. So you hover. You watch the output scroll, you read every intermediate step, you open the files as they change. It feels responsible. It is mostly a way of spending attention without buying anything with it.

The alternative is a rhythm of checking in, unblocking and stepping away. Check in at sensible intervals, not continuously. Look at where each running thread has got to, whether it is waiting for you, and whether it is heading somewhere plausible. If something is blocked on a question, answer it. If something is drifting, redirect it with a sentence. Then step away and do something that is not watching: review a finished piece of work, write tomorrow's brief, think about a decision, or simply take a walk.

The unblocking is the valuable part, and it is worth thinking about why. Agents get stuck on things that need a decision they cannot make: which of two approaches you prefer, whether a constraint really applies, what to do about something unexpected. An hour lost waiting for that decision is an hour of agent time wasted. A minute spent making it is the highest-leverage minute in the middle of the day. So when you check in, look first for anything waiting on you, and clear it before anything else.

Hovering spends attention. Unblocking invests it.

How often should you check in? It depends on the work and on how well it was briefed. A tightly specified task with clear checks can run for a long time without you. An exploratory task with fuzzy edges needs more frequent looks, because the chance of a wrong turn is higher. A useful rule is to check in at roughly the interval at which a wrong turn would become annoying to undo. For many tasks that is every fifteen or twenty minutes. For some it is once an hour.

Many agent tools now notify you when a thread needs input or finishes. Use those notifications, but tune them. If everything notifies you, nothing does. Ideally, the only interruptions you receive in the middle of the day are ones that need a decision, and everything else waits for your next check-in or your close.

The middle of the day is also where the morning's choices are tested. If today's three are well chosen and the briefs are clear, the middle is calm: you check in, unblock, step away, and review finished work as it arrives. If the middle feels frantic, that is usually a symptom of something upstream. Too many threads were started. A brief was vague. An outcome was too big. Make a note of the symptom for tomorrow's review rather than trying to fix the cause in the middle of the storm.

This week, set a timer for your check-ins instead of watching continuously. When it goes, look, unblock, and step away again. See how much of the day it returns to you. The agents will not notice the difference. You will.

Check in, unblock, step away Check in where is each thread? Unblock clear waits first high leverage Step away review, write, walk Not: hovering reading every step buys nothing CHECK-IN INTERVAL: WHEN A WRONG TURN WOULD GET ANNOYING TO UNDO Tight brief, clear checks: hourly Fuzzy, exploratory: every 15 min Hovering spends attention. Unblocking invests it.
Fig 17 · The Middle of the Day. The middle-of-day loop of checking in, unblocking and stepping away, instead of hovering.
Chapter 18 · Part II

Interrupts and the Queue

Ideas arrive at inconvenient times. You are halfway through reviewing a pull request and you think of a feature. You are writing a brief and you remember an email you meant to send. A message arrives asking for something small. An agent, mid-run, suggests an improvement to something else entirely. Each of these is a small interrupt, and each one, if acted on immediately, costs not just the time it takes but the time it takes to find your place again afterwards.

Agents make interrupts more dangerous, because they make acting on them so cheap. In the old world, a new idea mid-afternoon would have to wait because you were busy. Now you can open a new thread and start it in thirty seconds, and the thirty seconds feel free. They are not. The new thread will need briefing, watching, reviewing and deciding, and it has just quietly taken a slice of today's attention that was already promised to today's three.

The answer is a queue. Somewhere simple, a text file or a list in your register, where interrupts go to wait. When something arrives, ask one question: does this need to happen today, at a cost to today's outcomes? If yes, do it now, and consciously trade it against one of the three. If no, and it is almost always no, write it in the queue in one line and go back to what you were doing. It will be there in tomorrow's review, where it can compete fairly with everything else.

An idea written down is not lost. An idea acted on immediately often is.

The queue works because it removes the anxiety of forgetting. Much of the urge to act on an interrupt comes from fear that the idea will vanish if it is not seized. Once you trust that the queue will hold it and that the review will look at it, the urge fades. You can let a good idea go into the queue with the same calm you let a letter go into a postbox.

It is worth being honest about the urgent category. Very few things genuinely need to happen today. A production failure does. A promise to a client with a deadline today does. A message from someone who is waiting on you to proceed might. Almost nothing else does, no matter how urgent it feels, and the feeling of urgency is often just the feeling of novelty in disguise.

Agents can help with the queue, too. A good habit is to ask agents, when they notice something outside the current task, to note it rather than fix it. If you find other problems, list them at the end; do not change them is a line that prevents a great deal of scope creep. Those notes then go into the queue, and the current task stays the current task.

This week, keep a queue and use it for every interrupt that is not truly urgent. At the end of the week, look at what accumulated. Some of it will be valuable and will get scheduled. A surprising amount will look, on Friday, like it was never that important. That surprising amount is the attention you saved.

Every interrupt gets one question a new idea an email a message an agent's idea Needed today, at a cost? yes, rarely no, usually Do it now outage, client deadline and trade it against one of today's three, on purpose Into the queue a text file in the register one line, then back to work nothing is lost Tomorrow's review competes fairly Agents: list it, do not fix it An idea written down is not lost. An idea acted on at once often is.
Fig 18 · Interrupts and the Queue. Every interrupt meets one question; nearly all go to the queue for tomorrow's review.
Chapter 19 · Part II

The Finish Command

Every session should end with the same small ritual. Think of it as a finish command: a fixed sequence that turns a stretch of work into something settled. Verify, ship, note, stop. It takes five minutes. It is the difference between a session that finished and a session that merely ended because you ran out of time.

Verify first. Before anything else, check that what you believe is done is actually done. Run the checks one more time if the last run was before the final changes. Open the page, read the document, look at the data. Agents are good at reporting completion; they are not infallible at it, and the end of a session, when you are keen to move on, is exactly when you are most likely to take the report on trust.

Then ship. If the work is ready, put it where it goes: merge the pull request, publish the post, send the email, update the record. Do not leave finished work sitting in a branch or a draft folder overnight. It will not improve there, and it will become an item on tomorrow's review that requires you to rebuild your confidence in it from scratch. If it is not ready, decide explicitly that it is not, and why.

Then note. Write a short handoff: what was done, what is open, what the next move is. Put it where the next session will look, whether that is the project's memory file, the register, or a note at the top of the thread. Write it for someone who has forgotten everything, because by tomorrow that someone will be you. Many operators ask an agent to draft this note and then correct it, which is a fine division of labour as long as the correction actually happens.

A session that ends without a note ends twice: once now, and again tomorrow when you work out what happened.

Then stop. Close the thread, or leave it running on purpose with a clear task if it is a long background job. Close the tabs. Get up. The stop is part of the command, not an afterthought, because a session that does not stop leaks into the next one and blurs it.

Some operators formalise this as an actual command: a short saved instruction they paste or invoke at the end of a session that asks the agent to run the final checks, summarise the state, list open questions and draft the handoff note. That is a good idea, and it makes the ritual easy enough to do every time. The command is only as good as your willingness to read what it produces, though. The point is not to automate the finish. It is to make sure the finish happens.

This week, end every session with the four steps, in order, even when you are tired, even when the session went badly. Especially when it went badly. A bad session with a good note is recoverable tomorrow in five minutes. A bad session with no note tends to stay bad for days. The finish command is what turns an hour of effort into an hour of progress, and it costs less than the effort it protects.

The finish command Verify step 1 re-run the checks open the page look at the data Ship step 2 merge or publish or say why not Note step 3 what was done what is open the next move Stop step 4 close the tabs get up Saved as a command /finish: checks, state, open questions, draft note Bad session, good note recoverable in five minutes Bad session, no note stays bad for days A session that ends without a note ends twice.
Fig 19 · The Finish Command. The four-step finish command: verify, ship, note and stop.
Chapter 20 · Part II

Closing the Day

The close is the last twenty minutes of the operating day, and it is the part most operators skip. It feels optional. The work is done, or not done, and nothing will change if you leave it until morning. But the close is not about today. It is about tomorrow's morning review, and about the evening that sits between them.

Start by recording what shipped. Look at today's three and mark each one. Note anything else that finished. This takes a minute and has an effect out of proportion to its size: it turns a vague sense of having been busy into a concrete record of what changed. On good days the record is satisfying. On bad days it is informative. On both kinds of day it is better than the feeling you would otherwise have taken to bed.

Then record what is still open. Every thread that is mid-flight, every review that is half-done, every question you are waiting on someone to answer. For each, a line: where it stands and what the next move is. If you have been writing handoff notes at the end of each session, this step is mostly a matter of gathering them. If you have not, this is where you discover why you should.

Then write tomorrow. Not a full plan, which is the morning review's job, but a nudge: the first thing you intend to look at, and perhaps a candidate for one of tomorrow's three. Starting the morning with a suggestion from your past self makes the review far quicker. It also lets you hand off any lingering worry. If something is bothering you, write it down in this section, so that it has somewhere to live that is not your head.

Write down tomorrow so that tonight can be tonight.

Then stop. Decide which agents, if any, will keep working overnight, and make sure each has a clear, bounded task and a check it can run. Pause anything that does not need to run. Close the laptop. The stop is not a reward for having finished; it is a structural part of the machine. An operator who never stops does not get more done. They get the same amount done, worse, with less judgement, and they discover the cost only when a tired review lets something through that should have been caught.

There is a temptation, particularly when agents are running, to check in one more time after the close. The thread might have finished. There might be something interesting. Resist it. Whatever has happened will still be there in the morning, and the morning review is designed to deal with it calmly. Checking late at night replaces a calm morning review with an anxious evening one, which is a bad trade on every measure.

Tonight, do a proper close. Twenty minutes: shipped, still open, tomorrow, stop. Then notice how tomorrow's morning review goes. Most people find it is faster, calmer and more decisive. The close does not just end the day. It starts the next one, quietly, while you are not looking.

Twenty minutes that start tomorrow Close · Tuesday 20 min SHIPPED done: importer fix merged partly: pricing page Mark today's three done, partly or not done STILL OPEN export thread: tests left waiting on Sam: scope A line per open thread where it is, next move TOMORROW first look: the export PR worry parked: invoice date A nudge, not a plan and a home for the worry STOP overnight: data clean-up everything else paused Decide what keeps running bounded task with a check Write down tomorrow so that tonight can be tonight.
Fig 20 · Closing the Day. A close note records shipped, still open, tomorrow and stop, each with its purpose.
Part III

Briefs Agents Can Follow

Goal, context, constraints and done.

Chapter 21 · Part III

A Brief Is a Contract

A brief is the document that turns your intention into an agent's work. It can be a single sentence or a page, typed into a chat box or saved as a file, but in every case it does the same job: it says what you want, what the agent needs to know, what it must not do and how you will judge the result. When the brief is good, the work tends to be good. When the brief is poor, the work is often impressively, confidently poor.

It helps to think of a brief as a contract rather than a request. A request is something you hope will be fulfilled. A contract is something both parties can check against afterwards. The difference shows up at review time. If your brief said improve the landing page, you have no way of knowing whether the result fulfils it, because almost any change can be argued to be an improvement. If your brief said the landing page headline fits on one line on a phone, and the sign-up button is visible without scrolling, you can check in ten seconds.

A useful contract has four parts, stacked from top to bottom. The goal: one sentence, phrased as a result. The context: whatever the agent needs to know that it cannot easily discover, such as who the audience is, which files matter, or what was tried before. The constraints: what must not change, what tools or approaches are off-limits, how far the work may reach. And the definition of done: the checks that will tell both of you the work is finished. The last layer is the foundation, which is why it sits at the bottom. Everything above it depends on it.

If you cannot check it, you did not ask for it. You hoped for it.

Not every brief needs all four parts spelled out. A small task in a familiar project might need only a goal and a check, because the context and constraints already live in the project's memory. But it is worth running through the four parts mentally every time, because the one you skip is usually the one that causes trouble. Missing context leads to plausible work aimed at the wrong target. Missing constraints lead to work that sprawls into places you did not want touched. Missing done leads to arguments with yourself at review time.

Writing briefs this way also changes how you think about the work before it starts. Often, when you sit down to write the definition of done, you discover that you do not know what done is. That is not a failure of the brief. It is the brief doing its job early, before an agent has spent an hour building towards a target you had not chosen.

This week, take the next three tasks you would normally fire off in a sentence and write them as four-part contracts instead. Keep each part short, a line or two. Then compare the results to what you usually get. The agents have not become cleverer. You have simply told them, for the first time, what clever would look like.

Four layers of a brief IF THIS LAYER IS MISSING Goal one sentence, phrased as a result no target at all Context what it cannot discover itself plausible work, wrong target Constraints what must not change or be used sprawl into places you did not want Done looks like the checks both of you can run arguments with yourself at review foundation Request: improve the page anything counts as improved Contract: headline fits a phone checkable in ten seconds If you cannot check it, you did not ask for it. You hoped for it.
Fig 21 · A Brief Is a Contract. A brief stacks goal, context, constraints and done, and each missing layer has a cost.
Chapter 22 · Part III

Start With Done

If you only change one habit in how you brief agents, change this one: write the definition of done before you write the task. Not after, as an afterthought, and not implicitly, assuming the agent will know. First. It is the single most reliable way to improve the quality of what comes back.

The reason is that done is where your judgement lives. The task description says what to work on; the definition of done says what you will accept. An agent can work on almost anything competently, but it cannot read your standards from the air. When it does not know what you will accept, it makes a reasonable guess, and reasonable guesses are wrong just often enough to be expensive.

Writing done first also protects you from a specific trap. If you write the task first, the definition of done tends to be shaped by the task: add a filter to the table becomes the table has a filter. That is circular. If you write done first, you start from the result you actually care about: a user can find last month's orders in under ten seconds. The task then follows from it, and might turn out to be a filter, or a search box, or a default sort order. Starting with done leaves room for a better task than the one you first thought of.

The task is how. Done is why. Write the why first.

Good definitions of done share some features. They are observable: someone could check them without reading your mind. They are specific: they name a number, a behaviour or a state rather than a quality. And they include at least one check that would fail if the work were missing. All tests pass is weak, because all tests may have passed before you started. A new test covers the empty-file case, fails on the current code and passes after the change is strong, because it proves the change did something.

For non-code work the principle is identical. A report is done when it answers the three questions it was commissioned to answer, cites a source for each claim, and fits on two pages. A design is done when it works at phone width, has no text smaller than a stated size and uses only the agreed colours. An email is done when it can be read in under a minute and asks for exactly one thing. None of these is hard to write. All of them are routinely left out.

Once you have the definition of done, give it to the agent and ask the agent to check against it before reporting back. Many agents will now do this naturally if asked: run the checks, quote the results, flag anything that does not meet the bar. That turns your review from an open-ended inspection into a confirmation, which is faster and far less tiring.

This week, write Done when: as the first line of every brief, and fill it in before anything else. If you find you cannot, stop and think, because the work is not ready to be delegated yet. It is ready to be decided.

Write done before the task TASK FIRST Add a filter Done: table has a filter circular: proves nothing DONE FIRST A user finds last month's orders in under ten seconds A filter a possible task A search box a possible task A default sort a possible task room for a better task than the first idea Weak: all tests pass they passed before, too Strong: a new test fails, then passes proves the change did something The task is how. Done is why. Write the why first.
Fig 22 · Start With Done. Starting from done opens several possible tasks and demands a check that can fail.
Chapter 23 · Part III

Context, Not Biography

Agents need context. They need to know things about your project, your audience and your situation that they cannot discover by looking. The difficulty is that you know a very great deal, and most of it is irrelevant to any particular task. The art of the brief is giving the context the task needs and no more: context, not biography.

Picture two overlapping circles. One holds everything you know: the history of the project, your preferences, the client's quirks, the three approaches you rejected last spring, your opinion on semicolons. The other holds everything the task needs to be done well. The brief should contain the overlap and very little else. Too little, and the agent works blind. Too much, and the agent gives weight to things that do not matter, or loses the important detail in the noise.

Over-sharing is the more common problem among conscientious operators. They worry that the agent will miss something, so they include everything. The brief grows to a page of history, caveats and asides. The agent, which takes everything in front of it seriously, now tries to honour all of it. A passing mention that the client once disliked blue becomes a design constraint. An old abandoned approach, described for background, gets partly resurrected. The work comes back shaped by the biography rather than the task.

Every sentence in a brief is an instruction, whether you meant it as one or not.

A good test for any piece of context is to ask: if this sentence were removed, would the result plausibly be worse? If yes, keep it. If you are not sure, keep it but make its status clear: for background only, do not act on this. If no, cut it. You can always add context in a follow-up if the agent asks or the result shows it was needed.

The other half of the skill is noticing what the agent cannot find for itself. Agents working in a repository can read the code; they do not need it described to them. They cannot read your conversations with the client, the decision you made in the shower, or the fact that the staging server is down this week. Those are exactly the things to include. Point at what they can find, tell them what they cannot.

It is also worth separating stable context from task context. Stable context, the things that are true across many tasks, belongs in the project's memory, where every session will see it. Task context, the things that matter only here, belongs in the brief. Mixing the two means repeating yourself in every brief, or worse, forgetting to and getting inconsistent results. A later part of this book covers memory in detail.

This week, take a brief you have already written and go through it sentence by sentence with the removal test. Cut everything that fails. Then send the shorter brief and compare. Most operators are surprised that less context produces sharper work. Agents, like people, do better when they are told what matters rather than everything that has ever happened.

Context, not biography What you know leave it out project history your preferences last spring's rejects views on semicolons What it can find point at it the code itself the file layout the test suite The brief client calls your decisions staging is down only you can say Removal test if this sentence went, would the result be worse? stable facts go to memory
Fig 23 · Context, Not Biography. The brief holds only what the task needs and the agent cannot discover for itself.
Chapter 24 · Part III

Constraints Are a Kindness

It can feel ungenerous to fill a brief with restrictions. Do not change the public interface. Do not add dependencies. Do not touch the billing code. Keep it under three hundred words. Use only the existing colours. It reads like a list of prohibitions, and many people instinctively soften it or leave it out, hoping the agent will use good sense. But constraints are not a lack of trust. They are a kindness, to the agent and to your future self.

An agent without constraints has to guess where the edges of the task are. Should it refactor the messy function next door while it is there? Should it update the dependency that triggers a warning? Should it rewrite the introduction, since the introduction is also weak? Each of these is a reasonable thing to do, and each one expands the change, the review and the risk. Agents tend to be helpful, and helpfulness without edges sprawls.

With constraints, the agent can focus. It knows what it may touch and what it may not, so it spends its effort inside the boundary. The work comes back smaller, more coherent and much easier to review. You do not have to spend twenty minutes working out which parts of a large change were the task and which were enthusiasm.

A fence is not a cage. It is a description of the garden.

The most useful constraints fall into a few families. Scope constraints say what may change: these files, this section, this feature only. Interface constraints say what must not change: the API, the URL structure, the tone of voice, the data format. Tool constraints say how the work may be done: no new libraries, no network calls, no changes to configuration. And size constraints say how big the result may be: a word count, a line count, a number of options. Choose the ones that matter for the task. Usually two or three are enough.

There is a companion habit that makes constraints work even better: ask the agent to tell you when a constraint gets in the way. Sometimes the right fix really does require changing the interface or adding a library. You want to know that, but you want to decide it, rather than discover it in the diff. A line such as if you believe a constraint should be broken, stop and explain why before doing so turns constraints from walls into checkpoints.

Constraints are also a record. When you look back at a brief in a month's time, the constraints tell you what you were protecting and why. That is often more useful than the task description, because it captures the parts of the project you considered fragile or important at the time.

This week, add a Do not line to each brief, with two or three specific constraints. Notice how the size of the resulting changes shrinks, and how much faster the reviews become. You have not limited the agent. You have pointed it, which is a different and much more useful thing.

Constraints describe the garden Scope these files only Interface API and tone stay Tools no new libraries Size under 300 words The task effort stays inside the fence: two or three is enough OUTSIDE THE FENCE Refactor next door Bump a package Rewrite the intro In the way? a checkpoint stop and explain then you decide A fence is not a cage. It is a description of the garden. add a "Do not:" line with two or three specifics to every brief
Fig 24 · Constraints Are a Kindness. Four families of constraints fence the task in and keep tempting extras outside.
Chapter 25 · Part III

Name the Files, Name the Checks

The quickest way to improve a brief is to replace descriptions with pointers. Instead of the bit of the code that handles uploads, write the file path. Instead of the usual test command, write the command. Instead of a style like our other pages, name a page. A pointer is shorter than a description, more precise, and impossible to misread.

Agents are very good at following pointers. Give one a path and it will open the file, read it and orient itself. Give it a description and it will search, find several candidates, choose the most plausible one and proceed. Usually it chooses right. When it chooses wrong, it does so confidently, and you discover the mistake at review time, after the work has been built on the wrong foundation. A path costs you five seconds to copy. It saves the agent's search and your doubt.

The same goes for examples. If you want a new component to look like an existing one, point at the existing one. If you want a report in the same structure as last month's, point at last month's. If you want the tone of a particular email, paste it or link it. Examples carry a vast amount of implicit instruction that would take paragraphs to describe and would still be described imperfectly.

And it goes, above all, for checks. Name the exact command, page, query or test that will confirm the work. Not make sure it works, but run npm test and the new test in upload.test.ts must pass. Not check it looks right, but load /pricing at phone width and confirm nothing overflows. A named check is something the agent can actually run, and something you can actually confirm. An unnamed check is an aspiration.

A pointer is a sentence the agent cannot misunderstand.

This habit pays off twice. Once in the quality of the first attempt, because the agent starts in the right place with the right model. And again in the review, because you know exactly where to look. If the brief named three files and a check, your review starts with those three files and that check. You are not hunting through a large change trying to reconstruct what it was meant to do.

It also exposes gaps in your own understanding. If you sit down to name the files and realise you do not know which ones are involved, that is valuable to know before the agent starts. You might ask the agent to find out first, as a separate small task: list the files involved in uploads and describe each in one line. Then brief the real work with pointers in hand. Two small tasks with good pointers often beat one large task with vague ones.

This week, read each brief before you send it and underline every description that could be a pointer instead. Replace each one: a path, an example, a command. It feels pedantic for about two days. After that it feels like the obvious way to write, and the vague briefs you used to send start to look like riddles you were setting for no good reason.

Replace descriptions with pointers Description Pointer Files the bit that handles uploads src/upload.ts Examples a style like our other pages match /pricing Checks make sure it works npm test: upload.test.ts Result Agent searches, guesses confidently wrong, sometimes Starts in the right place and you know where to look A pointer is a sentence the agent cannot misunderstand.
Fig 25 · Name the Files, Name the Checks. Descriptions swapped for pointers to files, examples and checks the agent can run.
Chapter 26 · Part III

The One-Paragraph Brief

Most briefs fit in a paragraph. Not all of them, and the exceptions are real, but most. A sentence of goal, a sentence or two of context, a sentence of constraints and a sentence of done. Five or six sentences, a hundred words or so, and the agent has everything it needs. If your briefs are routinely running to a page, it is worth asking why.

The usual reason is that the thinking was not finished before the writing started. A long brief is often a record of the operator working out what they want: options considered, worries expressed, half-decisions left open. All of that thinking was necessary. None of it needs to be in the brief. The brief is the conclusion, not the working.

So treat the one-paragraph brief as a discipline. Write freely first if you need to, in a notebook or a scratch file. Then distil: take everything you know, keep only what the task needs, and compress that into a paragraph. If you cannot compress it, there are two likely explanations. Either the task is really several tasks and should be split, or you have not yet decided something that the agent will need decided. Both are worth discovering before work begins.

A long brief is often a short brief that has not finished being written.

Here is the shape, in prose rather than a template. Goal: users can export their saved items as a spreadsheet from the account page. Context: the account page lives in app/account, saved items come from the existing items query, and we already use a library for spreadsheet export in reports/. Do not: add new dependencies or change the items query. Done when: a new export button appears, the exported file opens in a spreadsheet program with one row per item, and a test covers the empty case. That is the whole brief. An agent can do excellent work from it, and you can review that work in ten minutes.

There are honest exceptions. A brief for a large, multi-step piece of work may need more structure: a short list of phases, each with its own done. A brief for creative work may need examples that take space. A brief for a delicate piece of work in an unfamiliar area may need more context than usual. But even then, the one-paragraph version is a useful first step. Write it, and then add only what the paragraph genuinely cannot carry.

Short briefs have another virtue that is easy to miss: they are reusable. A paragraph can be saved, adapted and sent again next month with a few words changed. A page of tangled thinking cannot. Over time, an operator who writes one-paragraph briefs builds up a library of them, and a library of good briefs is one of the most valuable assets a one-person operation can have.

This week, set yourself a limit of one paragraph for every brief. When you exceed it, stop and ask which of the two explanations applies: too many tasks, or not enough decisions. Fix that instead of writing more. The paragraph will thank you, and so will the agent.

Distil the thinking into one paragraph Scratch notes options considered worries expressed half-decisions project history asides… distil One paragraph ~100 words Goal export saved items as a sheet Context app/account, items query Do not add deps or change the query Done when opens, one row each, test Won't compress? two likely reasons Several tasks split it Something undecided decide it first or A long brief is often a short brief that has not finished being written.
Fig 26 · The One-Paragraph Brief. Scratch thinking distils into a four-part paragraph, or reveals a split or a decision.
Chapter 27 · Part III

Flag the Deviation

No brief survives contact with the work entirely intact. The agent opens the file and discovers that the function you named does not exist any more. The library you asked it to use does not support the format you need. The constraint you set, quite reasonably, makes the task impossible as specified. At that moment the agent has a choice: quietly do something different, or tell you. You want it to tell you, every time, and you have to ask for that explicitly.

Left to their own devices, agents tend to adapt. That is generally a strength; you do not want an agent that stops at every small surprise. But adaptation becomes a problem when it crosses a line you cared about. If the agent decides on its own to add a dependency, change an interface, skip a check that was failing or reinterpret the goal, you will discover it only at review time, if you discover it at all. The work will look complete. It will simply be a different piece of work from the one you asked for.

The fix is a single standing instruction, which belongs in every brief or, better, in the project's memory: if you need to depart from the brief, flag it clearly at the top of your report, explain why, and where the departure is significant, stop and ask before proceeding. It is a small line that changes the shape of the collaboration. The agent still adapts to small surprises, but the significant departures arrive as questions rather than as faits accomplis.

A silent deviation is a decision someone else made for you.

What counts as significant? That depends on the project, but some departures should always be flagged: changing anything listed as a constraint, adding or removing dependencies, altering a public interface, skipping or weakening a test, changing the goal's interpretation, or touching files outside the stated scope. You can list these in your memory file once and refer to them forever.

When a deviation is flagged, treat it as a decision point, not an annoyance. Sometimes the agent is right and the brief was wrong: the function was renamed, the library cannot do it, the constraint was based on an old assumption. Approve the departure and, if it matters, update the brief or the memory so the same surprise does not recur. Sometimes the agent is wrong: it has misunderstood the constraint or taken a shortcut. Redirect it. Either way, you made the call, and the work that comes back is the work you chose.

Over time, flagged deviations become one of your best sources of information about a project. They tell you where the brief and the reality diverged, which usually means your mental model of the project is out of date somewhere. An operator who reads deviations carefully learns more about their own system than one who simply reviews diffs.

This week, add the deviation instruction to every brief, or put it in your project memory once. Then watch what comes back. You will see the agent's judgement in a new way: not hidden inside the work but laid out in front of you, where you can agree with it, overrule it, or learn from it.

Departures arrive as questions Operator Agent brief + limits + flag rule Surprise found function renamed deviation, at the top: why You decide make the call approve, or redirect the work you chose Always flag list once in memory a constraint changed a dependency added a dependency removed a public interface a test weakened the goal reinterpreted files out of scope A silent deviation is a decision someone else made for you.
Fig 27 · Flag the Deviation. A deviation is flagged at the top of the report so the operator makes the call.
Chapter 28 · Part III

Examples Beat Adjectives

Adjectives are the weakest words in a brief. Make it clean. Make it professional. Make it friendly but not too casual. Make it modern. Every one of these words means something to you and something slightly different to the agent, and the gap between the two meanings is where disappointing work comes from. The cure is to replace adjectives with examples wherever you can.

An example carries far more information than an adjective. Friendly but not too casual could describe a thousand tones. Pasting two emails you have written that strike the right note describes exactly one. Clean layout is a vague aspiration. Pointing at a page whose layout you admire is a specification. The agent can see the example, notice its features, and reproduce them, including features you would never have thought to describe because you were not consciously aware of them.

Think of the possible ways to ask for something as falling on two axes. One is precision: how specific the request is. The other is showing versus telling: whether you describe what you want or point at an instance of it. Vague telling is where adjectives live, and it produces the most variable results. Precise showing, a real example with a note about what to keep, produces the most reliable ones. That corner is where you want your briefs to sit.

You cannot describe a taste. You can share a sample.

There are a few ways to use examples well. Point at the example and say what to take from it: match the structure and tone of this report, but not its length. Give more than one example when you want the agent to infer a pattern rather than copy a single instance. Give a counter-example when there is a common failure you want to avoid: not like this one, which is too formal. And when you have no example, ask the agent to produce two or three short options first, then choose one and use it as the example for the real work.

Examples work beyond writing and design. For data, show a sample of the output format you want. For code, point at an existing function written in the style you prefer. For research, show a summary you found useful and ask for more like it. In each case, the example does the job that a paragraph of adjectives would attempt and fail at.

It is worth building a small collection of examples you return to often: a few pieces of writing in your voice, a few designs in your style, a few reports in your preferred shape. Keep them somewhere easy to reach. When you write a brief, reach for them first. Over time they become a kind of portable taste, a way of transferring your standards to an agent in seconds rather than describing them anew each time.

This week, find every adjective in your next three briefs and ask whether an example could replace it. Where one can, swap it in. The work that comes back will feel more like yours, because, in a small but real way, it will be built from pieces of yours.

Examples beat adjectives Precise Vague Telling Showing A spec list precise telling long and still imperfect misses what you never noticed you wanted A real example plus what to keep match this report: structure and tone, not length plus a counter-example Adjectives vague telling clean, modern, professional friendly but not too casual the most variable results A loose reference vague showing something like that site copies the wrong features no example? ask for 2 or 3 short options, pick one
Fig 28 · Examples Beat Adjectives. Precision against showing: a real example with notes beats a pile of adjectives.
Chapter 29 · Part III

Briefs That Outlive the Session

Some briefs are one-offs. Many are not. If you find yourself writing roughly the same brief for the third time, preparing a weekly summary, reviewing a pull request, drafting a client update, cleaning a dataset, then the brief is no longer a brief. It is a recipe, and recipes deserve to be written down properly and kept.

Most agent tools now offer some way to save and reuse instructions: saved prompts, custom commands, skills, templates or project files that an agent can load on request. The names and mechanics vary and will keep changing. The principle does not. A recurring task should have a recurring brief, stored somewhere you and the agents can find it, refined each time it is used.

The trick is not to write the recipe too early. A recipe written after one run is a guess about what will matter. A recipe written after three or four runs by hand captures what you actually learned: the context that kept being needed, the constraint you kept having to add, the check that caught the problem twice. So the cycle runs in a particular order. Do the task by hand a few times with ordinary briefs. Then distil what worked into a recipe. Then reuse it, noticing where it falls short, and feed those notes back into the next version.

Do it three times before you write it down. Then never write it again.

A good recipe looks like a good brief with the variable parts marked. The goal is fixed, or nearly so. The context is mostly stable, with a slot for the specifics of this run. The constraints are fixed. The definition of done is fixed. When you use it, you fill in the slots, and everything else comes for free. Many operators keep recipes as plain text files with a short header saying when to use them, which is a format that will outlast any particular tool.

Recipes also carry your standards forward. Once a recipe includes the check that catches a common mistake, every future run includes it, whether or not you remember to ask. This is one of the quieter ways an operation improves over time. Each mistake caught becomes a line in a recipe, and each recipe makes the next run a little safer than the last.

Be careful of recipe rot. A recipe written six months ago may refer to files that have moved, tools that have changed or standards you no longer hold. When a recipe produces a poor result, check the recipe before blaming the agent. Often the fix is a line or two, and making it improves every future run. A short review of all your recipes once a quarter, deleting the unused ones and updating the stale, keeps the collection honest.

This week, look back over your recent briefs and find one you have written three or more times. Turn it into a recipe: goal, context with slots, constraints, done. Save it where you will find it. Next time the task comes up, use the recipe instead of writing from scratch, and see how much of your attention it returns to you.

From brief to recipe Run by hand 3 or 4 times ordinary briefs Write the recipe fixed parts + slots what you learned Reuse it fill in the slots notice shortfalls notes from each run improve it recipes/weekly-summary.md when: every Friday goal: fixed context: {this week's projects} do not: fixed done when: fixed Recipe rot the six-month problem files moved, tools changed check the recipe before blaming the agent quarterly: prune, update Do it three times before you write it down. Then never write it again.
Fig 29 · Briefs That Outlive the Session. Run a task by hand a few times, write it as a recipe with slots, then reuse and refine.
Chapter 30 · Part III

The Library of Recipes

Over months, a working operator accumulates a set of recipes: reusable briefs for the tasks that come up again and again. At first they live wherever they were written, scattered across folders and threads. At some point it becomes worth gathering them into one place, a library, and treating that library as one of the central assets of the operation. It may well be the most valuable thing you own that is not a client relationship.

Why does a library matter more than the individual recipes? Because it changes how you think about new work. When a task arrives, your first move is no longer to write a brief. It is to check the library. Often there is a recipe that fits exactly, or one that fits with a small change. You fill in the slots and send it. The task gets the benefit of every lesson that went into the recipe, and you spend your attention on the parts that are actually new.

A library has a natural shape. Recipes cluster around the main activities of the operation: reviewing work, releasing it, researching questions, drafting documents, maintaining tools. You might have a dozen recipes in total or several dozen, but they usually fall into four or five families. Organising by family makes them easier to find and shows you where the gaps are. If you have six drafting recipes and no review recipes, that tells you something about where your standards are written down and where they are only in your head.

A library of good recipes is the operation, written down.

Keep the library in plain, portable form. Text files in a folder, versioned if possible, with a short index listing each recipe and when to use it. Tools come and go; plain text survives. If your tools support loading recipes as commands or skills, by all means wire the library into them, but keep the source in a form you could move to a different tool tomorrow without loss.

Look after the library as you would any important asset. When a recipe is used, note whether it worked. When it falls short, fix it straight away while the problem is fresh. When two recipes overlap, merge them. When a recipe has not been used in months, retire it to an archive. A library that is curated stays useful. A library that only grows becomes a junk drawer, and nobody consults a junk drawer when they are busy.

There is one more benefit worth naming. A library of recipes is how a one-person operation can be understood by someone else. If you ever bring in help, a collaborator or a contractor, the library is the fastest way to show them how the work is done. It is also how you will understand your own operation after a holiday, when you have forgotten most of the details and need to pick them up again quickly. The library remembers on your behalf.

This week, create the library if you do not have one: a folder, an index file, and whatever recipes you already have, moved into it. Sort them into families. Note the gaps. Then, the next time a task arrives, open the library first. That small reflex is the beginning of an operation that improves itself.

The library, sorted into families Recipe library index.md, plain text Review a gap: standards in your head Release release-notes ship-site Research compare-tools source-check Drafting client-update weekly-digest proposal cold-email Tools dep-update memory-prune used: note it · short: fix it now · overlap: merge · stale: archive A library of good recipes is the operation, written down.
Fig 30 · The Library of Recipes. A recipe library sorted into families shows where your standards are not yet written.
Part IV

Memory and Notes

What the machine should already know.

Chapter 31 · Part IV

Memory Is Infrastructure

Agents begin most sessions knowing nothing about you. They do not remember yesterday's conversation, the decision you made last week or the reason the build script has that strange line in it. Whatever continuity your operation has, you have to supply. That makes memory, the written record of what the machine should already know, not a nice extra but infrastructure, in the same sense that plumbing is infrastructure. Nobody admires it. Everything stops working without it.

Most agent tools now offer some form of persistent memory: a project file the agent reads at the start of every session, a store of facts it can update, instructions that apply to every conversation in a workspace. The details vary by tool and will keep changing. What they share is the idea that some knowledge should be loaded automatically, so that you do not have to repeat it in every brief. Used well, this is the single biggest improvement in agent quality available to an operator. Used badly, it is a quiet source of confusion.

Think of memory as a stack of layers. At the bottom, plain text files: the most durable, most portable form, readable by any tool and by you. Above them, notes and logs: your running record of what happened and why. Above those, project memory: the curated set of facts and rules that each project's agents load every time. And at the top, your judgement, which decides what goes into each layer and what gets taken out. The project memory layer is where most of the leverage is, because it is what the agents actually see.

What you do not write down, you will explain again. And again.

Treating memory as infrastructure changes a few habits. You give it maintenance time, rather than updating it only when something goes wrong. You review changes to it with the same care you would give to a change in code, because a bad line in memory affects every future session. You keep it in plain text where you can, because infrastructure should outlive the tools built on top of it. And you design it for its readers: the agents first, and you second, when you return to a project after months away.

There is also a useful mental shift. Every time you find yourself explaining the same thing to an agent twice, that is a memory bug. The fix is not to explain it a third time but to write it down in the right layer so that it is known from now on. Every time an agent makes a mistake because it did not know something, that is a memory bug too. Collect these. They are the work orders for your memory infrastructure.

This week, pick your most active project and look at what its agents load at the start of a session. If the answer is nothing, start a memory file. If there is already one, read it as if you were a new agent, and mark anything that is wrong, stale or missing. Then fix those. It will feel like housekeeping. It is closer to laying pipe, and the water will flow better for months.

Memory as a stack of layers Your judgement decides what goes in and out Project memory what the agents actually see Notes and logs your running record of why Plain text files durable, portable, any tool top base Memory bugs your work orders the same thing explained twice; a mistake from not knowing a fact Treat it as plumbing scheduled upkeep review edits like code plain text first written for its readers What you do not write down, you will explain again. And again.
Fig 31 · Memory Is Infrastructure. Memory as layers, with project memory the one agents read and memory bugs as work orders.
Chapter 32 · Part IV

What the Agent Should Already Know

A project memory file is a short document that an agent reads at the start of every session in that project. Its job is to answer, in advance, the questions any capable newcomer would ask on their first morning. What is this project for? How do I build it, test it and run it? What are the rules here? Where are the traps? If those four questions are answered well, most sessions start in the right place without you having to say anything.

Start with purpose. One or two sentences on what the project is and who it serves. This sounds unnecessary, since the agent can read the code, but code tells you what something does, not what it is for. Knowing that a tool serves a handful of internal users rather than the public changes a great many small decisions, from error messages to how much care goes into performance.

Then commands. The exact commands to install, build, test, lint and run the project, and anything unusual about them. This is the most practically useful part of the file, because without it agents guess, and guessed commands are a frequent source of wasted time. If the test suite needs a particular environment variable or a running database, say so here.

Then rules. The conventions that are not obvious from the code: how things are named, where new files go, what the house style is for writing, which patterns are preferred and which are being phased out. Keep these to the ones that actually matter and that an agent would plausibly get wrong. A memory file with fifty rules is not followed better than one with ten. It is usually followed worse, because the important rules are diluted by the trivial ones.

The memory file is the induction you give every new hire, every morning, for free.

Finally, pitfalls. The places where things go wrong: the fragile module, the misleading function name, the test that is flaky on Tuesdays, the folder that looks unused and is not. Pitfalls are where memory pays for itself most clearly, because each one prevents a specific, recurring mistake. Many of the best lines in a mature memory file began as a note in an incident write-up.

Keep the file short. A page or two is plenty for most projects. Agents read the whole thing every session, so everything in it costs a little attention every time, and irrelevant material can nudge work in odd directions. If the file grows long, split it: a core file that is always loaded, and topic files that are pointed to when relevant. Many tools support this kind of layering directly.

Write it in plain, direct sentences, the same register as a good brief. Avoid aspirational statements that no one can check, such as write high-quality code. Prefer specific ones, such as every new function has a test in the matching file under tests/. The agent can follow the second. It can only nod at the first.

This week, write or rewrite the memory file for one project using the four headings: purpose, commands, rules, pitfalls. Keep it under two pages. Then start a fresh session and watch how much less you have to explain. That silence is the file doing its job.

The induction, every morning, for free memory.md one or two pages PURPOSE internal tool for six staff not public: plain errors are fine What is this for? code says what, not why COMMANDS test: npm test (needs the DB) lint: npm run lint How do I build and test? or the agent guesses RULES each new function: a test dates are stored in UTC What are the rules here? ten, not fifty PITFALLS billing/legacy.ts is fragile sync test flaky on Tuesdays Where are the traps? each prevents a repeat Write specific rules the agent can follow, not hopes it can only nod at.
Fig 32 · What the Agent Should Already Know. A project memory file answers purpose, commands, rules and pitfalls before anyone asks.
Chapter 33 · Part IV

The Operator's Notebook

Project memory files are for the agents. You need something for yourself as well: a running notebook in which you record what you did, what you noticed and what you are thinking about. It is the least glamorous tool in the operator's kit and one of the most powerful, because it is the only place where the whole operation, across every project, is visible in one stream.

The format matters less than the habit. A single text file with a dated heading for each day works perfectly well. So does a paper notebook, if you prefer. What matters is that it is one place, that you write in it every working day, and that you can search or skim it later. Many operators keep it open all day and jot a line whenever something happens: a decision made, a surprise encountered, an idea parked, a thread started or finished.

What goes in? Mostly short lines. Shipped importer fix; empty rows now handled. Agent kept reaching for the old config; added a note to memory. Client wants the report a week early; moved it up. Idea: weekly digest for overdue threads; queued. None of these is a work of literature. Together they form a record of the operation that is enormously useful in ways you cannot predict when you write them.

The notebook does not remember for you. It remembers instead of you.

The notebook earns its keep in three ways. First, it feeds the close and the review. At the end of the day, skimming today's entries tells you what shipped and what is open far faster than reconstructing it from threads. Second, it feeds the weekly and quarterly reviews, which a later part of the book covers. Looking back over a week of entries reveals patterns that are invisible day to day: the project that keeps stalling, the kind of task that keeps going wrong, the hour of the day when nothing good happens. Third, it is a record you can search when something comes back. When a client asks why a change was made in March, the notebook often has the answer in one line.

The chain that matters is: note it, re-read it, decide. A notebook that is written but never read is a diary. Useful, perhaps, but not an operating tool. Build the re-reading into your rhythm: a skim at each close, a proper read at each weekly review. Each read should end in at least one decision, even if it is only add this pitfall to the memory file or stop starting threads after four o'clock.

Agents can help here too, but carefully. An agent can summarise a week of notebook entries, extract recurring themes or draft a list of open items. That is useful. What an agent should not do is write the notebook for you, because the act of writing a line is itself a small moment of reflection, and that moment is part of the value.

This week, start the notebook if you do not have one. One file, one heading per day, one line per thing that happened. At Friday's close, read the week's entries from top to bottom and write one decision at the bottom. That is the whole practice. It compounds quietly, the way good habits do.

One stream for the whole operation ONE LINE EACH a decision made a surprise an idea parked a thread done Operator's notebook one file, dated headings shipped importer fix old config: to memory report moved earlier idea queued: digest Daily close skim what shipped Weekly review patterns that stall Search later why, back in March? Note it a line, as it happens Re-read it at close and weekly Decide at least one per read written but never read is a diary, not a tool The notebook does not remember for you. It remembers instead of you.
Fig 33 · The Operator's Notebook. Notebook lines feed the close, the weekly review and later search, and end in decisions.
Chapter 34 · Part IV

A Decision Log

Decisions are the most valuable thing an operator produces and the least likely to be recorded. Code is versioned. Documents are saved. Pull requests have descriptions. But the reasons behind them, why this approach and not that one, why this project and not the other, why we dropped the feature, usually live only in your head, where they decay at an alarming rate. Three months later you look at something and genuinely cannot remember why it is the way it is.

A decision log fixes this cheaply. It is a single file, per project or for the whole operation, in which you record significant decisions as you make them. Each entry is short: the date, the decision in one sentence, the reason in one or two, and the main alternative you rejected. That is all. A few lines that will save you hours.

Why record the rejected alternative? Because it is the part you will most need later. When you or an agent revisit a decision, the first question is usually why didn't we just do the obvious other thing? If the log says considered using the hosted search service; rejected because results needed to work offline, the question is answered in seconds and the old debate does not have to be reopened. Without that line, you may well spend an afternoon rediscovering the reason, or worse, reverse the decision without remembering that there was one.

Recording what you decided is useful. Recording why is the part that saves you.

What counts as significant? A rough test is whether someone could reasonably ask about it later. Architectural choices, changes in scope, dropping or parking a project, choosing a tool, changing a standard, accepting a known risk. Small day-to-day choices do not need entries; that is what the notebook is for. If you are logging more than a few decisions per project per week, you are probably logging too much.

The decision log also helps the agents. If it lives alongside the project, you can point agents at it when they work on something affected by an earlier decision, or include a short summary in the project's memory. An agent that knows the search service was rejected for offline reasons will not suggest it again, and will not quietly reintroduce it while fixing something else. This is one of the few places where giving agents history, rather than just current state, genuinely improves their work.

There is a further benefit that is easy to overlook. Writing down the reason forces you to have one. Operators sometimes discover, when they sit down to log a decision, that their reasoning is thinner than they thought. That is not a reason to stop logging. It is the log doing its most useful work, before the decision has had any consequences.

This week, start a decision log for your busiest project. Backfill the three most significant decisions you can remember from the last month, with their reasons and rejected alternatives. Then add entries as new decisions arrive. In three months, when someone asks why, you will have the unusual pleasure of simply knowing.

Record why, and what you rejected You, in March decides Decision log one file You, in June or an agent decision + reason date: 3 March did: local search why: works offline not: hosted search why not the obvious? answered in seconds Debate stays shut log it if someone could ask later: architecture, scope, tools, risk
Fig 34 · A Decision Log. A decision log entry with its reason and rejected option answers a question months later.
Chapter 35 · Part IV

Notes That Get Read

Operators write a great deal: briefs, handoff notes, memory files, notebook entries, decision logs, incident write-ups. The question that matters about all of it is not how much was written but how much was read again, and how much of what was read led to a better decision. Picture it as a funnel. A lot goes in at the top. Less is read again. Less still changes anything. The goal is not to write less, necessarily, but to make the funnel narrower at the top and wider at the bottom: fewer notes, more of them useful.

Notes that get read have a few things in common. They are written for a specific reader at a specific moment: tomorrow's morning review, the next session on this project, an agent starting work, you in three months. Knowing who will read it and when changes what you write. A handoff note for tomorrow can be terse. A decision log entry for three months from now needs its reasons spelled out, because by then the context will have evaporated.

They lead with the conclusion. The most important line is the first one: what is the state, what was decided, what needs doing. Detail follows, for anyone who needs it. Notes that begin with narrative and end with the point are rarely read to the end, which means the point is rarely read at all.

Write the note for the tired person who will read it. That person is usually you.

They are in a predictable place. A note that cannot be found does not exist. Decide where each kind of note lives, the notebook in one file, handoffs at the top of the project memory, decisions in the log, and stick to it. If you have to search for a note, you will often not bother, and the note will have been wasted.

They are short. Long notes get skimmed, and skimmed notes lose their details. If a note needs to be long, give it a one-line summary at the top that could stand alone. Many operators adopt a simple rule: no note longer than a screen without a summary line.

And they are pruned. A handoff note that has been superseded should be removed or marked as old, or the next reader will act on stale information. A memory file full of outdated advice is worse than a short one with current advice. The next chapter takes up pruning in more detail, but it belongs here as well, because notes that are never pruned are notes that eventually stop being trusted.

There is a way to test your notes. Once a week, pick one note you wrote at least a fortnight ago and read it cold. Ask whether it told you what you needed to know quickly, whether anything in it was wrong or out of date, and whether it led you to do anything. If the answer to all three is good, keep writing that kind of note. If not, adjust.

This week, look at the last ten notes you wrote of any kind and ask of each: who was this for, and did they read it? If you cannot say, the note was probably written for nobody. Write fewer of those and more for somebody in particular.

Notes that get read Everything written briefs, handoffs, memory, logs Read again by its reader, on time Acted upon a better decision Read notes are five habits for a named reader conclusion first in a set place short, summed up pruned when stale Weekly test read one fortnight-old note cold: quick? true? acted on? Write the note for the tired person who will read it. Usually you.
Fig 35 · Notes That Get Read. Of everything written, less is read and less acted on; five habits widen the bottom.
Chapter 36 · Part IV

Pruning the Memory

Memory accumulates. Each time an agent makes a mistake, you add a rule. Each time a new convention emerges, you note it. Each time a pitfall bites, you record it. All of this is good practice, and none of it is ever removed, because removing things feels risky and adding things feels responsible. After six months the memory file is three times the length it was, a third of it is out of date, and the agents are following instructions written for a project that no longer quite exists.

Stale memory is worse than no memory. An agent with no memory asks questions or explores. An agent with stale memory acts confidently on wrong information. It uses the old command, follows the abandoned convention, avoids the pitfall that was fixed months ago by working around it in an unnecessarily complicated way. The errors are subtle because they look like deliberate choices, and in a sense they are: they were your deliberate choices, once.

The remedy is a cycle: add, review, prune. Adding happens naturally in the course of work. Reviewing and pruning need to be scheduled, because they will never happen spontaneously. Once a month, or at each quarterly review at the very least, read every memory file in full with a red pen in mind. For each line, ask whether it is still true, whether it is still needed, and whether it is in the right place.

Memory that only grows becomes a museum. Agents should work in a workshop.

Some lines will be simply wrong: commands that changed, files that moved, rules that were abandoned. Fix or delete them. Some will be true but no longer necessary: a pitfall that has been fixed at the source, a convention that is now enforced by tooling. Delete them; the tooling remembers on your behalf. Some will be true and necessary but in the wrong place: a project-specific rule in a general file, or a one-off instruction that has been promoted to a permanent one. Move them.

A useful trick is to ask an agent to help. Give it the memory file and the current state of the project and ask it to identify lines that appear outdated, contradictory or unsupported by what it can see. It will not catch everything, and it will flag a few things that are actually fine, but it is a good first pass, and it is particularly good at catching contradictions between rules written months apart.

Pruning also keeps memory short, which matters for its own sake. Every line in a memory file is read by every session. A file that has been pruned to its essential half is not just more accurate but more effective, because the remaining rules are not competing for attention with dead ones.

This week, take your largest memory file and prune it. Aim to remove at least a fifth of it. Read every line, delete what is wrong or unnecessary, move what is misplaced. Then run a normal session and see whether anything goes wrong. In most cases nothing will, and you will have learned something useful: a good deal of what you were carrying was weight, not knowledge.

Every line, three questions Each line in memory Still true? no Fix or delete command changed yes Still needed? no Delete it tooling enforces yes Right place? no Move it wrong file yes Keep it true, needed, placed Add during work Review monthly, scheduled Prune aim for a fifth Memory that only grows becomes a museum.
Fig 36 · Pruning the Memory. Each memory line faces three questions on a scheduled add, review and prune cycle.
Chapter 37 · Part IV

One Source of Truth

Every fact in your operation should have exactly one home. The test command lives in one place. The house style lives in one place. The current status of each project lives in one place. When a fact lives in two places, the two copies will eventually disagree, and when they do, you and your agents will waste time working out which is right. Often you will not even notice the disagreement until something goes wrong because of it.

Duplication creeps in for innocent reasons. You put the build instructions in the memory file for the agent and in the readme for humans. You keep a project list in your notebook and another in a spreadsheet. You note a decision in the decision log and also in the pull request description, then later update one but not the other. Each copy was helpful when made. The trouble is that facts change, and you will only ever update the copy you happen to be looking at.

The fix is to choose a home for each kind of fact and to point everywhere else. The memory file says see the readme for build commands rather than repeating them, or the readme says see the memory file, depending on which you maintain more carefully. The project register is the single place where project status lives; everything else links to it. The decision log is where decisions are recorded; the pull request description refers to the log entry.

Two copies of a fact are one fact and one future bug.

Pointing has a cost, which is that the reader has to follow the pointer. For agents this cost is usually tiny; they can open a file in an instant. For you it is a little larger, but far smaller than the cost of acting on a stale copy. The only real exception is a short summary of something that lives elsewhere, which is sometimes worth keeping for convenience. If you do keep one, label it clearly as a summary and say where the source is, so that anyone in doubt knows which to trust.

One source of truth matters especially in a world of many agents. If three threads are working on the same project and each has a slightly different idea of the conventions because they were briefed from slightly different copies, you will get three slightly different kinds of work. Pulling the conventions into one memory file that every thread loads is the simplest way to make the fleet behave consistently.

It also matters for you, because duplication is a tax on your memory as well as the agents'. If you have to remember that the status is in the notebook and also in the register and also in the spreadsheet, you will forget one. If the status is only in the register, there is nothing to remember except where the register is.

This week, pick one fact you know is written in more than one place and consolidate it. Choose the home, update it, and replace the other copies with pointers. Then do the same for one more next week. It is unglamorous tidying, and it removes a whole category of confusion from the operation, permanently.

Every fact gets one home Before: copies After: one home, pointers memory file test: npm test readme test: make two copies drift apart register status: active spreadsheet status: done which one is right? you update only the copy you happen to be looking at memory file see readme readme test: npm test notebook see register spreadsheet removed register status: active one place to update; agents follow pointers instantly Two copies of a fact are one fact and one future bug. a summary is fine if it is labelled and names its source
Fig 37 · One Source of Truth. Facts copied in several places drift; one home plus pointers keeps them consistent.
Chapter 38 · Part IV

Secrets Stay Out

Memory files, notebooks, briefs and logs are designed to be read widely. Agents read them every session. They get committed to repositories, synced to cloud storage, pasted into threads and occasionally shared with collaborators. That is exactly what makes them useful, and exactly why they must never contain secrets. Memory is not a vault.

By secrets, think broadly. Passwords and access keys, obviously. But also tokens, private connection strings, personal details about clients or users, anything covered by a confidentiality agreement, and anything you would be uncomfortable seeing appear in a public search result. The rule is simple: if it would cause harm in the wrong hands, it does not go into anything an agent reads by default.

A helpful way to think about any piece of information is to place it on two axes. One is sensitivity: how much harm its exposure would cause. The other is reach: how many places and readers it will end up in. Memory files and briefs have very high reach. Anything sensitive that lands in a high-reach document is in the wrong quadrant. Keep it out, and put it somewhere designed for sensitive material instead.

If the agent can read it every morning, assume the world can read it eventually.

The practical alternatives are well established. Credentials belong in a proper secrets store, environment variables or the secret management that your tools provide, and agents are given access to them through those mechanisms, never by pasting the value into a brief. Personal data about clients belongs in whatever system you use for client records, and when an agent needs to work with it, it should get only the minimum required for the task. Confidential material should be referred to by location rather than copied: the contract terms are in the client folder rather than the terms themselves.

It is also worth thinking about what agents might write into memory on their own. Some tools let agents save facts they learn during a session. That is useful, but it means a value seen once during debugging could be saved permanently. Review automatic memory periodically, and include a standing instruction that agents must never store credentials, tokens or personal data in memory or notes. Most agents will honour this reliably once told. Some will need reminding.

If a secret does end up somewhere it should not, treat it as an incident rather than a tidying job. Rotate the credential, not just delete the line, because a secret that has been in a repository or a synced file should be assumed to have been seen. Then write a short incident note and add whatever guardrail would have prevented it. The same principle appears again in the chapter on never committing secrets, because it is one of the very few rules in this book that has no exceptions.

This week, search every memory file, notebook and recipe you own for anything that looks like a key, a password, a token or someone's personal details. Remove what you find and rotate anything that was a live credential. Then add the standing instruction to your memory. It takes half an hour and closes a door that would otherwise stay quietly open.

If the agent reads it daily, assume the world will Harm if exposed low Reach: few readers many readers Secrets store where secrets belong keys in env vars or a vault client data in its system agents get access, not values Keep it out the wrong quadrant a token pasted in a brief client details in memory a password in a notebook Scratch rarely matters throwaway, low stakes Memory, briefs, notes high reach by design read every session synced, committed, pasted refer by location instead leaked? rotate it, do not just delete the line
Fig 38 · Secrets Stay Out. Sensitive material in high-reach notes is in the wrong quadrant and belongs in a store.
Chapter 39 · Part IV

Session Handoffs

Every session ends and, for any substantial piece of work, another session will pick it up later. That might be you tomorrow, a fresh agent thread, or a parallel thread that needs to know what this one did. The quality of that pickup depends almost entirely on one small document: the handoff. Write it well and the next session starts in minutes. Skip it and the next session starts with archaeology.

A good handoff answers four questions. What was the goal of this session? What is the current state, specifically, including what was finished and what was not? What is the next move? And what should the next session know that it would not otherwise discover: a dead end already explored, a decision made, a surprise found? Four short paragraphs, or four lines, depending on the size of the work.

The most important of the four is the next move. A handoff that describes the state beautifully but does not say what to do next leaves the next session to work it out, which wastes time and risks a different conclusion. Next: add the test for the empty case, then run the full suite is worth more than a page of description, because it lets the next session begin acting immediately.

The best handoff is the one that lets the next session start with a verb.

The second most important is the dead ends. Agents, and people, are prone to rediscovering the same wrong path. If this session spent half an hour discovering that the obvious fix does not work because of some hidden constraint, say so in the handoff. Otherwise the next session will spend its own half hour on the same discovery, and may not even reach the same conclusion.

Where does the handoff live? Somewhere the next session will look without being told. For many operators that is a short section at the top of the project's memory file, replaced at the end of each session. For others it is a note in the project register next to the project's line. For work in a repository, it might be the pull request description, updated as the work progresses. The location matters less than consistency: pick one and always use it.

Agents are good at drafting handoffs, and you should let them. At the end of a session, ask the agent to write the handoff using the four questions. Then read it and correct it. The correction step is not optional. Agents tend to be optimistic about state, describing something as nearly finished when it is half finished, and they do not always know which of their discoveries will matter later. Your edit is where the handoff becomes reliable.

This week, end every session on a multi-session piece of work with a four-question handoff. Put it in the same place every time. Then notice how the next session goes. You will likely find that the first ten minutes, which used to be spent working out where things were, are now spent doing things. The handoff is a small courtesy to your future self, and your future self is a demanding client.

A handoff, written and corrected This session the agent You operator Next session any thread Draft it 4 questions Correct it fix optimism File it same place Start work with a verb Goal of this session State done and not done Next move most important Dead ends do not rediscover
Fig 39 · Session Handoffs. A handoff drafted by the agent, corrected by you and read by the next session.
Chapter 40 · Part IV

The Filing Cabinet Mind

There is a popular idea that the solution to information overload is a second brain: a vast, linked, lovingly tagged personal knowledge system in which everything you have ever read or thought is stored for later. Some people build these and find them genuinely useful. Many more build them and find they have created a beautiful archive that they never consult. For an operator, a better model is humbler: not a second brain but a filing cabinet.

The difference is in what you optimise for. A second brain optimises for capture: getting everything in, connecting it, making it rich. A filing cabinet optimises for retrieval: being able to put your hand on the right thing at the right moment. Capture without retrieval is hoarding. Retrieval without capture is impossible. What you want is the overlap, the information you captured that you can actually find and use when you need it. That overlap is usually much smaller than the archive.

Operators with agents have a particular reason to favour retrieval. Agents are very good at searching, reading and summarising, which means that well-organised plain files are now dramatically more useful than they used to be. You do not need to tag and link everything by hand if an agent can search the folder and pull out what is relevant. What you do need is for the files to be in predictable places, named sensibly and written clearly enough that a search finds them.

A note you cannot find is a note you did not write.

So keep the cabinet simple. A small number of top-level folders that match how you think about your work: projects, recipes, decisions, reference, archive. Inside each project folder, the same few files every time: memory, register entry, decision log, handoffs. Names that a search will find. Plain text wherever possible. And a firm habit of archiving what is finished, so that the active drawers contain only what is active.

Capture selectively. Not everything needs to be kept. A useful test is whether you can imagine a specific future moment when you would want this: a client question, a recurring task, a decision to revisit. If you can, file it where that moment will look. If you cannot, let it go. The fear of losing something is real, but most of what operators save is never opened again, and the clutter it creates makes the useful things harder to find.

Retrieval improves with practice. When you need something, try to find it before asking an agent to search. Notice where you looked first. That is where your mind expects it to be, and if it was somewhere else, consider moving it. Over time, the cabinet reshapes itself around how you actually think, which is the only organisation scheme that reliably works.

This week, ask an agent to find three things in your files that you know exist: a decision from last quarter, a recipe you rarely use, an old client note. Time each search. Where the search was slow or failed, ask why, and fix the filing. You are not building a brain. You are building a cabinet that opens at the right drawer, and that is a far more achievable thing.

A cabinet, not a second brain projects memory, log, handoffs recipes reusable briefs decisions the why reference things you look up archive finished, out of the way Capture Retrieval Useful the overlap is small Capture test before filing can you picture the moment you'd want it? else let it go Search yourself first then move what was misplaced A note you cannot find is a note you did not write.
Fig 40 · The Filing Cabinet Mind. A simple filing cabinet of drawers, built for retrieval rather than capture.
Part V

The Register and the Threads

Projects, delegation and parallel work.

Chapter 41 · Part V

The Project Register

An operator with agents can have a remarkable number of things in flight. A client project, two internal tools, a book draft, a research question, a migration that has been ninety per cent done for three weeks, an idea you started on a Sunday. Each one may have several threads running against it. Without a single place that lists all of it, the operation becomes impossible to see whole, and what cannot be seen whole cannot be steered.

That single place is the project register. It is a list of every project currently in your life, with one line for each, a status and a next move. It is not a project management system, a task tracker or a kanban board, although it can live in any of those if you like. It is the top-level view: the answer to the question what am I actually running?

The register sits at the top of a small hierarchy. Beneath it are the projects, one line each. Each line points to a next move, the single most important thing that should happen on that project. And beneath those are the threads and sessions doing the actual work. The register is the part you look at every morning. The threads are the part the agents look at. Keeping the two levels separate is what lets you think about the operation without drowning in its details.

If it is not on the register, it is not a project. It is a distraction with ambitions.

What goes on the register? Anything that will need more than one session, involves a commitment to someone, or will produce something that ships. One-off tasks do not belong; they belong in today's three or the queue. Ongoing responsibilities, such as keeping a tool updated, can go on the register as standing projects with their own simple status.

Keep it in one place, in plain text if you can, and keep it short enough to read in a minute. Many operators keep it as a single file with a heading per status and a line per project. Some keep it as a table. A few keep it on paper by the desk. The format matters far less than the discipline of having exactly one register and updating it at every close.

The register does three jobs at once. It feeds the morning review, because today's three are chosen from it. It keeps you honest about capacity, because the number of active lines is a visible measure of how much you have taken on. And it shows you what is drifting, because a line whose next move has not changed in a fortnight is a line that needs a decision, not more work.

It also lets agents help with oversight. An agent can read the register alongside the threads and tell you which projects have had no activity this week, which next moves are stale and which threads are not attached to any project. That is a useful morning digest, as long as the register itself remains yours to edit.

This week, write your register. Every project, one line, a status and a next move. You may be surprised by the length of the list. That surprise is the first useful thing the register will ever tell you.

The register sits above the threads The register what am I running? Importer rewrite active Ship empty-rows fix next move fixing testing Client report waiting Chase draft two next move no live threads Course outline parked Revisit at quarter next move threads stopped you read the top two levels; the agents work the bottom one feeds the review shows your load shows what drifts
Fig 41 · The Project Register. The register lists each project with a status and next move, above the working threads.
Chapter 42 · Part V

One Line Per Project

The register works only if each project fits on one line. That sounds like a formatting rule. It is actually a thinking rule, because compressing a project into a single line forces you to know what its status is and what should happen next. If you cannot write the line, you do not currently understand the project well enough to steer it.

A good line has four parts: the project's name, its status, its next move and, if relevant, a date. Importer rewrite: active; next, ship the empty-rows fix; by Thursday. Client report: waiting; next, chase feedback on draft two. Course outline: parked; next, revisit at quarterly review. Each line is a sentence you could say aloud in under five seconds, and each tells you exactly what to do if you chose to work on that project now.

The compression is where the value lies. Every project has far more information associated with it than one line can hold: history, context, open questions, risks, ideas. All of that belongs somewhere, in the project's memory file, decision log or handoff notes. The register line is not a summary of all that. It is a distillation of what matters for steering: where is it, and what is next?

If you cannot say it in one line, you do not know it yet.

Writing the next move is the hardest part, and the most useful. It must be a concrete action, not a theme. Work on the importer is a theme. Ship the empty-rows fix is a move. Think about pricing is a theme. Draft two pricing options and pick one is a move. Themes do not tell you what to do this morning. Moves do, which is why a register of moves is something you can run a day from and a register of themes is merely a list of worries.

When a project's line is hard to write, take that seriously. Usually it means one of three things. The project is actually several projects and should be split into separate lines. The project is waiting on a decision you have not made, in which case the next move is to make it. Or the project has lost its purpose and is continuing out of habit, in which case the next move may be to park it or end it. All three are worth knowing.

Agents can draft register lines from a project's handoff notes, and this is a handy shortcut at the close. But edit each line yourself. The next move in particular is a decision, and agents are inclined to propose the most obvious continuation rather than the most valuable one. Sometimes the most valuable next move is to stop, and an agent rarely suggests that unprompted.

This week, rewrite every line on your register in the four-part form: name, status, next move, date. Where a next move is a theme, turn it into an action. Where a line will not compress, find out why. By the end, you should be able to read the whole register aloud in under a minute and know exactly what each project needs. That is the register doing its job.

One line, four parts NAME Importer rewrite STATUS active NEXT MOVE Ship the empty-rows fix DATE Thursday THEME: A WORRY MOVE: AN ACTION Work on the importer Ship the empty-rows fix Think about pricing Draft two options, pick one IF THE LINE WILL NOT COMPRESS Several projects split into lines Waiting on a decision make it: the move Lost its purpose park it or end it If you cannot say it in one line, you do not know it yet.
Fig 42 · One Line Per Project. A register line splits into name, status, next move and date; moves beat themes.
Chapter 43 · Part V

Status Words That Mean Something

Every project on the register has a status, and the vocabulary you use for statuses matters more than you might expect. Too many status words and they blur together: in progress, ongoing, underway, started, in development. Too few and they hide important differences. What you want is a small, honest vocabulary in which each word implies a specific action, and the most useful set for most operators is four words: idea, active, parked and done.

An idea is something you might do but have not committed to. It sits on the register, or more often in a separate list, so that it is not lost, but nothing is expected of it. Ideas are cheap and should stay cheap. The danger with agents is that ideas become active by accident: you start a thread to explore something, the thread produces something promising, and suddenly a project exists that you never decided to start. Keep a firm line between idea and active, and cross it only on purpose.

Active means you have committed to it and it is getting attention this week. Active projects have a next move that is being worked on, and they compete for today's three. The number of active projects is the single most important number on the register, because it is the closest thing you have to a measure of load. Most solo operators can run three to five active projects well. More than that, and each one gets less review than it needs.

A status is a promise about what will happen next. Keep the promises few.

Parked means deliberately paused. It is not failed, not forgotten, not abandoned. It is a project you have chosen not to work on for now, with a note of why and when you will revisit it. Parked is the most underused status in most operations and the most valuable, because it gives you an honest way to reduce load without pretending projects do not exist. The next chapter is about it.

Done means finished and shipped. It is the status every project is meant to reach, and it should be marked clearly, with a date and a line about what shipped. Done projects can move to an archive section of the register, but keep a record. Looking back over what got done in a quarter is one of the more encouraging things an operator can do.

You may want one more status for projects that are blocked on someone else: waiting. That is fine, as long as each waiting project has a named person or event it is waiting for and a date when you will chase it. Without those, waiting becomes a polite synonym for stuck.

Resist adding more. Every new status word adds a decision about which word to use, and every ambiguous word becomes somewhere for projects to hide. If you find yourself wanting a status like mostly done or in progress but slow, that is usually a sign that a project needs a decision rather than a new label.

This week, go through your register and assign each project exactly one of the four words. Where one does not fit, decide which it should be. Count the actives. If the number surprises you, the next chapter will help.

Four honest status words Idea costs nothing Active this week Done shipped, dated Waiting who, until when Parked why, until when Ended kindly not wanted now on purpose never by accident finished pause condition met blocked answered at review count the actives: three to five is what one person reviews well A status is a promise about what happens next. Keep the promises few.
Fig 43 · Status Words That Mean Something. Projects move between idea, active, waiting, parked and done on deliberate transitions.
Chapter 44 · Part V

Parking Without Guilt

Every operator has more projects than they can run well. The natural response is to keep them all nominally active, giving each a little attention, feeling guilty about the ones that get less, and making slow, scattered progress on everything. The better response is to park. Parking is the act of deliberately pausing a project, writing down why and when you will look at it again, and then not thinking about it until then.

The test for parking is simple: can this project make meaningful progress this week, given everything else? If yes, keep it active. If no, because it is waiting on something, because other projects matter more, or because you simply do not have the attention, park it. The answer will often be no for projects you care about. That is fine. Caring about something is not the same as having capacity for it this week.

Parking properly involves three steps. First, write a handoff note, so that when you return you can pick up where you left off without archaeology. Second, write the reason for parking and the condition for unparking: parked until the client confirms scope, or parked until the quarterly review, or parked until the importer is done. Third, stop any threads attached to the project, or leave them in a clean, finished state. A parked project with live threads is not parked. It is neglected.

Parked is a decision. Drifting is the absence of one.

The guilt is the hard part. Parking feels like admitting defeat, particularly for projects you started with enthusiasm. But the alternative, keeping too many projects active, guarantees that some of them will drift, and drifting projects carry far more guilt than parked ones, because their state is uncertain. A parked project sits quietly with a clear note attached. A drifting project nags at you every time you see its name.

Parking also improves the projects that stay active. With fewer lines competing for today's three, each active project gets more review, better briefs and faster finishes. Operators who park aggressively often find that their total throughput goes up, not down, because finishing a few things is far more productive than advancing many.

Agents make parking more important rather than less. Because starting is so cheap, the number of things you could be working on grows constantly. Without a habit of parking, the register swells until it is unreadable. With it, the register stays a manageable size, and the ideas that would have swollen it wait in an orderly queue where they can be assessed calmly.

Review parked projects at each quarterly review, or when their unparking condition is met. Some will come back to active. Some will turn out, in the cold light of a few months, to be projects you no longer want, and these can be ended kindly, which a later chapter covers.

This week, look at your active projects and park at least one. Write the note, state the condition, stop the threads. Then notice how the remaining actives feel. Usually lighter. The parked project will not mind. It has a note.

Park on purpose an active project Can it move this week? yes no Keep it active gets more review 1 · Write the handoff no archaeology later 2 · Reason, condition until scope confirmed 3 · Stop its threads live threads: neglect Parked a decision + note condition met or end it kindly at review Parked is a decision. Drifting is the absence of one.
Fig 44 · Parking Without Guilt. Projects that cannot move this week are parked in three steps until a condition is met.
Chapter 45 · Part V

Threads Are Units of Delegation

Most agent tools organise work into threads: separate conversations or sessions, each with its own context, history and running state. It is easy to think of a thread as a chat window, a place where you happen to be talking to an agent. It is more useful to think of it as a unit of delegation, a bounded job handed to a worker, with a brief, a scope, a set of checks and a handoff at the end.

That shift in thinking changes how you start threads. A chat window is opened casually, whenever a question occurs to you. A unit of delegation is opened deliberately, when there is a piece of work worth delegating and you know what done looks like. The first habit produces dozens of half-used threads, each holding a fragment of context, none clearly finished. The second produces fewer threads, each with a purpose and an end.

A thread as a unit of delegation has four parts, gathered around it like a hub. The brief that started it, which states the goal and done. The scope, which says what this thread may touch and what it may not. The checks it must run before reporting back. And the handoff it leaves when it finishes or pauses. If you can name all four for a thread, it is a well-formed piece of delegation. If you cannot, it is a conversation that might turn into work, which is a different and less reliable thing.

A thread without a brief is a conversation. A thread with one is a job.

Thinking in units also helps you decide when to start a new thread and when to continue an old one. Continue when the work is the same job and the context is still relevant. Start fresh when the job has changed, when the context has grown cluttered with dead ends, or when you want a clean pair of eyes. A long thread accumulates history, and history is not always helpful. Agents can become anchored to an early approach that you have since abandoned. A fresh thread with a clean brief and a good handoff note often outperforms a long thread that knows too much.

Units of delegation are also what make the register work. Each thread should belong to a project on the register. If you find a thread that does not, either it is part of a project you have not registered, which you should fix, or it is a stray, which you should finish or close. A quick audit of open threads against the register is a good weekly habit, and it almost always finds a few strays.

When you describe your work to yourself in terms of threads as jobs, you naturally start thinking like someone who delegates, which is what you are. You ask whether the brief was clear, whether the scope was right, whether the checks were sufficient. Those are the questions that improve delegation. Why did the chat go wrong? is not.

This week, before opening any new thread, write its four parts in a line or two each. If you cannot, do not open the thread yet. Count, at the end of the week, how many threads you opened compared with the week before. Fewer, almost certainly. Better, almost certainly too.

A thread is a job, not a chat BELONGS TO A PROJECT ON THE REGISTER One thread a delegated job Brief goal and done Scope may touch, may not Checks run before reporting Handoff left when it pauses a fresh thread beats one that knows too much Chat window opened casually, half-used Unit of delegation opened on purpose, has an end
Fig 45 · Threads Are Units of Delegation. A thread as a unit of delegation, with a brief, scope, checks and handoff.
Chapter 46 · Part V

One Thread, One Outcome

Each thread should aim at exactly one outcome. Not one outcome and a few related tidy-ups. Not two outcomes that happen to touch the same files. One. This is the single most effective rule for keeping parallel agent work reviewable, and it is broken constantly, usually with the best intentions.

The temptation is efficiency. You have a thread open on the importer, the agent has the context loaded, and you remember that the exporter has a similar bug. Why not ask the same thread to fix that too? It is right there. It knows the code. It will take five minutes. And so the thread now has two outcomes, braided together, and when it comes back you have one large change that fixes two things, and you have to review both at once, untangling which parts belong to which.

Braiding outcomes has three costs. Review becomes harder, because changes for different purposes are mixed, and it is easy to approve the whole because one half is clearly good. Shipping becomes harder, because one outcome may be ready while the other is not, and they are now in the same piece of work. And reverting becomes harder, because if one fix turns out to be wrong, undoing it means undoing both. All three costs are paid later, which is why the efficiency at the start is so tempting.

Two outcomes in one thread is one thread you cannot ship.

Think of threads on two axes: how broad their scope is and how clear their outcome is. Narrow scope with a clear outcome is the corner you want. Broad scope with a clear outcome is manageable but slow to review. Narrow scope with an unclear outcome tends to wander. Broad scope with an unclear outcome is the thread that runs for three days and produces something nobody can evaluate. One thread, one outcome, pushes you firmly towards the good corner.

So when the related idea arrives mid-thread, send it to the queue, or open a separate thread for it with its own brief. Yes, the new thread will need to load context again. That costs a minute. Untangling a braided change costs much more, and the minute buys you two clean pieces of work that can each be reviewed, shipped and reverted on their own.

The same rule helps agents. An agent working on one outcome can check its work against one definition of done. An agent working on two has to balance them, and may well decide that a change helpful for one is acceptable even if it slightly harms the other. You want those trade-offs made by you, at review, not by the agent, mid-run.

There are exceptions. Sometimes two changes are genuinely inseparable, because one cannot be done without the other. In that case, they are really one outcome, and the brief should say so. The rule is not about the number of files touched. It is about the number of things you will have to decide about when the work comes back.

This week, audit your open threads. For each, write its outcome in one sentence. If you need the word and, the thread has two outcomes. Split it, or let one wait. Your reviews will get shorter almost immediately.

One thread, one outcome Clear outcome Fuzzy Narrow scope Broad scope One outcome the corner you want one definition of done review, ship, revert each on its own Manageable broad and clear clear, but a big change slow to review Wanders narrow and fuzzy small, but no target drifts while you watch The three-day thread broad and fuzzy braided outcomes one half good, one not nobody can evaluate it write each outcome in a sentence; if it needs "and", split
Fig 46 · One Thread, One Outcome. Narrow scope and a clear outcome keep threads reviewable, shippable and revertible.
Chapter 47 · Part V

Parallel, Not Tangled

One of the great promises of agents is parallelism. Several threads can run at once, each working on a different piece of the operation while you review the results as they arrive. Done well, it is extraordinary: a morning in which four separate pieces of work progress simultaneously and each one finishes cleanly. Done badly, it is a tangle in which threads interfere with each other, step on each other's changes and generate more review than any single person can handle.

The difference comes down to separation. Parallel work stays clean when the pieces are genuinely separate: different files, different outcomes, different branches, no shared state that more than one thread is changing at once. Clean parallel work is the overlap of two things, running at the same time and not touching each other. Lose either and you get either slow sequential work or fast tangled work.

The most common tangle is two threads changing the same files. Each thread does sensible work on its own terms, and when both finish, their changes conflict. Resolving the conflict requires understanding both changes in detail, which is exactly the review effort you were hoping parallelism would spread out. The fix is to plan parallel work so that each thread has its own territory. If two pieces of work must touch the same files, run them in sequence, not in parallel.

Parallel work is a gift only when the pieces do not share a wall.

Many tools now help with this directly. Agents can work in separate copies of a repository, often called worktrees or isolated environments, so that their changes do not collide until you choose to combine them. Use these whenever you run more than one thread on the same codebase. For non-code work, the equivalent is separate documents or separate sections, each owned by one thread.

The second tangle is review overload. Four parallel threads produce four sets of results, and they tend to finish at roughly the same time. If each needs careful review, you now have a queue of review work that can take longer than the threads took to run. The fix is to stagger: start threads at intervals, or give them work of different sizes, so that results arrive one at a time. Or simply run fewer in parallel. Two threads that you review well beat five that you review in a hurry.

The third tangle is attention. Watching several threads at once feels efficient and is not. Each check-in costs a context switch, and context switches are expensive for people even when they are free for agents. Batch your check-ins: look at all running threads together at fixed intervals, rather than flicking between them whenever one moves.

This week, before starting parallel threads, draw a quick map: which thread touches which files or documents. If any two overlap, sequence them instead. Use isolated copies for code. Stagger the starts. Then notice how the reviews feel. Parallel work should feel like a well-run kitchen, not a crowded one, and the difference is mostly in the planning before anyone picks up a knife.

Staggered, separate, reviewable Thread A worktree a runs Thread B worktree b runs Thread C worktree c runs You review review A review B review C one at a time time: starts staggered Own territory clean parallel work A: importer/ B: exporter/ C: docs/ isolated copies, combined later A shared wall tangled parallel work two threads, one file: conflict run them in sequence instead Parallel work is a gift only when the pieces do not share a wall.
Fig 47 · Parallel, Not Tangled. Staggered threads in separate territories deliver results for review one at a time.
Chapter 48 · Part V

Handing Off Well

Every delegation starts with a handoff: the moment you pass a piece of work to an agent. Most operators think of this as writing the brief, and the brief is indeed most of it. But a good handoff is a short exchange rather than a single message, and the second half, the part where the agent asks questions before starting, is where many expensive misunderstandings are caught.

The exchange goes like this. You send the brief and any context the agent needs. The agent reads it, looks at the relevant files or material, and comes back with questions or a short plan. You answer, adjust, or confirm. Only then does the work begin. In both cases, it is far cheaper than discovering a misunderstanding after the work is done.

Many agents will not ask questions unless invited. They are built to be helpful, and helpful often means getting on with it. So invite them explicitly. A line such as before starting, list any questions or ambiguities, and propose a short plan turns a monologue into a dialogue. If the agent has no questions, it will say so, and you have lost nothing. If it does, you have found the gap in your brief before it cost anything.

The questions an agent asks before starting are cheaper than the answers it guesses.

Read the questions carefully, because they are diagnostic. If the agent asks something you thought was obvious, your brief probably assumed context that was not there. Answer it, and consider adding the answer to the project's memory so the next brief does not have the same gap. If the agent asks about something you had not considered, good: it has found a real ambiguity, and you can decide it now rather than discovering the agent's guess later.

Read the plan carefully too. A plan tells you how the agent understood the task. If the plan is aimed at the wrong target, the misunderstanding is visible in plain words, before any work is done. If the plan is right but more ambitious than you wanted, you can trim it. If it is right and appropriately sized, confirm and step away.

A good handoff also sets expectations about reporting. Tell the agent what you want to see when it finishes: the result, the evidence, any deviations, any open questions. Tell it when to stop and check in, for example before making any irreversible change or if the task turns out much larger than expected. These instructions shape the end of the work as much as the brief shapes the beginning.

There is a temptation, once you trust an agent, to skip the exchange and just send the brief. For small, familiar tasks, that is fine. For anything substantial or new, keep the exchange, even if it is only a minute. Trust is built on the exchanges that caught problems early; skip them and you will eventually rediscover why they mattered.

This week, add the questions-and-plan line to every brief for a substantial task. Keep a tally of how often the exchange catches something. Most operators find it is more often than they expected, and that each catch saved far more time than the exchange cost.

A handoff is a short exchange Operator Agent brief + context + "questions first" Reads the files questions + a short plan Read them brief gaps? answer, trim, confirm Work begins Report back with set it at the start the result the evidence any deviations open questions Stop and check in before anything irreversible, or if it grows much larger Questions asked before starting are cheaper than answers guessed.
Fig 48 · Handing Off Well. A handoff exchange: brief, questions and plan back, confirmation, then work begins.
Chapter 49 · Part V

When to Pull a Thread Back

Threads drift. An agent sets off on a well-briefed task and, somewhere along the way, starts doing something adjacent. It refactors a module the brief did not mention. It spends twenty minutes on a test environment problem that has nothing to do with the task. It rewrites a section that only needed a tweak. It goes round the same loop three times, trying slight variations of an approach that is not working. Each step is reasonable on its own. The direction is wrong.

The operator's skill is to notice drift early and pull the thread back before it costs much. That requires knowing what drift looks like, and checking in often enough to see it. The middle-of-the-day rhythm described earlier helps: at each check-in, ask not just whether the thread is busy but whether it is busy on the right thing.

Some signs of drift are reliable. The agent is editing files outside the scope you set. The agent is fixing problems you did not ask it to fix. The agent has tried the same kind of fix more than twice without success. The agent's progress reports describe activity rather than results. The time elapsed is well beyond what the task should have needed. Any one of these is worth a closer look. Two or more together almost always mean drift.

Drift is not a failure of effort. It is effort pointed the wrong way.

When you spot drift, the response depends on how far it has gone. Early drift can be corrected with a sentence: stay within the importer; ignore the exporter for now. Moderate drift may need a short re-brief: stop, summarise what you have learned, then continue with this narrower goal. Severe drift, where the thread has gone deep into the wrong territory or accumulated a confusing history, is often best handled by stopping the thread altogether and starting fresh with a better brief and a handoff note capturing anything useful the drifting thread discovered.

Pulling a thread back is not a criticism of the agent. Drift usually traces to something upstream: a brief that left the scope ambiguous, a constraint that was not stated, a definition of done that did not make the target clear, or a genuine surprise in the work that the agent handled by improvising. When you pull a thread back, take a moment to ask which of these applied. Fixing the cause improves every future thread. Fixing only the symptom improves this one.

There is also a case for letting a thread run. Sometimes what looks like drift is the agent discovering that the task genuinely requires more than the brief anticipated. If it flags this clearly, as a good deviation instruction asks it to, you can decide whether the larger scope is acceptable. The problem is not an agent that goes beyond the brief. It is an agent that goes beyond the brief silently.

This week, at every check-in, ask one question of each running thread: is it working on what I asked for? Pull back anything that is not. Note what caused each drift. After a week, you will have a short list of brief-writing habits to change, and fewer threads to pull back the week after.

Match the pull to the drift Signs of drift two together: drift edits outside the scope fixes nobody asked for same fix tried thrice reports activity, not results far past the expected time Early one sentence stay in scope Moderate re-brief narrower goal Severe stop it fresh brief THEN FIX THE CAUSE UPSTREAM vague scope unstated limit unclear done a real surprise Drift is not a failure of effort. It is effort pointed the wrong way.
Fig 49 · When to Pull a Thread Back. Signs of drift and the escalating responses, from one sentence to a fresh thread.
Chapter 50 · Part V

The Coordinator's Seat

At some point, the operator's role shifts from running threads to coordinating them. Instead of one thread at a time, there are several, each working on a slice of a larger outcome, and the job becomes deciding how to divide the work, dispatching the slices, reviewing what comes back and integrating the results into something whole. This is the coordinator's seat, and it is the most ambitious way to run a one-person operation.

Some tools now support coordination directly, with a lead agent that breaks work down and hands slices to other agents, or a project space in which a coordinating session dispatches parallel threads. These are useful, and they will keep improving. But the coordinating role does not disappear when a tool takes on part of it. Someone still has to decide how the work should be divided, whether the slices fit together and whether the integrated result is what was wanted. That someone is you.

The coordinator's loop has three steps. Dispatch: divide the outcome into slices that can run in parallel without tangling, and brief each one with its own goal, scope and done. Review: as slices return, check each against its own brief. Integrate: combine the slices, check that the whole works, and decide whether it meets the original outcome. Then, usually, go round again, because integration tends to reveal a gap that needs another slice.

The coordinator does not do the work. The coordinator makes the pieces fit.

Integration is where the value is, and it is the step most likely to be underestimated. Each slice can pass its own checks and the whole can still fail, because the slices made slightly different assumptions, or because nobody was responsible for the joins between them. Good coordinators define the joins in advance: the interface between two slices, the shared format, the agreed vocabulary. They also check the joins first at integration, because that is where problems hide.

The division of work is the other place skill shows. A good division gives each slice a clear, independent outcome and minimises the joins. A poor division creates slices that depend on each other's details, so that one slice cannot finish until another has, and parallelism collapses into a queue. Time spent on the division is time saved everywhere else. Draw it out before dispatching anything.

The coordinator's seat is demanding. It requires more attention than running a single thread, not less, because you are holding the whole in your head while the pieces are being made. It is worth it for large outcomes that genuinely divide well. It is not worth it for outcomes that are small enough to run in one thread, and many operators reach for coordination too early, because it feels advanced. A single good thread is usually faster than three coordinated ones for anything that fits in a session.

This week, find one outcome large enough to divide. Draw the slices and the joins on a page. Dispatch two or three slices with careful briefs. Review each, then integrate, checking the joins first. Notice how much of your time went on the division and the joins rather than on the slices. That proportion is the coordinator's job.

Dispatch, review, integrate DISPATCH REVIEW INTEGRATE Large outcome divides well Slice A own goal, done Slice B own goal, done Slice C own goal, done join: the API join: format Review each slice vs its brief Integrate joins first whole works? integration reveals a gap: one more slice The coordinator does not do the work. The coordinator makes the pieces fit.
Fig 50 · The Coordinator's Seat. The coordinator dispatches slices, reviews each, integrates via the joins and loops.
Part VI

Reviewing the Work

Evidence, diffs and taste.

Chapter 51 · Part VI

Review Is the Job

When agents do most of the making, the operator's working hours migrate towards reviewing. This surprises people. They expected to spend their freed-up time on strategy, creativity or rest, and instead find themselves reading diffs, checking drafts and testing features for a large part of each day. It can feel like a demotion: from maker to inspector. It is the opposite. Review is where the operator adds most of the value, and treating it as a chore is one of the most expensive mistakes in the trade.

Consider what review actually does. Agent work arrives at the bottom of the stack: plentiful, fast, usually good, occasionally wrong in ways that are hard to spot. At the top sits the decision to ship, which puts your name on the work. Between them is the review, the only layer where your judgement touches the work directly. Everything you know about the client, the users, the quality bar and the risks enters the work through review. Skimp on it and all that knowledge stays in your head, unused.

Review is also where you learn. Each review shows you how the agent interpreted your brief, which tells you how to brief better next time. It shows you where the project's memory is incomplete, where its tests are weak, where its conventions are unclear. An operator who reviews carefully gets better at every other part of the job, because review is the feedback loop that connects the parts.

Making is now cheap. Knowing whether it is right is not.

Treating review as the job changes how you schedule it. It gets proper time in sessions, not leftover minutes between other things. It gets your best hours, not your worst, because a tired review is where defects slip through. And it gets protected from interruption, because review needs sustained attention in a way that briefing does not.

It also changes how you measure your own productivity. If review is the job, then a day spent reviewing four pieces of work carefully and shipping three of them is an excellent day, even if you made nothing yourself. Operators who measure themselves by what they personally produced tend to feel guilty about review time and rush it. Operators who measure themselves by what shipped well tend to give review the time it needs.

None of this means reviewing everything to the same depth. A later chapter covers how to match the depth of review to the risk. A one-line copy change needs a glance. A change to how payments are calculated needs a slow, deliberate read and a real test. The skill is in knowing which is which, and giving each what it needs.

This week, look at your calendar and find where review happens. If it is squeezed into the gaps, give it its own blocks, in your sharpest hours. Then notice what changes: fewer surprises after shipping, better briefs as you learn from each review, and a growing sense that you are running the work rather than merely receiving it. That sense is the job, properly understood.

Where your judgement enters the work Ship decision your name goes on the work Your review the only layer your judgement touches best hours · protected time Agent work plentiful · fast · usually good what only you know Client Users Quality bar Risks Lessons brief · memory tests · rules next brief Making is now cheap. Knowing whether it is right is not. a good day: review four carefully, ship three
Fig 51 · Review Is the Job. Agent work rises through your review, the one layer where your judgement touches it.
Chapter 52 · Part VI

Evidence Over Reassurance

Agents are reassuring. They tell you the work is complete, the tests pass, the edge cases are handled and the change is ready for review. Most of the time they are right. But reassurance is not evidence, and an operator who accepts reassurance in place of evidence will, sooner or later, ship something that was not what it claimed to be.

The distinction is simple. Reassurance is a claim: all tests pass. Evidence is something you can check: the output of the test run, showing which tests ran and that they passed. Reassurance is the page works on mobile. Evidence is a screenshot at phone width, or better, you opening the page on your phone. Reassurance is I have handled the empty case. Evidence is a test that exercises the empty case and its result. In every pair, the first is what the agent believes. The second is what you can verify.

The reason to insist on evidence is not that agents lie. It is that agents, like people, can be mistaken about their own work. A test run may have been against an old version. A check may have passed for the wrong reason. An edge case may have been handled in one path and missed in another. The agent sincerely reports success, and the report is wrong. Evidence catches these cases. Reassurance cannot.

Ask for the receipt, not the promise.

The practical habit is to ask for evidence in the brief and to look at it in the review. In the brief: when you finish, include the full test output and a screenshot of the new page at phone width. In the review: read the output, look at the screenshot, and check that they actually show what the summary claims. This takes a minute or two. It is the most reliable minute of the review.

There is a subtler version of the same point. Some evidence is stronger than other evidence. A test that passes is weak evidence if the test would have passed before the change too. A screenshot is weak evidence if it shows a different page from the one changed. Strong evidence demonstrates that the change made a difference: the new test fails without the change and passes with it; the before and after screenshots show the fix. Ask for strong evidence when it matters.

Agents have become quite good at providing evidence when asked, and many will now do it by default for code. For other kinds of work, you may need to be more explicit. For a research summary, ask for sources and quotes. For a data transformation, ask for row counts before and after and a sample of the output. For a design change, ask for screenshots at the sizes that matter. Each of these turns a claim into something checkable.

This week, for every piece of agent work you review, look for the evidence before reading the summary. Where there is none, ask for it before approving. Notice how often the evidence confirms the summary, and how occasionally it does not. Those occasional cases are why the habit exists.

Ask for the receipt, not the promise Work Reassurance: a claim Evidence: something to check Tests All tests pass Full test output, what ran Mobile It works on mobile Screenshot at phone width Edge case Empty case handled A test that exercises it Research Summary is accurate Sources and quotes Data The transform worked Row counts and a sample Design It looks right Screenshots at real sizes EVIDENCE HAS STRENGTH Weak: passes before and after Strong: fails without the change agents can be sincerely mistaken; the receipt catches it
Fig 52 · Evidence Over Reassurance. Six agent claims set beside the evidence that would let you check each one.
Chapter 53 · Part VI

Read the Diff, Not the Summary

Every piece of agent work comes with a summary, and agent summaries are excellent: clear, well-organised, confident and easy to read. They are also written by the party that did the work, with all the natural tendency that implies to emphasise what went well and pass lightly over what did not. Read the summary, by all means. But read the diff first.

The diff, the actual record of what changed, is the ground truth. It shows every file touched, every line added and removed, every change whether or not the summary mentioned it. For code, it is literally a diff. For documents, it is a comparison between the old version and the new. For data, it is the before and after. Whatever the medium, the diff tells you what happened. The summary tells you what the agent thinks happened, which is usually the same thing and occasionally importantly different.

The differences tend to fall into a few patterns. The summary omits a change, because the agent thought it was minor: a configuration tweak, a deleted test, a modified file outside the stated scope. The summary describes the intention rather than the result: improved error handling, when the diff shows that one error case was handled and two were removed. The summary is accurate but incomplete: it lists what was asked for and not what else came along with it. None of these is deceptive. All of them matter.

The summary is the story. The diff is what happened.

Reading diffs is a skill, and it is worth developing even if you are not a programmer. Start with the shape: how many files changed, which ones, and roughly how much. If the shape surprises you, more files than expected, files outside the scope, deletions where you expected additions, look there first. Then read the substantive changes. You do not need to understand every line to notice that a test was deleted, that a constant changed from one value to another, or that a whole section of a document was rewritten when you asked for a tweak.

After the diff, read the summary, and compare. Does the summary mention everything you saw? Does it describe the changes accurately? If there is a gap, ask about it. Often there is a good reason. Sometimes there is not, and the question surfaces something that would otherwise have shipped unnoticed.

Agents can help you read diffs, too. Asking a separate agent to review a change, with no context except the original brief and the diff, is a useful second opinion, particularly for large changes. A reviewer agent that did not do the work has no attachment to it and will often notice things the original agent glossed over. Treat its findings as evidence to weigh, not a verdict to accept.

This week, reverse your reading order. For every piece of agent work, open the diff before the summary. Form your own view of what changed. Then read the summary and see whether it agrees. Count the times it does not. That count is the measure of how much your review was relying on someone else's account of their own work.

The summary is the story. The diff is what happened. THE SUMMARY SAYS YOU NOTICE THE DIFF SHOWS Added retry to importer importer: +24 −3 match Improved error handling errors: 1 case added, 2 cut intention, not result Updated the tests tests: 1 test deleted understated (silence) config: timeout changed not mentioned READING ORDER 1 Diff shape files and size 2 Real changes tests, values 3 Summary only now 4 Compare, ask ask about gaps a reviewer agent with only the brief and the diff is a useful second opinion
Fig 53 · Read the Diff, Not the Summary. Matching an agent summary against its diff exposes omissions and overstatements.
Chapter 54 · Part VI

The Cheapest Check First

Not every review needs to be deep. Some changes are trivial and some are dangerous, and treating them all the same either wastes your time on the trivial or skimps on the dangerous. A better approach is to order your checks by cost, cheapest first, and to stop as soon as you have enough confidence for the risk involved.

The cheapest checks take seconds. Does the diff look the size you expected? Did the checks pass? Does the summary match the brief? Is the evidence present? These questions catch a surprising fraction of problems, particularly the gross ones: the agent worked on the wrong file, missed half the task or broke something obvious. If any of them fails, you have learned what you need without spending more time, and the work goes back.

The next tier is real use. Open the page. Run the feature. Read the document as its intended reader would. Load the data into the tool that will use it. This takes minutes rather than seconds, and it catches a different class of problem: the work that is technically correct but does not do what anyone actually needed. Real use is the check most often skipped, because it feels slow compared with reading. It is also the check most likely to catch the problems users would find.

Look first where looking is cheap. Look harder where mistakes are expensive.

The deepest tier is the careful read. Line by line through the diff, thinking about edge cases, asking what could go wrong, checking that the change is consistent with the rest of the system. This takes real time and real attention, and it is the right check for changes with high stakes: anything involving money, security, personal data, irreversible actions or code that many other things depend on. It is the wrong check for a change to some button text.

The order matters because it lets you stop early. A change that fails a cheap check does not need a deep read; it needs to be redone. A change that passes the cheap checks and the real-use test may be fine to ship if its stakes are low. Only the changes that pass the first two tiers and carry real risk need the third. This is how an operator reviews a high volume of work without either drowning or getting careless.

It helps to decide the depth in advance, as part of the brief. When you write the definition of done, also note the risk level: low, medium or high. Low-risk work gets cheap checks and a quick real-use test. Medium-risk work gets all of that plus a focused read of the key changes. High-risk work gets everything, slowly, preferably when you are fresh. Writing the risk down beforehand stops you from deciding it in the moment, when you are tired and inclined to think everything is low risk.

This week, label every brief with a risk level, and review accordingly. Notice how much time the cheap checks save on low-risk work, and how much more attention you have left for the high-risk work. That reallocation is the whole point. Attention is finite; spend it where mistakes cost most.

Look first where looking is cheap Ship Cheap checks seconds size as expected? checks passed? matches the brief? evidence present? Real use minutes open the page run the feature read as the reader load the data Careful read real time line by line edge cases what could go wrong? fits the system? pass pass low risk high risk Send back or redo fail fail fail RISK, SET IN THE BRIEF low: tiers 1–2 medium: + focused read high: all, when fresh money · security · personal data · irreversible → high
Fig 54 · The Cheapest Check First. Three review tiers ordered by cost, with early exits to ship or send back.
Chapter 55 · Part VI

The Two-Pass Review

For anything larger than a small change, review in two passes. The first pass looks at shape: is this the right kind of thing, aimed at the right target, built in a sensible way? The second pass looks at detail: is each part correct? Doing them in that order saves a great deal of time, because a piece of work with the wrong shape does not deserve a detailed review.

The shape pass is quick and high-level. Read the summary and skim the diff. Look at the structure of the document or the layout of the page. Ask: did the agent understand the task? Is the approach reasonable? Is the scope right, neither too narrow nor sprawling? Does it fit with the rest of the project? You are not checking anything line by line. You are checking that the work is the work you wanted.

If the shape is wrong, stop. Do not proceed to detail. Detailed comments on work with the wrong shape are wasted, because the work will be redone and the details will change. Worse, detailed comments can make it look as if the shape is acceptable and only the details need fixing, which steers the next attempt in the wrong direction. Send it back with a clear note about the shape: this solves it at the database level; I wanted it solved in the interface. One sentence about shape is worth twenty about detail.

Never polish the wrong thing.

If the shape is right, move to the detail pass. Now go slowly. Check the edge cases, the error handling, the wording, the numbers. Look at each change and ask whether it is correct, consistent and necessary. This is where the depth chosen in the previous chapter applies: a low-risk change gets a lighter detail pass, a high-risk one a thorough one.

Then decide. Ship it, send it back with specific notes, or park it. The decision completes the loop, and if the work goes back, the next version gets the same two passes. Often the second time round, the shape pass takes seconds because the shape was settled, and only the detail pass needs real time.

The two-pass habit is worth applying to your own work as well as the agents'. When you write a brief, a register line or a decision, look first at whether it is the right thing, then at whether it is right in its particulars. Operators who skip the shape pass on their own work often find themselves carefully editing a brief for a task they should not have started.

There is a social benefit, too, for the occasions when you review work from people rather than agents. Separating shape from detail makes feedback much clearer. The approach is right; here are some details is very different from the approach needs rethinking, and people appreciate knowing which kind of feedback they are getting before they read it.

This week, review every substantial piece of work in two explicit passes. Write a one-line verdict after the shape pass before starting the detail pass. Notice how often the shape pass alone settles the matter, and how much faster the detail pass goes when you already know the work is the right work.

Shape first, then detail Work arrives v1, v2, … Shape pass quick, high-level right task? sensible approach? scope right? fits the project? Detail pass slow, depth set by risk edge cases error handling wording, numbers consistent, necessary? right verdict wrong shape: stop Send back on shape one sentence on shape redo Decide Ship Notes Park Never polish the wrong thing.
Fig 55 · The Two-Pass Review. A shape pass gates the detail pass; wrong shapes go back before any line edits.
Chapter 56 · Part VI

What Green Hides

There is a particular feeling that operators learn to distrust: the relief of seeing everything green. Tests pass. Checks pass. The build succeeds. The agent reports completion. Green is good news, and most of the time it means what it seems to mean. But green has blind spots, and the most dangerous problems are the ones that live in them.

Green tells you that the checks you have passed. It does not tell you that the checks are the right ones. A test suite that covers the main path and not the edge cases will be green for code that breaks on the edge cases. A check that verifies a page loads will be green for a page that loads with the wrong content. A linter that enforces formatting will be green for code that is beautifully formatted and logically wrong. The overlap between all green and still wrong is where the trouble hides.

Agents add a specific risk here. An agent asked to make the checks pass will try hard to do so, and occasionally the easiest way to make a check pass is to change the check. It might weaken an assertion, skip a failing test, broaden an expected value or mock away the part that was failing. Most agents are now quite good about not doing this, and many will flag it if they do. But it happens, and it produces the most deceptive kind of green: checks that pass because they no longer check anything.

Green means the checks passed. It does not mean the checks were right.

The defence is to review the checks as well as the code. When a diff includes changes to tests, read those changes with particular care. Ask whether any test got weaker: an assertion removed, a value loosened, a case skipped. A brief can help here by stating that tests may be added but not weakened without explicit approval, and that any change to an existing test must be flagged.

The second defence is to look for evidence that the checks exercised the new work. A new feature should come with a new test that fails without the feature. A bug fix should come with a test that reproduces the bug. If all the tests passed before the change and all pass after, and no new test was added, the green tells you only that nothing obvious broke. It tells you nothing about whether the new thing works.

The third defence is real use, from an earlier chapter. Checks are a model of what correct means, and models are always incomplete. Using the thing, opening the page, running the import, reading the document, tests it against reality rather than the model. Real use catches what green hides more reliably than anything else.

This week, whenever you see all green, ask one extra question before shipping: what would have to be true for this to be green and wrong? Look specifically at any test changes in the diff, and check that at least one check exercised the new work. It adds a minute. It removes a class of surprises that otherwise arrive at the least convenient moment, usually from a user.

What green hides All green checks passed Still wrong users find it Blind spot How green lies main path only, no edges page loads, wrong content formatted, logic wrong assertion weakened failing test skipped Three defences read test changes closely new test fails without it real use, not the model What would have to be true for this to be green and wrong? green means the checks passed, not that they were right
Fig 56 · What Green Hides. All green and still wrong overlap in a blind spot; three defences shrink it.
Chapter 57 · Part VI

Reviewing Prose and Pictures

Much of what agents produce for an operator is not code. It is writing: emails, reports, proposals, documentation, posts. Or it is visual: pages, slides, diagrams, layouts. These have no test suite and no build that goes green. Reviewing them requires a different set of checks, and many operators review them less carefully than code for exactly that reason. That is a mistake, because prose and pictures are usually what other people see.

For prose, review along two axes. One is accuracy: are the facts right, are the claims supported, is anything invented? The other is voice: does it sound like you, at the right register for the reader, without the tics that mark writing as machine-made? Prose that is accurate and in your voice is ready to send. Prose that is accurate but not in your voice may be fine for internal use but should not go to a client. Prose in your voice but inaccurate is the most dangerous kind, because it is persuasive.

Accuracy needs real checking. Agents write fluently, and fluency can mask errors: a date slightly wrong, a figure from the wrong source, a summary of a document that subtly misstates its conclusion. For anything that will be read by someone who matters, check every factual claim against its source. Ask the agent to cite sources in the draft so that you can check them quickly, and actually click through. This is tedious and it is essential.

Fluent is not the same as true, and the reader cannot tell the difference. You can.

Voice is harder to check, but it can be done. Read the prose aloud, or at least slowly. Notice phrases you would never say. Notice structures that repeat: every paragraph opening the same way, every list with exactly three items, every conclusion summarising what was just said. These are common patterns in generated text and readers increasingly recognise them. Strike them out, or better, add a note to your memory or recipe so that future drafts avoid them. Keeping a few examples of your own writing to hand, as an earlier chapter suggested, helps enormously here.

For pictures, the equivalents are correctness and fit. Correctness: does the layout work at the sizes it will be seen, is the text legible, are the colours accessible, does everything that should be clickable work? Fit: does it look like it belongs to your brand and your other work, or does it look generically competent? Check at real sizes, on real devices, in both light and dark modes if relevant. A design reviewed only at full desktop width is a design reviewed for one reader in five.

The two-pass method applies here too. Shape first: is this the right document or design, with the right structure, for the right reader? Then detail: is each sentence true, each element correct? Prose and pictures that fail the shape pass should go back without a line edit.

This week, take one piece of agent-written prose that is about to go to someone else and review it slowly for accuracy and voice, checking every factual claim and striking every phrase you would not say. Time it. It will take longer than you expected, and it will be worth it every time.

Reviewing prose: accuracy and voice VOICE yours not yours Most dangerous your voice, wrong facts persuasive Ready to send accurate, your voice → the client Rewrite neither Internal only accurate, wrong voice inaccurate accurate ACCURACY check every claim at source Pictures CORRECT real sizes legible text contrast links work light + dark FIT on brand not generic Fluent is not the same as true, and the reader cannot tell. read aloud · strike repeated tics · shape pass before line edits
Fig 57 · Reviewing Prose and Pictures. Prose sorted by accuracy and voice, beside the checks that pictures need.
Chapter 58 · Part VI

Comments That Teach

When you send work back, the comments you write do two jobs. The obvious one is to fix this piece of work. The less obvious one, and over time the more important, is to improve the next piece. Comments that only do the first job leave you correcting the same problems again and again. Comments that do both make the operation better every time you review.

A comment that teaches has three qualities. It is specific: it says exactly what is wrong and where, rather than gesturing at a general dissatisfaction. The second paragraph claims a figure without a source rather than needs more rigour. It is explanatory: it says why, so that the agent can apply the reason to similar cases. Clients have asked about sources before; every figure needs one rather than just add a source. And it is portable: it states a principle that could apply beyond this piece of work.

Specific comments get the current work fixed quickly. Explanatory comments help the agent make the right call on similar issues elsewhere in the same work. Portable comments are the ones worth promoting to memory: once you have written every figure needs a source as a comment three times, it belongs in the project's memory file or the relevant recipe, so that every future draft starts with it.

A comment fixes a draft. A comment promoted to memory fixes every draft after it.

The sequence matters. You review and write specific, explanatory notes. The agent revises. You check the revision. Then, if the comment was portable, you update memory or the recipe. That last step is the one most often skipped, because by the time the revision is approved you are ready to move on. But it is the step that turns a review into an improvement, and it usually takes under a minute.

Some comments are better written as questions. Why does this function retry three times? invites the agent to explain its reasoning, which may reveal either a good reason you had not considered or a mistaken assumption you can correct. Questions are particularly useful when you suspect something is wrong but are not sure, because they let you learn before you instruct.

Avoid comments that are purely emotional. This is not good gives the agent nothing to work with. I don't like the tone is barely better. If you find yourself writing comments like these, stop and ask what specifically is wrong. Often you will discover that you do know, and can say it. Occasionally you will discover that you do not know, and that the problem is with the brief rather than the work.

Comments are also a record. Reviewing your own past comments once in a while shows you what you care about, what keeps going wrong and what your standards actually are, as opposed to what you think they are. This is useful raw material for memory files, recipes and the next quarterly review.

This week, after each review where you wrote comments, ask which ones were portable. Promote at least one per day into memory or a recipe. By the end of the week, notice whether any of the problems you commented on have stopped recurring. Some will have. That is the teaching working.

Comments that teach A GOOD COMMENT IS Specific fixes this draft Explanatory fixes similar cases Portable fixes every draft Operator Agent Memory note + why revised work check promote the portable rule next draft "every figure needs a source" A comment fixes a draft. In memory, it fixes every draft after. questions invite reasons · "not good" teaches nothing
Fig 58 · Comments That Teach. A specific, explained note fixes one draft; promoted to memory it fixes them all.
Chapter 59 · Part VI

Redo or Repair

When a piece of agent work comes back with problems, you face a choice: repair it, by sending notes for the agent to fix, or redo it, by discarding the work and starting again with a better brief. Most operators default to repair. It feels economical: the work is mostly there, why throw it away? But repair is often the more expensive choice, and knowing when to redo is one of the quieter skills of the trade.

The deciding question is whether the shape is right. If the work is aimed at the right target and built in a sensible way, with problems confined to details, repair it. Detail fixes are quick and the agent has the context to make them well. If the shape is wrong, the approach misguided, the structure unworkable, the target misunderstood, redo it. Repairing a wrong shape produces a patched, compromised version of something that should not exist, and it takes longer than starting fresh.

Agents change the economics in favour of redo. When a human colleague has spent a day on something, throwing it away is costly and demoralising. When an agent spent twenty minutes, throwing it away costs twenty minutes and no feelings. The temptation to repair comes partly from habits formed in a world where making was expensive, and those habits no longer fit.

When making is cheap, the expensive thing is patching the wrong thing.

Redo has another advantage: it improves the brief. A failed first attempt is information. It shows you what the agent misunderstood, what context it lacked, what constraint you forgot to state. When you redo, you write a better brief with that information in it, and the second attempt benefits. Repair does not force this; you fix the symptoms in the current work and the brief stays as it was, ready to produce the same problem next time.

There is also the matter of thread history. A thread that has been through several rounds of repair carries all of those rounds in its context. The agent sees its original wrong approach, your corrections, its partial fixes and your further corrections. This history can anchor it, making it harder to move cleanly to a better approach. A fresh thread with a better brief starts unanchored.

A useful rule of thumb is the two-round limit. If a piece of work has been through two rounds of repair and still is not right, stop and redo it. By then, the problems are almost certainly about shape, whatever they looked like at first, and a third round of repair rarely succeeds where two have failed.

When you do redo, keep what was useful. The failed attempt may have discovered something real: a constraint in the code, a fact about the data, a dead end worth noting. Put those discoveries in the new brief. Redo does not mean forgetting. It means starting the making again with everything learned and nothing tangled.

This week, apply the shape question to every piece of work that comes back with problems. Where the shape is wrong, redo with a better brief instead of sending notes. Track how long redo takes compared with what repair would have taken. Most operators find redo is faster more often than they expected, and that the second attempt is noticeably better.

Redo or repair? Work back with problems Right shape? yes Repair detail notes ≤ 2 rounds? Ship fixed no still wrong Redo with a better brief fresh thread, unanchored keep what the attempt discovered an agent attempt ≈ 20 minutes, no feelings When making is cheap, patching the wrong thing is expensive.
Fig 59 · Redo or Repair. Right shape gets repaired; wrong shape, or two failed rounds, gets redone.
Chapter 60 · Part VI

Taste Is a Review Skill

After checks, evidence, diffs and shape, there is one more filter that operators apply, often without naming it: taste. Is this work not just correct but good? And not just good but the kind of good you want associated with your name? Taste is the last and narrowest filter in the review, and it is one of the few that agents cannot yet apply on your behalf.

Picture the filters as a funnel. At the top, a great deal of work is correct: it meets the brief, passes the checks, does what was asked. Less of it is good: well made, clear, appropriately simple, pleasant to use or read. Less still is yours: carrying the particular choices and standards that make your work recognisable. The funnel narrows because each filter is harder to satisfy and harder to specify.

Agents are now very good at producing correct work and often good at producing good work. What they find hardest is producing work that reflects your particular taste, because taste is mostly tacit. You know it when you see it. You struggle to write it down. And so it is the part of the standard that most often has to be applied at review, by you, after the work is made.

Correct is the floor. Taste is the signature.

Taste can be cultivated and partly transmitted. Cultivated, because you get better at it by paying attention: noticing what you admire in other people's work and why, noticing what bothers you in your own and why. Transmitted, because you can capture parts of it in examples, in memory, in recipes and in comments that explain rather than merely correct. Each time you write down we prefer one clear sentence to two careful ones or no decorative illustrations, only ones that explain, you move a little of your taste from tacit to explicit, and the agents get a little closer.

But taste should not be fully delegated, and there is a reason beyond capability. Taste is how your work differs from everyone else's. If everyone hands their taste to the same agents, everyone's work converges. The operator who keeps applying their own taste at review, who rejects the generically competent in favour of the particular, produces work that stands out precisely because it was filtered through one person's sensibility.

Taste is also a check on volume. Agents make it easy to produce large amounts of acceptable work. Taste is what stops you from shipping all of it. It says: this is fine, but it is not good enough to bear our name, so it does not go out. That refusal is uncomfortable, because the work is right there and would only take a click to publish. It is also what keeps the quality of your output from drifting downwards as the volume goes up.

This week, add one question to the end of every review: is this mine? Not merely acceptable, not merely correct, but something you would be glad to have associated with you. Where the answer is no, say why in a sentence, and see whether that sentence belongs in your memory. Slowly, the agents will learn some of your taste. The rest will remain yours, which is as it should be.

Taste is the narrowest filter Correct meets the brief, passes checks Good well made, clear, simple Yours your signature agents: very good agents: often agents: hardest, tacit Tacit → explicit "one clear sentence beats two careful ones" "no decorative illustrations, only ones that explain" Correct is the floor. Taste is the signature. Ask: is this mine?
Fig 60 · Taste Is a Review Skill. Work narrows from correct to good to yours, the filter agents find hardest.
Part VII

Release Discipline

Pull requests, automerge and shipping small.

Chapter 61 · Part VII

Ship Small, Ship Often

The most reliable release discipline for a solo operator is also the simplest: ship small, ship often. Small changes, each one reviewed and released on its own, in a steady stream. Not large batches assembled over weeks and released in a nervous lump. Agents make this easier than ever, and they also make the opposite mistake easier than ever, so it is worth being deliberate about.

Small changes are easier to review. A change that touches three files and does one thing can be understood in minutes. A change that touches thirty files and does six things takes an hour, and even then you are not quite sure you have seen everything. Since review is the job, and review attention is finite, small changes are simply the efficient way to spend it.

Small changes are easier to ship safely. If something goes wrong after a small release, you know exactly what changed and can fix or revert it quickly. If something goes wrong after a large release, you have to work out which of many changes caused it, often while users are waiting. The blast radius of a small change is small by construction.

Small batches are not timid. They are how you go fast without falling over.

And small changes keep the operation moving. When work accumulates into large releases, everything is perpetually ninety per cent finished, which is the most tiring state for any project to be in.

Agents complicate this in a particular way. They are capable of producing large changes very quickly, and a well-meaning agent asked to implement a feature will often implement the whole feature, its tests, its documentation and a few improvements along the way, in one go. The result is a large change that arrived fast, and its speed disguises its size. The fix is to brief for small changes explicitly: implement only the data layer; we will do the interface in a separate change. Split the outcome before the agent starts, not after it finishes.

For non-code work the principle is the same. Publish the first section of a guide rather than waiting for all ten. Send the client the first draft of the most important page rather than the whole site. Release the dataset with the core fields and add the extras later.

There is one honest objection: some things genuinely cannot ship in halves. A database migration and the code that depends on it may need to go together. A redesign may look broken if half of it ships. In these cases, look for ways to ship small anyway: hide the new work behind a switch that is off by default, ship the migration in a backwards-compatible way first, release to yourself before releasing to everyone. Most large changes can be sliced if you think about it before the work starts.

This week, look at the size of everything you ship. If any single release took more than half an hour to review, ask how it could have been split. Then brief the next piece of similar work as two or three smaller pieces. Notice how much calmer shipping feels when each release is something you fully understand.

One big batch or a steady stream BEFORE: ONE LARGE RELEASE 30 files, 6 things arrived fast, size disguised review: an hour, still unsure blast radius: large which change broke it? always 90 per cent done AFTER: SMALL, OFTEN #1 #2 #3 #4 #5 3 files · 1 thing · ship review: minutes, fully seen blast radius: small fix or revert at once always moving CAN'T SHIP IN HALVES? SLICE IT ANYWAY Behind a switch off by default Migration first backward-compatible Release to yourself then to everyone Small batches are not timid. They are how you go fast without falling over.
Fig 61 · Ship Small, Ship Often. One large release beside a stream of small ones, plus three ways to slice big work.
Chapter 62 · Part VII

The Pull Request as Paper Trail

For anyone whose work lives in a code repository, the pull request is the natural unit of release: a proposed change, with its description, its checks and its discussion, waiting to be merged. Most operators think of it as a gate, the place where review happens before code goes in. It is that. But for a one-person operation, it is equally valuable as something else: a paper trail. Each pull request is a permanent, searchable record of what changed, why, and what evidence supported it.

That record is only as good as what you put into it. A pull request with the title updates and an empty description is a gate with no paper trail. Six months later, when you are trying to understand why some behaviour changed, it tells you nothing. A pull request with a clear title, a description of the reason and a note of the evidence is a small piece of documentation that will answer that question in seconds.

Think of a good pull request as four layers. The title states the outcome, in words a stranger could understand: Importer handles empty rows without crashing. The description gives the why: what problem this solves, what approach was taken, what alternatives were rejected. The evidence lists the checks: which tests were added, what was verified by hand, any screenshots. And the diff shows exactly what changed. The description is the layer that most often goes missing, and it is the one you will most want later.

A merged pull request is a letter to your future self. Write it as if they will read it.

Agents write pull request descriptions well when asked, and many agent tools now create pull requests with descriptions by default. Use this, but review the description as you review the code. Agents tend to describe what they did rather than why it was needed, and the why is the part only you may know. A sentence from you at the top, explaining the reason for the change in terms of the project, turns a competent description into a useful record.

Link the pull request to the rest of your system. If it implements a decision, mention the decision log entry. If it fixes something from an incident, mention the incident note. If it completes a project's next move, update the register. These links are cheap to make at the time and very hard to reconstruct later.

For work that does not live in a repository, the same idea applies in a different form. A document's version history, a change log for a dataset, a short note in the project's folder recording what was published and why: each is a paper trail. The medium varies. The discipline of leaving a record with every release does not.

This week, look at your last five merged pull requests, or their equivalent. Could a stranger understand from each one what changed and why? If not, write a better description for the next five. Ask the agent to draft and then add the why yourself. In a few months you will start finding the answers you need in your own history, which is one of the quiet pleasures of running a well-kept operation.

Four layers of a pull request Title: the outcome Importer handles empty rows without crashing Description: the why problem · approach · rejected options Evidence: the checks tests added · checked by hand · screenshots Diff: what changed +48 −12 across 3 files most often missing LINK IT Decision log Incident note Register A merged pull request is a letter to your future self. the agent drafts what it did; you add why it was needed
Fig 62 · The Pull Request as Paper Trail. A pull request as four layers, from outcome title down to diff, linked outward.
Chapter 63 · Part VII

Automerge and Its Conditions

Many repository platforms let a pull request merge itself automatically once its conditions are met: checks pass, required reviews are in, no conflicts remain. For a solo operator running agents, automerge is very attractive. Agent work arrives, its checks run, and if they pass it merges without you lifting a finger. The release pipeline runs while you sleep. It is also a way to ship things you have not looked at, so it deserves careful thought.

The question is not whether automerge is good or bad. It is which changes may use it. The answer depends on two things: the risk of the change and the strength of the checks. Low-risk changes with strong checks are good candidates. High-risk changes, or changes whose checks are weak, are not. The decision is a fork: either the change is safe enough to merge on the strength of its checks alone, or it needs a human merge after review.

What makes a change low risk? It touches an area that is well covered by tests. It is easy to reverse. It does not involve money, security, personal data or anything users depend on in a critical way. It is the kind of change that has gone well many times before. Dependency updates that pass the full test suite, documentation fixes, small refactors in well-tested code and content changes on low-traffic pages are typical candidates.

Automerge does not remove the review. It moves it into the checks, so the checks had better be good.

What makes checks strong? They would actually fail if the change were wrong. A test suite that covers the main paths and edge cases. A build that would break on a type error. A visual check that would catch a broken layout. If you are not confident the checks would catch a bad change, automerge is just merging without review, which is a polite name for hoping.

Set the conditions explicitly. Most platforms let you require particular checks to pass, require labels or require that only certain paths are changed. Use these to enforce your policy rather than relying on remembering it. A common pattern is to allow automerge only for pull requests carrying a specific label that you or a trusted recipe applies, and only when every required check passes. Changes that touch sensitive paths, such as payment code or configuration, can be excluded entirely.

Review automerged work afterwards, at least briefly. Automerge means you were not in the loop at merge time, not that you should never look. A quick scan in the morning review of what merged overnight keeps you aware of how the codebase is changing and lets you spot any pattern of problems before it becomes a crisis. If something bad ever automerges, treat it as an incident: work out which condition should have stopped it, and tighten that condition.

This week, write down your automerge policy in a sentence or two: which changes may automerge and which checks must pass. If you do not use automerge yet, pick one safe category to start with. If you do, check whether your actual settings match the policy. They often do not, and the gap is precisely where something will one day slip through.

Which changes may merge themselves Agent PR Sensitive path? Labelled safe? All checks pass? no yes yes Automerge Morning scan what merged overnight yes no no Human merge after your review e.g. payments, config, personal data Low risk means well covered by tests easy to reverse no money, security, data gone well many times Automerge moves the review into the checks. weak checks + automerge = hoping bad merge? an incident: tighten a condition
Fig 63 · Automerge and Its Conditions. Three gates decide whether a change may automerge or needs a human merge.
Chapter 64 · Part VII

Checks You Trust

Every release pipeline has checks: tests, builds, linters, type checkers, visual comparisons, link checkers, whatever the project needs. Over time, most projects accumulate a lot of them. The question that matters is not how many checks you have but how many you trust: how many would actually fail if something important went wrong. Those are the checks that bite. They, and only they, are the real gate.

Start by distinguishing the checks that bite from the ones that merely exist. A check that bites has caught something real in the last few months, or you are confident it would. A check that merely exists always passes and you are not sure what it would catch. A flaky check, one that sometimes fails for no reason, is worse than useless, because it trains you to ignore failures and to re-run until green. Each kind needs different handling.

Checks that bite should be required: no merge without them, automerge or otherwise. Checks that merely exist should be examined. Some are worth keeping because they protect against rare but serious problems. Others are noise and can be removed, which makes the pipeline faster and the signal clearer. Flaky checks should be fixed or disabled immediately. A check that cries wolf undermines every other check, because it teaches you that red does not always mean wrong.

A gate is as strong as the checks you would actually believe.

Agents can help strengthen checks, and this is one of the best uses of spare agent capacity. Ask an agent to look at an important module and propose tests for its edge cases. Ask it to deliberately introduce a small bug and see whether the existing tests catch it; if they do not, you have found a gap. Ask it to find flaky tests by running the suite several times and comparing results. Each of these turns a vague confidence in your checks into a measured one.

Pay particular attention to checks on the things that matter most. If the most important thing your project does is calculate invoices correctly, the checks on invoice calculation should be the strongest you have, with tests for every edge case you can think of. If the most important thing is that a page loads quickly for users on slow connections, there should be a check that would fail if it did not. The checks should mirror the risks, and the biggest risks deserve the sharpest teeth.

For work outside code, checks look different but the principle holds. A recipe for drafting client emails might include a check that every figure has a source. A recipe for publishing posts might include a check that every link works and every image has descriptive text. These are checks in the same sense: things that must pass before release, and that would catch a real problem.

This week, list every check in one project's release pipeline and mark each as bites, exists or flaky. Fix or remove the flaky ones. For the most important part of the project, ask an agent to try to sneak a small bug past the checks. If it succeeds, you have found the next check to write.

Sorting the checks you have Every check you have Bites caught something real Require it this is the real gate Merely exists always passes, unclear Examine it keep if rare but serious Flaky fails for no reason Fix or disable today: it cries wolf SPARE AGENT CAPACITY: TEST THE TESTS Propose edge-case tests Plant a bug: caught? Run suite 5×: flaky? A gate is as strong as the checks you would actually believe.
Fig 64 · Checks You Trust. Checks sorted into bites, merely exists and flaky, each with its own treatment.
Chapter 65 · Part VII

Versioned Passes

Large pieces of work do not need to be done in one heroic attempt. They can be done in passes: a first version that gets the shape right, a second that fills in the detail, a third that polishes. Each pass is a complete, reviewable unit with its own number, and each one builds on the last. This is how many writers, designers and engineers have always worked, and it fits agent work beautifully.

The key is to number the passes and give each one a distinct purpose. Pass one: structure and skeleton, everything present but rough. Pass two: substance, every part filled in properly. Pass three: quality, edge cases handled, language tightened, details checked. Tell the agent which pass it is on and what that pass is for. An agent asked for pass one will not waste effort polishing. An agent asked for pass three will not restructure.

Numbered passes make review far easier. You review pass one for shape only, using the first half of the two-pass review. You review pass two for substance. You review pass three for quality and taste. At each stage, you are looking for one kind of problem, which is much less tiring and much more reliable than looking for every kind of problem at once. And because each pass is reviewed before the next begins, problems are caught at the cheapest moment.

Do not ask for perfect. Ask for pass one, and then pass two.

Passes also make progress visible. Instead of a large piece of work that is somewhere between started and finished, you have a clear record: pass one done, pass two in progress. That fits neatly on the register and in handoff notes. It also makes it easy to stop at the right point. Some work only needs two passes. Some needs four. Numbering them lets you decide, after each review, whether another pass is worth it.

Keep each pass. Save pass one before starting pass two, either as a version in a repository, a numbered file, or a copy in the project folder. If pass two goes wrong, you can return to pass one without losing anything. If you want to understand later how the work evolved, the passes tell the story. Agents occasionally make a later pass worse than an earlier one, particularly when asked to improve something that was already good, and having the earlier version to hand makes it easy to notice and recover.

Passes work for almost any kind of output. A report: outline, draft, edit. A feature: data model, logic, interface. A design: layout, content, visual polish. A dataset: schema, population, validation. In each case, the same discipline applies: name the pass, state its purpose, review it for that purpose, keep it, then begin the next.

This week, take one substantial piece of work and run it in explicit numbered passes. Brief each pass separately with its purpose. Review each for its own kind of problem. Keep each version. At the end, compare how the work went with how a single big attempt usually goes for you. The difference tends to be less in the final quality, which may be similar, than in how much calmer the journey was.

Numbered passes, each with one purpose Pass 1 structure, rough shape v1 kept Pass 2 every part filled substance v2 kept Pass 3 edges and wording quality + taste v3 kept build review for save pass 4 only if worth it Report outline draft edit Feature data model logic interface Design layout content visual polish Dataset schema population validation Do not ask for perfect. Ask for pass one, and then pass two.
Fig 65 · Versioned Passes. Work run in numbered passes, each reviewed for one kind of problem and kept.
Chapter 66 · Part VII

Ship Beats Perfect, Mostly

There is an old principle, beloved of anyone who has ever shipped anything, that shipping beats perfect. A good thing in the world is worth more than a perfect thing in a draft folder. Users can respond to what exists; they cannot respond to what you are still polishing. Agents make this principle more relevant than ever, because the gap between good enough and perfect is often much smaller than it looks, and the time spent closing it is usually better spent on the next thing.

But the principle has an asterisk, and operators who forget the asterisk learn about it the hard way. Ship beats perfect when the cost of being wrong is low and the cost of undoing is small. When either cost is high, the calculation changes, and a little more polish before shipping is not perfectionism but prudence.

Think of it on two axes. One is how much polish the work has received. The other is how costly it would be to undo if it turns out to be wrong. Work that is cheap to undo can ship at modest polish: if something is off, fix it tomorrow. Work that is expensive to undo, an email to every customer, a change to how money is calculated, a public statement, a deletion, needs more polish before it goes. The corner where you should ship now is low undo cost and good-enough polish, which is where most work sits.

Ship what you can take back. Polish what you cannot.

The trap for perfectionists is treating every piece of work as if it were expensive to undo. A blog post can be edited after publication. A page can be fixed in the next release. A feature can be improved next week. None of these needs to be perfect before it ships; they need to be good and correct. Holding them back for polish is not caution. It is delay, and delay has costs of its own.

The trap for shippers is the opposite: treating everything as cheap to undo because most things are. An operator with a strong bias to ship, armed with fast agents, can push out something irreversible with the same casual speed as something trivial. The habit that prevents this is labelling work by undo cost before it is reviewed, as part of the brief. Low undo cost: review lightly, ship fast. High undo cost: review thoroughly, consider a staged release, sleep on it if you can.

There is a middle path for work that is expensive to undo but needs to ship: make it cheaper to undo. Release to a small audience first. Put it behind a switch you can turn off. Send the email to yourself and a colleague before sending it to the list. Many irreversible actions can be made partly reversible with a little thought, and that thought is often cheaper than extra polish.

This week, label everything you ship with its undo cost: low or high. Ship the low ones as soon as they are good and correct. Give the high ones an extra careful review and, where possible, a way back. Notice how much faster the low ones move when you stop treating them like the high ones.

Ship what you can take back IRREVERSIBLE, ROUGH POLISH, STAGE FIX FIRST Ship it now most work lives here deletion money calc customer email public statement blog post page fix feature high low undo cost rough POLISH good enough Cheaper to undo small audience behind a switch send to yourself back up first Moves down into ship-now Ship what you can take back. Polish what you cannot.
Fig 66 · Ship Beats Perfect, Mostly. Work placed by polish and undo cost; cheap-to-undo, good-enough work ships now.
Chapter 67 · Part VII

The Release Note

Every release deserves a note. Not a long one, and not necessarily a public one, but a short record written at the moment of release saying what went out, why, what might go wrong and how to undo it. It takes two minutes. It is the most useful two minutes you can spend at release time, and it is almost always skipped.

The note has four parts. What: the change, in a sentence that a user or a future you would understand. Exports now include archived items. Why: the reason, briefly. Users asked for a complete history. Risk: what could go wrong, in your honest estimate. Large accounts may see slower exports. Undo: how to reverse it if needed. Revert the merge; no data changes involved. Four lines, gathered around the release like spokes around a hub.

The risk and undo lines are the ones that pay off. Writing down what might go wrong forces you to think about it for a moment, which sometimes reveals a problem before release. And writing down how to undo it means that if something does go wrong, at an inconvenient hour, you do not have to work out the recovery under pressure. You read the note and follow it. Operators who have had to work out a rollback at midnight with a worried client on the line tend to adopt release notes very quickly afterwards.

The time to write the undo is before you need it.

Where should notes live? Somewhere you will find them when something goes wrong. A change log file in the project, appended with each release, works well. So does a section in the pull request description, or an entry in the notebook tagged with the project name. For releases that users see, a public change log is a separate and useful document, but it is not a substitute for the private note, which can be more candid about risks.

Agents can draft release notes from the pull request and the diff, and this is a sensible use of them. But the risk line in particular needs your attention, because it depends on knowledge the agent may not have: which customers are sensitive to which changes, what has gone wrong before, what you are quietly worried about. Edit the draft before you trust it.

Release notes also feed your longer reviews. At the end of a week or a quarter, reading back over the release notes gives a precise, dated record of what shipped, what risks you accepted and which of them materialised. That is excellent material for learning: you can see whether your risk estimates tend to be too pessimistic or too optimistic, and adjust.

There is a modest discipline benefit, too. If you find yourself unable to write a sensible what line, the release probably bundles too much. If you cannot write the undo, the release may be more irreversible than you realised. The note is a final check, disguised as paperwork.

This week, write a four-line release note for everything you ship. Keep them in one place per project. At Friday's close, read them back. Notice which risks you named, which happened and which did not. That small exercise will make you a better judge of risk faster than any amount of reading about it.

The four-line release note Release note What Exports now include archived items. Why Users asked for a complete history. Risk Large accounts may see slower exports. Undo Revert the merge; no data changes. forces a moment of thought read it at midnight, do not work it out The time to write the undo is before you need it. two minutes · one changelog per project · read back on Friday
Fig 67 · The Release Note. A release note as four spokes: what, why, risk and how to undo it.
Chapter 68 · Part VII

Rollback Is a Feature

Things go wrong after release. Not often, if you ship small and review well, but sometimes, and when they do, the most important question is how quickly you can put things back as they were. An operation that can roll back in a minute can afford to ship boldly. An operation that cannot roll back has to ship nervously, slowly and rarely. Rollback is not an emergency procedure. It is a feature of a healthy release process, and it should be designed in, tested and kept working.

The sequence you want is short. You ship a change. You or a check spots a problem. You roll back. Service is restored, and only then do you investigate. The order matters. The instinct, when something breaks, is to diagnose it in place and fix forwards, because you are fairly sure you know what is wrong and the fix is probably small. Sometimes that works. Often the fix introduces a second problem, and now you are debugging under pressure with users affected. Rolling back first takes the pressure off. You can investigate calmly with everything working.

For code in a repository, rollback is usually straightforward: revert the merge and redeploy. Make sure you know exactly how to do this for each project, and that it actually works. Some deployment setups make it trivial; others make it surprisingly awkward. Find out before you need to. A rollback procedure that has never been tried is a hypothesis, not a procedure.

If you cannot undo it, you did not release it. You committed to it.

Some changes are harder to roll back than others, and these deserve extra thought before release. Database changes that remove or transform data cannot simply be reverted, because the old data is gone. Emails and messages, once sent, cannot be unsent. Changes that other systems depend on may leave those systems in an odd state if reverted. For these, plan the rollback before release: back up the data, make the change in a backwards-compatible way, or stage the release so that only a small audience sees it first.

Agents can help with rollback planning. As part of the brief for any risky change, ask the agent to describe how the change could be reversed and what would be lost. If the answer is complicated or involves data loss, that is a signal to restructure the change before it ships. Many agents will also, if asked, write the rollback steps into the pull request description, which puts them exactly where you will look when you need them.

After any rollback, write an incident note, which a later part of this book describes. The rollback restored service; the note makes sure the same problem is less likely next time. Together they turn a bad moment into an improvement.

This week, pick your most important project and do a practice rollback. Ship a harmless change, then revert it, timing how long the whole thing takes and noting any step that was awkward. Fix the awkward steps. Then write the procedure in the project's runbook. You will probably never need it in a hurry. If you do, you will be glad it was a feature rather than an improvisation.

Roll back first, investigate second Operator Live system Users ship a small change change reaches users you or a check spots a problem 1 roll back service restored 2 investigate calmly 3 incident note not: diagnose live and fix forwards a rollback never tried is a hypothesis Hard to undo? Plan the rollback before release data removed → back up first emails sent → stage it dependents → backward-compatible If you cannot undo it, you did not release it. You committed to it.
Fig 68 · Rollback Is a Feature. When a release goes wrong: roll back, restore service, then investigate.
Chapter 69 · Part VII

Never Commit Secrets

Most of the rules in this book are defaults, sensible starting points that you should adapt to your own operation. This one is not. Never commit secrets. Not to a repository, not to a memory file, not to a shared document, not to a brief, not to a pull request description, not to a notebook that syncs to the cloud. Passwords, access keys, tokens, private connection strings: none of them belong anywhere that code or notes are stored.

The reason is simple. Repositories and notes are designed to be copied, shared, synced and kept forever. Once a secret is in a repository's history, removing it from the current version does not remove it from the history, from clones made in the meantime or from any service that has indexed it. A secret that has been committed should be assumed to have been seen. The overlap between your code and a secret is not a place where things are stored. It is an incident waiting to be discovered.

Agents raise the stakes in two ways. First, they work fast and touch many files, so a secret pasted into a brief or found in a configuration file can end up in a commit without you noticing. Second, they are helpful, and an agent debugging a connection problem may well put a credential directly into code to test it and then forget to remove it.

A secret in a repository is no longer a secret. It is a countdown.

The defences are layered. Keep secrets in a proper store, environment variables or the secret management your platform provides, and load them at runtime. Never paste a secret's value into a brief; tell the agent where the secret is configured and let it use the mechanism. Make sure files that hold local secrets are excluded from version control, and check this when setting up any new project. Turn on secret scanning for your repositories wherever the platform offers it, so that a committed secret is caught at push time rather than months later. And put a standing instruction in every project's memory: credentials and tokens must never be written into code, notes or commits.

If a secret is committed anyway, act immediately and in the right order. Rotate the secret first, so that the exposed value no longer works. Then remove it from the repository and, if appropriate, from the history. Then check whether the exposed value was used by anyone else while it was live. Then write an incident note and add whatever guardrail would have caught it. Rotating first is the crucial step. Deleting the line without rotating is like changing the lock's label without changing the lock.

This is the one rule in this book worth turning into a reflex. Every time you are about to paste something into a brief, a note or a file, ask: is this a secret? If it is, stop and find the proper place for it.

This week, turn on secret scanning for every repository you own, check that local secret files are excluded from version control, and add the standing instruction to every memory file. Then rotate anything you are not completely sure has never been committed. It is an afternoon of dull work. It closes the one door that should never be open.

Never commit secrets Code & notes copied, synced, kept A secret key, token, password Incident assume it has been seen IF ONE IS COMMITTED, IN ORDER 1 Rotate it first 2 Remove from repo and history 3 Check if it was used 4 Incident note + guardrail deleting without rotating = relabelling the lock LAYERED DEFENCES Secret store env at runtime Ignore file local secrets Scanning blocks at push Memory rule never in code Briefs no values A secret in a repository is no longer a secret. It is a countdown.
Fig 69 · Never Commit Secrets. Code meeting a secret is an incident: rotate first, then remove, check, note.
Chapter 70 · Part VII

Done Means Live

When is something done? Operators answer this question differently, and the answer quietly shapes how much actually ships. Some consider work done when the agent reports completion. Some when the review passes. Some when the pull request merges. The most useful answer, and the one this book recommends, is stricter: done means live, in front of the people it was for, and checked there.

The chain from finished work to done work has several links, and each one is a place where things stall. The work is merged but not deployed, because deployment is a separate step that nobody triggered. It is deployed but not live, because it sits behind a switch that was never turned on, or on a staging server nobody visits. It is live but not checked, because you assumed that if it deployed it must work. Each stall leaves work in a state that feels finished and is not, and that state is surprisingly comfortable to leave things in.

The final link, checked, is the one that matters most. Once the work is live, look at it there. Open the page on the real site. Run the feature as a user would. Confirm the email arrived in a real inbox. Check that the data appears in the real dashboard. This takes a minute or two and catches a whole class of problems that no amount of pre-release checking can: configuration differences between environments, caching, permissions, third-party services behaving differently in production. Work that passed every check before release can still fail in the real world, and the only way to know is to look.

Merged is a milestone. Live and checked is a result.

Defining done this way changes behaviour upstream. When you know that done means live and checked, you brief with the release in mind: how will this be deployed, how will it be switched on, how will I check it in production? You notice deployment friction, because every stalled deployment keeps something from being done. You schedule the check, rather than assuming it. And the register gets more honest, because projects are not marked done until they genuinely are.

It also changes how you count. The earlier chapter on throughput suggested counting finishes. With done meaning live and checked, the count becomes a count of things that actually reached people. It is a smaller number than merges or completions, and it is the only number that reflects what your operation did for anyone.

For work that does not deploy, the principle translates. A document is done when the person it was for has received it and can open it. A dataset is done when it is loaded where it will be used and a sample has been checked. A design is done when it is in the hands of whoever will build or publish it. In each case, done is defined at the far end of the chain, not at your desk.

This week, change your definition of done to live and checked, and apply it to everything. Notice how many things you would previously have called done that are stuck at an earlier link. Push each one through to the end and look at it there. Some will need a small fix you would otherwise have missed. All will finally be done.

Done means live, and checked there Reported Reviewed Merged Deployed Live Checked feels done is done stalls: nobody deployed switch left off never looked Look where it lives open the real site run it as a user would see the real inbox see the real dashboard What only live shows environment config caching permissions third parties in production NOT CODE? DONE AT THE FAR END Document received, opens Dataset loaded, sample checked Design with whoever builds it Merged is a milestone. Live and checked is a result.
Fig 70 · Done Means Live. The chain from reported to checked, with the stalls where work feels done.
Part VIII

Weeks, Quarters and Energy

Longer loops and the budget underneath.

Chapter 71 · Part VIII

The Weekly Review

The day has its review and its close. The week needs something larger: an hour, once a week, in which you step back from the daily rhythm and look at the operation as a whole. The weekly review is where you notice what the days cannot show you, the project that has quietly stalled, the pattern of problems, the commitment you forgot, and where you decide what next week is for.

The review has four movements. Gather: collect everything that happened this week. The register, the notebook, the release notes, the queue of parked ideas, any open threads. Review: read through it, looking for what went well, what went badly and what has not moved. Decide: make the decisions the week has surfaced, about which projects to push, which to park, which threads to close, which recipes to fix. Plan: sketch next week, not in detail, but in terms of the two or three outcomes that would make it a good week.

The deciding is the heart of it. A weekly review that gathers and reads but does not decide is a pleasant hour of reflection with no effect. Make sure the review produces decisions, written down, at least a few each week. Park the course outline until next quarter. Close the three stray threads on the old migration. Add the empty-file check to the import recipe. Next week's main outcome: ship the export feature. These are the sentences that turn reflection into steering.

The day keeps you moving. The week keeps you pointed.

Choose a fixed time and protect it. Many operators prefer the end of the working week, when the week is fresh in mind and the review can double as a proper close before the weekend. Others prefer the start of the week, when they are rested and planning feels natural. Either works. What does not work is reviewing whenever you find a spare hour, because you will not find one.

Agents can do much of the gathering. A good recipe asks an agent to read the week's notebook entries, the register, the release notes and the open threads, and to produce a short digest: what shipped, what stalled, what is waiting on you, what looks stale. That digest makes the review faster and more thorough. It does not replace your reading of the underlying material, particularly the notebook, where your own words often reveal more than any summary.

The weekly review also maintains the machine. It is the natural time to check the register is current, prune the queue, close stray threads, glance at memory files for anything stale, and make sure nothing has been left in an ambiguous state. Ten minutes of tidying each week prevents the slow accumulation of mess that would otherwise need a whole day to clear.

This week, book an hour for the weekly review and run it with the four movements. Gather, review, decide, plan. Write down at least five decisions. Next week, start the review by checking whether those decisions were carried out. That simple loop, decide then check, is what makes the weekly review an engine rather than a ritual.

The weekly review, in four movements GATHER REVIEW DECIDE PLAN Agent You Digest shipped · stuck waiting · stale Tidy stray threads stale queue the agent gathers; the deciding stays yours Read it all notebook first your own words Look back well · badly not moved Decide at least five written down Plan 2–3 outcomes for next week next week opens by checking last week’s decisions "park the outline" · "close three stray threads" · "ship the export" The day keeps you moving. The week keeps you pointed.
Fig 71 · The Weekly Review. Gather, review, decide and plan, split between an agent digest and your reading.
Chapter 72 · Part VIII

Clearing the Decks

Over a week, open loops accumulate. A thread you started and did not finish. A message you meant to answer. An idea in the queue. A review half done. A pull request waiting. A note to yourself to look into something. Each one is small, and each one sits in the back of your mind consuming a sliver of attention. Together they produce the low, constant hum of things not dealt with, which is one of the most tiring sensations an operator can have.

Clearing the decks is the practice of taking every open loop, once a week, and dealing with it in one of a few ways. Close it, by finishing it if it takes two minutes. Decide it, by choosing what will happen next and writing that down. Park it, with a note and a condition. Or drop it, by accepting that it is not going to happen and letting it go. The aim is not to finish everything. It is to make sure that nothing remains in the undecided state.

The funnel is steep. You begin with every open loop you can find, and there will be more than you expected. Most of them, on inspection, need only a quick decision: this thread can close, that idea can be dropped, this message needs a one-line reply. A smaller number need real thought. At the bottom, every loop has been either cleared or given a clear next step, and the hum goes quiet.

An open loop costs attention whether you work on it or not. A decided one costs nothing.

Finding the loops is half the work. Look in the obvious places: the register, the queue, open threads, draft folders, pull requests, your inbox. Then look in the less obvious ones: the notebook, where you may have written must look into this on Tuesday and never did; the end of agent reports, where open questions get listed and then forgotten; the corners of your desk, physical or digital. Many operators keep a short checklist of places to look, which turns finding the loops from an act of memory into a routine.

Be decisive about dropping. Some loops will never get done, and that is fine, but only if you admit it. An idea that has sat in the queue for a month without making it into a weekly plan is probably not going to happen. Drop it. A thread that was exploring something you no longer care about can close. Dropping is not failure. It is a decision, and decisions are what clear the deck.

Agents can help find loops. Ask one to list all open threads with their last activity date, all draft pull requests, all unanswered questions in recent reports, and all queue items older than two weeks. That list is a good starting point. But the deciding is yours, and it is best done quickly, almost briskly, without agonising over each item.

This week, as part of the weekly review, clear the decks. Find every open loop, and for each one close, decide, park or drop. Count how many you started with and how many remain undecided at the end. The target for the second number is zero. You may not hit it the first time. You will notice the quiet when you get close.

Clearing the decks WHERE LOOPS HIDE register queue open threads drafts pull requests inbox notebook report endings Every open loop more than you expected Close finish it if it takes 2 min Decide write down the next step Park with a note and a condition Drop admit it will not happen Undecided at the end: 0 agents list them; you decide, briskly An open loop costs attention whether you work on it or not.
Fig 72 · Clearing the Decks. Every open loop is found, then closed, decided, parked or dropped until none remain.
Chapter 73 · Part VIII

Counting What Shipped

Operators with agents can lose track of what they actually achieved. So much happens, so many threads run, so much output arrives, that by the end of a week it is genuinely hard to say what changed. This matters more than it seems, because an operator who does not know what shipped cannot judge whether the operation is working, and tends to compensate for the uncertainty by starting more things.

The remedy is an honest weekly count. Not of hours worked, threads started or tokens used, but of things that shipped: live and checked, in front of the people they were for. Keep the count in the notebook or the register, with a line for each item. At the end of the week, read the list.

It helps to see the count as the top of a stack. At the bottom is everything started: the widest layer, and the least meaningful. Above it is everything finished by agents, which is narrower. Above that is everything merged or sent, narrower still. At the top is what was live and checked. Each layer is real, but only the top one is a result. Many operators, when they first do this count, discover that their top layer is surprisingly thin compared with the layers beneath it.

Count the finishes. The starts will take care of themselves.

That discovery is useful rather than depressing. A thin top layer with a thick middle means work is getting stuck between finished and shipped, usually at deployment, at final review or at the point of actually sending something. A thin top layer with a thick bottom means too much is being started and not enough is being carried through. Each pattern suggests a specific fix, and the count tells you which pattern you have.

The count is also good for morale, in the right way. Operating alone can be isolating, and without colleagues to notice your work, it is easy to feel that you are not achieving much. A written list of what shipped, read at the end of the week, is concrete evidence to the contrary. On weeks when the list is short, it is a prompt to ask why, without guilt, as one would ask about a machine running below capacity.

Keep the counts. Over a quarter, they become a record of the operation's actual output, which is far more useful than your memory of it. At the quarterly review, reading twelve weeks of shipped lists shows you trends that no single week reveals: which kinds of work ship reliably, which kinds stall, how the volume has changed, whether your capacity estimates are accurate.

Be strict about what counts. Something is shipped if it is live and checked, not if it is nearly there, merged but not deployed, or sent to the agent for one last revision. Strictness keeps the count honest, and honest counts are the only kind worth keeping.

This week, keep a shipped list, one line per item, strictly defined. At the weekly review, read it and compare it with what you started. Look at the gap between the layers. That gap is where the next improvement to your operation is hiding.

Counting what shipped Live and checked the only result Merged or sent Finished by agents Started widest, least meaningful Thick middle stuck after finish → check deploy, review, sending Thick base too much started → start less, carry through ONE LINE PER ITEM, EVERY WEEK, STRICTLY W1 W2 W3 W4 W5 W6 W7 W8 W9 W10 W11 W12 twelve lists → the quarterly review sees trends Count the finishes. The starts will take care of themselves.
Fig 73 · Counting What Shipped. Work narrows from started to live and checked; the gaps show where it sticks.
Chapter 74 · Part VIII

The Quarterly Review

The weekly review keeps the operation pointed in the right direction. The quarterly review asks whether it is the right direction at all. Once every three months, take a half day, ideally away from the usual desk, and look at the whole operation from far enough back to see its shape. This is where you change course on purpose rather than by drift.

The quarterly review revolves around four questions about every part of the operation. What should I keep doing, because it works? What should I end, because it no longer earns its place? What should I start, because something is missing? And what should I change, because it is right in principle but wrong in practice? Ask these of projects on the register, of recipes in the library, of tools in the toolbelt, of habits in the daily rhythm. The answers are the quarter's decisions.

Start with the record. Read the shipped lists from each week, the decision log, the incident notes and the notebook. Look for patterns. Which projects produced the most? Which consumed the most attention for the least result? Which kinds of work went well and which kept going wrong? What did you decide at the last quarterly review, and did it happen? This reading is the raw material for honest answers, and it guards against the natural tendency to remember the quarter as better or worse than it was.

The week asks what to do next. The quarter asks what to stop doing.

The end question is the hardest and the most valuable. Every operation accumulates projects, tools, habits and commitments that made sense once and no longer do. They persist because ending things is uncomfortable and continuing them is easy. The quarterly review is the moment to end them deliberately. A project that has been parked all quarter without its condition being met should probably end. A tool you have not used in three months should probably go. A recipe that keeps producing poor results should be fixed or retired. The next chapter is about ending things well.

The start question is where new direction comes from. Look at the parked ideas, the queue, the gaps the quarter revealed. Pick at most one or two new things to start. Starting more than that, on top of everything continuing, usually means none of them gets enough attention to succeed.

Write the decisions down, briefly, and set them where next quarter's review will find them. Then make the weekly reviews carry them out: each decision becomes a line in the register or a change to a recipe or a habit, and the weekly review checks progress. A quarterly review whose decisions are never acted on is merely a pleasant afternoon of thinking. Pleasant afternoons are nice. They are not steering.

This quarter, book a half day for the review. Read the record. Ask the four questions of every project, recipe, tool and habit. Write down the decisions, especially the endings. Then hand them to the weekly reviews to carry out. Three months from now, read them back and see how many happened. The proportion will tell you how much your operation is steered and how much it simply moves.

Four questions for the quarter The record shipped lists decision log incident notes notebook last quarter decisions Keep it works End no longer earns it the hard one Start something missing one or two at most Change right idea, wrong practice Decisions written down for next review Weekly review carries it out ASK OF EVERY project recipe tool habit The week asks what to do next. The quarter asks what to stop doing. a half day, away from the usual desk
Fig 74 · The Quarterly Review. The quarterly review reads the record, then asks keep, end, start or change.
Chapter 75 · Part VIII

Killing Projects Kindly

Some projects should end before they are finished. The market moved, the client changed their mind, the idea turned out to be less interesting than it looked, or you simply have better things to do with the attention. Ending a project is one of the most useful decisions an operator makes, and one of the most avoided, because it feels like failure. It is not. It is pruning, and a well-pruned operation grows more than an overgrown one.

The question to ask is not whether the project is good, but whether it is still worth what it costs. Every active or parked project costs something: attention in the register, guilt when you look at it, threads that need tending, a slot that could hold something else. If the project's likely value no longer justifies that cost, it should end. The answer will sometimes be keep going, and that is fine. When it is end it, the next question is how.

End it well. That means four things. Finish or close every thread, so nothing is left running. Write a short closing note: what the project was, what it achieved, why it ended, and anything worth keeping. Archive the material somewhere you could find it again, rather than deleting it. And tell anyone who needs to know, a client, a collaborator, users of a tool, with enough notice and explanation that they are not surprised.

Ending a project well is a skill. Letting one die slowly is a habit.

The closing note matters more than it seems. Projects that end abruptly leave behind confusion: half-finished threads, orphaned files, a vague sense of unfinished business. Projects that end with a note leave behind a clean record. If you ever want to revive the idea, the note tells you where you got to and why you stopped. If you never do, the note lets you stop thinking about it, which is its own kind of relief.

Look for what can be salvaged. Ended projects often contain useful pieces: a recipe that worked, a component that could be reused, a piece of research that answers a question elsewhere, a lesson worth adding to memory. Before archiving, spend ten minutes extracting these. They are the dividends of the project, and they often justify the effort even when the project itself did not reach its goal.

Agents can help with the mechanics, closing threads, archiving files, drafting the closing note, listing reusable pieces. They should not make the decision. Ending a project is exactly the kind of judgement that belongs to the operator, because it depends on priorities and commitments only you can weigh.

The hardest projects to end are the ones you started with enthusiasm and have invested a lot in. The sunk cost pulls at you: so much work, it would be a waste to stop now. But the work is spent either way. The only question is whether the next hour is better spent on this project or something else. Answer that question honestly and the sunk cost loses its grip.

This quarter, find one project that should end and end it well: threads closed, note written, material archived, people told. Notice how the register feels afterwards. Lighter, almost certainly. That lightness is attention returned to you.

Ending a project kindly What it costs attention guilt on sight threads to tend a slot for better Worth the cost? yes Keep going no End it well sunk cost: spent either way Close threads none left running Closing note what, why, keep Archive do not delete Tell people with notice Salvage first: recipes · components · research · lessons for memory Ending a project well is a skill. Letting one die slowly is a habit.
Fig 75 · Killing Projects Kindly. A project that no longer earns its cost is ended well in four steps, salvage first.
Chapter 76 · Part VIII

Energy Is the Real Budget

The usual way to think about capacity is in hours: how many do I have, how should I spend them? For an operator with agents, hours are the wrong unit. The agents can work any number of hours. What limits the operation is not your time but your energy: the quality of attention you can bring to deciding and reviewing. An hour of sharp attention and an hour of foggy attention are both hours. They are not the same budget.

Energy varies across the day, across the week and across longer stretches. Most people have a few hours each day when they think clearly, decide well and review carefully, and many more hours when they can do useful work but not their best work. The precise pattern differs by person. What does not differ is that the best hours are scarce and the work that needs them is the most valuable work in the operation.

So match the work to the energy. Think of the operator's tasks on two axes: how much value they create and how much energy they demand. High-value, high-energy work, making important decisions, reviewing risky changes, writing briefs for difficult tasks, belongs in your best hours. Low-value, low-energy work, tidying files, closing threads, answering routine messages, belongs in the hours when you are not at your sharpest. The most common mistake is the reverse: spending the first sharp hour of the morning on email and the last foggy hour of the afternoon on a high-risk review.

Hours are what you have. Energy is what you spend.

Agents make this matching easier, because so much routine work can now be delegated. But they also create a new energy drain: the constant low-level demand of monitoring threads, reading reports and responding to notifications. Each demand is small. Together they can consume a surprising share of your best energy if you let them into your best hours. Protect those hours. Batch the monitoring into the less precious parts of the day.

Energy also has a weekly and seasonal shape. Many operators find that they decide better early in the week and review better midweek, or that certain months are reliably more draining than others. Notice your own patterns and plan around them. Put the big decisions where you are strongest. Lighten the load where you know you will be weaker.

There is a temptation to treat energy as a matter of willpower, to push through low-energy periods by force. This works occasionally and fails as a strategy. Decisions made on low energy are worse, reviews done on low energy miss things, and the cost of those misses arrives later and larger. It is better to do less, well, than more, badly, particularly when the agents are perfectly happy to keep the doing going without you.

This week, rate your energy each hour on a simple scale of high, medium or low. At the end of the week, look at the pattern and identify your best hours. Then move your most important decisions and reviews into them, and move the routine into the rest. It is a small rearrangement. Its effect on the quality of your work is not small at all.

Match the work to the energy Best hours big decisions risky reviews hard briefs Any time quick yes/no short replies Foggy hours tidy files close threads routine messages Batch it monitoring notifications report skims high low value low ENERGY IT DEMANDS high The common mistake sharp hour: email foggy hour: risky review → swap them Hours vs energy agents work any hours your attention cannot do less, well RATE EACH HOUR FOR A WEEK M H H M L L M H M L Hours are what you have. Energy is what you spend.
Fig 76 · Energy Is the Real Budget. Tasks sorted by value and energy demand, so the best hours get the hardest work.
Chapter 77 · Part VIII

Attention Has a Ceiling

With agents, the number of things you could have running at once is effectively unlimited. Ten threads, twenty, fifty, each working on something useful. The constraint is not the agents but you, because each thread eventually needs your attention: to brief it, unblock it, review its output and decide what happens next. Your attention has a ceiling, and the ceiling is much lower than the number of threads you could run.

Picture three numbers. The threads you could run is very large, limited only by ideas and budget. The threads you can watch, keeping track of their progress and catching them when they drift, is much smaller, perhaps a handful at a time. The threads you can properly review, giving their output the careful attention it needs before it ships, is smaller still. That last number is the real capacity of the operation. Everything above it produces output that will either wait in a queue or ship under-reviewed.

Most operators, when they first get agents, run well above the ceiling. It feels efficient: more threads, more output, more progress. In practice it produces a backlog of unreviewed work, a growing sense of being behind and a creeping lowering of review standards as you try to keep up. The work does not get done faster. It gets done worse, and then some of it has to be done again.

Run as many threads as you can review. Not one more.

Finding your ceiling takes some honest observation. For a week, note how many threads you run each day and how many of their outputs you review properly versus skim. The point at which skimming starts is roughly your ceiling. For many operators it is somewhere between three and six substantial threads a day, though it varies a great deal with the kind of work and with how well the briefs are written.

Better briefs raise the ceiling. A thread with a clear definition of done and strong evidence requirements is quicker to review than one without, so you can review more of them. Small changes raise it too, because they take less review each. Good recipes raise it, because familiar work is quicker to assess. These are all ways to get more from the same attention, and they are much more effective than simply trying harder.

Delegating review raises it a little, but only a little. A reviewer agent can do a first pass on output, flagging problems and summarising changes, and this saves time. It does not remove the need for your review, because your review is where your judgement and ownership enter the work. Think of a reviewer agent as an assistant that prepares the papers, not one that signs them.

When you are above the ceiling, the fix is not to work longer but to run fewer threads. Park projects. Sequence work that was parallel. Let some ideas wait in the queue. It feels like slowing down. It is actually speeding up, because work that is reviewed properly ships once rather than twice.

This week, find your ceiling. Then, for the following week, run no more threads than it allows. Notice whether more or less actually ships. Most operators are surprised by the answer, and pleasantly.

Attention has a ceiling Threads you could run ideas and budget: no limit Threads you can watch a handful Threads you can review 3–6 a day your ceiling above it: a queue, or work shipped under-reviewed RAISE THE CEILING Better briefs quicker to review Small changes less per review Good recipes familiar work Reviewer agent only a little above it? run fewer: park, sequence, let ideas queue Run as many threads as you can review. Not one more.
Fig 77 · Attention Has a Ceiling. Threads you could run, can watch and can review; the last is the real capacity.
Chapter 78 · Part VIII

The Hours You Are Good For

Nobody is good for eight straight hours of high-quality judgement. Most people are good for a few, scattered through the day in a pattern that is surprisingly consistent once you notice it. An operator who plans the day around those hours gets more from them than one who treats every hour as interchangeable, and a great deal more than one who spends them on the wrong things.

The day naturally cycles between three kinds of time. Deep work: the hours when you can concentrate fully, make difficult decisions and review complex changes with care. Light work: the hours when you can do useful things that do not need full concentration, such as tidying, routine reviews, closing threads, writing simple briefs. And recovery: the time between, when you are not working at all, and your capacity for the next deep stretch is being rebuilt. All three are necessary. The mistake is to treat light work or recovery as failures to do deep work.

For most people, deep work comes in blocks of an hour or two, a few times a day at most. Some find their best block in the early morning; others late in the morning; a few in the evening. The pattern is personal, but it is usually stable, and it is worth learning. Once you know your deep hours, protect them. No messages, no monitoring, no light work that can be done later. Put the most demanding part of the day's three outcomes there.

You cannot work deeply all day. You can arrange the day so that the deep hours do the deep work.

Light work fills much of the rest, and agents have changed its character. Much of what used to be light work, formatting, searching, routine drafting, is now delegated. What remains is mostly monitoring and quick review: checking threads, approving simple changes, reading digests. This is real work and it should be scheduled, not allowed to leak into the deep hours where it does the most damage.

Recovery is the part people skip, and the part that determines whether the next deep block is any good. A walk, a meal, a conversation, an hour doing something unrelated. Checking threads on your phone during a break is not recovery; it is light work in a different chair. The agents will be fine without you for an hour. You will be better for having been away.

The cycle repeats through the day. Deep, light, recovery, then deep again if you have another block in you, light, recovery, close. Many operators find that two deep blocks a day is a good day and three is exceptional. Planning for more than that tends to produce blocks that are nominally deep and actually shallow.

This week, track which hours you do your best thinking. Then plan the following week around them: the hardest decisions and reviews in the deep blocks, routine work in the light ones, and genuine breaks in between. Treat the breaks as appointments, not as leftovers. At the end of the week, compare the quality of your reviews and decisions with the week before. The hours were the same. What you put in them was not.

A day shaped around the deep hours Deep Light Recovery 2 h 1.5 h morning review close 08 09 10 11 12 13 14 15 16 17 18 Deep work hardest decisions complex reviews no messages Light work monitoring quick approvals simple briefs Recovery walk, meal, people phone ≠ recovery an appointment Two deep blocks is a good day. Three is exceptional.
Fig 78 · The Hours You Are Good For. A day cycling between deep, light and recovery time, with two deep blocks.
Chapter 79 · Part VIII

Rest Is Maintenance

There is a quiet assumption, common among people who work alone and especially among people with tireless agents, that rest is what you earn after the work is done. The work is never done, so the rest never quite arrives. The agents keep running, the threads keep finishing, there is always one more thing to review, and the evenings and weekends gradually fill with small acts of checking. It feels diligent. It is a slow breakdown of the most important part of the machine.

Rest is not a reward. It is maintenance. The operator's judgement, the thing that makes all the other parts work, depends on rest in the same way that a machine depends on oil. Without it, decisions get worse, reviews miss things, briefs get vaguer, and the cost of those failures appears later, in rework, incidents and the special exhaustion of fixing problems you caused while tired. Good judgement lives in the overlap between work and rest. Remove the rest and the overlap disappears.

The agents make this harder in a particular way: they never need rest, so they never give you a natural stopping point. A human team goes home at the end of the day and the work pauses. Agents can run all night, and the knowledge that something might have finished creates a pull towards checking. The pull is strongest precisely when you most need to resist it, late in the evening and at weekends, when your judgement is lowest and the cost of acting on something you half-read is highest.

The machine can run without rest. The operator cannot. Plan accordingly.

So build rest into the structure, not into the leftovers. The daily close ends the working day explicitly. The weekend is a weekend: no reviews, no briefs, no checking, unless something is genuinely on fire, and almost nothing is. Holidays are holidays, prepared for in advance with parked projects and clear handoffs so that nothing needs you while you are away. These boundaries are not luxuries. They are part of the operating system.

Make the boundaries easier to keep by designing for them. Overnight and weekend agent work should be limited to well-briefed, low-risk tasks that can safely wait for review. Notifications should be off outside working hours. Anything that might genuinely need you urgently should be rare, well defined and routed through a single channel, so that silence on that channel means you can truly stop.

It helps to notice the difference between rest and distraction. Scrolling, half-watching something while thinking about work, or doing light admin on a Sunday are not rest. Rest is time in which the work is genuinely not in your head: exercise, people, sleep, absorbing activities that have nothing to do with the operation. It is when the mind does its slow background work of making sense of things, and many of the best decisions arrive, unbidden, on the Monday after a proper weekend.

This week, set a firm boundary: one evening and one full day with no checking, no briefing, no reviewing. Prepare for it at the close beforehand. Notice what happens to your first review afterwards. Most operators find it is noticeably sharper. That sharpness is the maintenance paying off.

Rest is maintenance Work briefs reviews decisions Rest sleep people exercise Judgement remove the rest and the overlap disappears Without rest worse decisions missed review defects vaguer briefs rework and incidents Not rest scrolling half-watching, thinking work Sunday admin BUILD IT INTO THE STRUCTURE Daily close ends the day Weekend no checking Holidays prepared Alerts off at night One channel fires only The machine can run without rest. The operator cannot.
Fig 79 · Rest Is Maintenance. Good judgement lives where work and rest overlap; remove rest and it goes.
Chapter 80 · Part VIII

A Pace You Can Keep

The most productive operators are not the ones who have the most impressive weeks. They are the ones who have good weeks, consistently, for years. A pace you can keep beats a pace that impresses, because the operation compounds. Recipes improve, memory deepens, the register stays healthy, judgement sharpens, and each year's work is built on the previous year's foundations. Bursts of heroic effort followed by recovery do not compound. They oscillate.

Agents tempt operators towards bursts. The capacity is right there: you could run twenty threads, work through the night and ship a whole product in a weekend. Sometimes that is the right call, for a genuine deadline or an opportunity that will not wait. As a habit, it is corrosive. Each burst borrows from the energy and judgement of the following weeks, and the borrowing is repaid with interest in the form of tired decisions, missed reviews and the slow accumulation of problems that were waved through in the rush.

A sustainable pace starts with a steady day: the review, sessions and close described earlier, with today's three as the target and the attention ceiling respected. Steady days add up to a steady week, with its review, its deck-clearing and its honest count of what shipped. Steady weeks add up to a steady year, with quarterly reviews that change direction on purpose and an operation that is noticeably better at the end than at the start. The chain is unremarkable at every link. The result at the end is remarkable.

Heroics make good stories. Habits make good years.

What does a sustainable pace feel like from the inside? Mostly, calm. You know what you are working on and why. You know what shipped last week. You are not behind on review, because you only run what you can review. You stop at the end of the day without anxiety, because the close has captured everything. You take weekends and holidays without dread, because the operation is designed to pause.

They usually are. The test is not how hard it feels but what it produces over months. A calm operator who ships three well-reviewed things a day, every day, outproduces a frantic one who ships ten things one week and nothing the next while recovering from the first. And the calm operator's work tends to be better, because it was reviewed by someone who was paying attention.

Pace is also something you can adjust deliberately. There are seasons when a little more push is right, and seasons when a little less is wise. The quarterly review is the natural place to decide this: what pace does the coming quarter need, and what will I do to make sure it is sustainable? Deciding the pace on purpose is very different from having it decided for you by whatever is shouting loudest.

This week, ask yourself honestly: could I keep up this week's pace for a year? If the answer is no, find the one thing that makes it unsustainable and change it. Then ask again next week. The goal is a yes you believe. When you get there, you will have the most valuable thing an operator can build: a machine that runs well, indefinitely, with you still in the chair.

Bursts oscillate, habits compound what shipped, cumulative weeks → Steady pace heroic bursts sprint, then recover borrowed judgement Steady day review · close Steady week review · count Steady year quarterly reviews test: could I keep this week’s pace for a year? Heroics make good stories. Habits make good years.
Fig 80 · A Pace You Can Keep. A steady pace compounds past heroic bursts; days build weeks build years.
Part IX

Tools, Mistakes and Incidents

Upkeep, scars and the runbook.

Chapter 81 · Part IX

Fewer Tools, Better Kept

There has never been a better time to collect tools. New agent products, extensions, connectors, plug-ins and services arrive every week, each promising to remove some friction from your day. Many of them are genuinely good. Most operators try a great many and keep more than they use. The result is a toolbelt so heavy that it slows them down: too many places to look, too many subscriptions to manage, too many integrations that might break, and too little depth of knowledge in any one of them.

The better approach is fewer tools, better kept. A small set that you know deeply, use daily and maintain properly will serve you better than a large set that you know shallowly and use occasionally. Depth matters because the value of a tool comes mostly from knowing it well: its shortcuts, its limits, its failure modes, the way it behaves when something goes wrong. That knowledge takes time to build, and it is spread thin when the toolbelt is wide.

Think of it as a funnel. At the top, every tool you have ever tried. Below that, the ones you actually use weekly. At the bottom, your toolbelt: the small set you rely on, know properly and would notice immediately if they broke. If yours has thirty, most of them are probably at the top of the funnel pretending to be at the bottom.

A tool you know well beats two you half know.

Every tool has a carrying cost, even if it is free. It needs updating. Its permissions need reviewing. Its integrations need watching. It occupies space in your mind and in your memory files. When it changes, you need to learn the change. When it fails, you need to notice and respond. For a tool you use daily, this cost is easily justified. For a tool you use once a month, it often is not.

So be slow to adopt and quick to drop. When a new tool appears, ask what specific problem it would solve that your current tools do not. If you cannot name one, let it pass, however impressive it looks. If you can, try it on a small, contained piece of work before letting it into your daily rhythm. And when a tool stops earning its place, remove it, cleanly, including its permissions and any memory or recipes that refer to it.

This principle sits comfortably alongside the small machine rule from the first part of this book. A small toolbelt is part of a small machine, and a small machine is one you can hold in your head on a bad day. When something goes wrong in an operation with five tools, you know where to look. When something goes wrong in an operation with thirty, you may spend the morning finding out which tool is involved.

This week, list every tool, service and integration your operation currently uses. Mark each with how often you used it in the last month. Anything you did not use is a candidate for removal. Anything you used only once or twice deserves a hard look. Then remove at least one tool, cleanly. Your toolbelt will be lighter, and you will know the rest of it a little better for the space.

Fewer tools, better kept Every tool you tried still installed, half known Used weekly worth a hard look Your toolbelt known deeply Carrying cost, even if free updates to apply permissions to review integrations to watch space in memory files changes to learn failures to notice subscriptions to pay SLOW TO ADOPT QUICK TO DROP A real need? tools miss it Small trial contained work Daily rhythm earned its place Remove cleanly access, memory A tool you know well beats two you half know.
Fig 81 · Fewer Tools, Better Kept. Tools narrow from everything tried to a small toolbelt known deeply.
Chapter 82 · Part IX

The Toolbelt Audit

Tools accumulate silently. A connector added for one project stays connected after the project ends. A browser extension installed for a trial is still running months later. A subscription renews because nobody looked at it. An agent integration still has access to a folder you no longer use. None of this is dramatic. All of it is clutter, some of it costs money, and a little of it is a security risk. The remedy is a regular audit.

The toolbelt audit is a short, structured review, once a quarter, of everything in your operation that is a tool: software, services, subscriptions, connectors, integrations, extensions, scripts and automations. It has three steps. List everything. Check how each is used. Decide what to keep, change or remove.

Listing is the step that reveals the most. Most operators, when they first do it, find tools they had forgotten existed. Look everywhere: installed applications, browser extensions, connected services in each of your main platforms, agent connectors and their permissions, scheduled jobs, automations, subscriptions on your statements. Agents can help with parts of this, particularly listing integrations and scheduled jobs, but you will need to look at some places yourself.

Every tool you forgot you had is a tool that can surprise you.

Checking use means asking, for each item: when did I last use this, for what, and would I notice if it disappeared? Also ask what it can access. A tool you rarely use that has broad access to your files or accounts is a much bigger concern than one you use daily with narrow access. Pay particular attention to anything connected to agents, because agents act on the access they are given, and stale access is an invitation for something to go wrong in a way you will not anticipate.

Deciding is the point of the exercise. Keep the tools you use and would miss. Change the ones you use but have configured badly: too much access, outdated settings, overlapping with another tool. Remove the ones you do not use, cleanly: revoke their access, cancel subscriptions, delete integrations, and update any memory files or recipes that mention them. Removing a tool from your life but leaving its access in place is not removal. It is abandonment.

The audit is also a good moment to check the health of the tools you keep. Are they up to date? Are their settings what you think they are? Have they changed in ways that affect how you use them? Is there anything about them you have been meaning to learn? A small amount of attention here keeps the toolbelt sharp rather than merely present.

Record the audit briefly in the notebook or decision log: what you removed, what you changed, anything surprising. Next quarter, start by reading the last audit. Over time, these records show you how your toolbelt evolves and whether it is growing or shrinking, which is a useful indicator of whether the small machine rule is holding.

This quarter, run the audit. List, check, decide. Aim to remove at least a few items and to tighten the access of at least one. It takes an afternoon. It leaves you with an operation that you understand better and that has fewer places for something unexpected to happen.

The quarterly toolbelt audit 1 List everything installed apps browser extensions connected services agent connectors scheduled jobs automations subscriptions 2 Check each one last used? for what? would I notice if it disappeared? what can it access? agents act on access 3 Decide Keep used, would miss Change access, settings Remove cleanly revoke, cancel USE × ACCESS narrow access broad access used daily rarely used fine tighten it remove? biggest concern Record the audit removed, changed, surprises read it first next quarter Every tool you forgot you had is a tool that can surprise you.
Fig 82 · The Toolbelt Audit. List, check and decide every tool; rarely used tools with broad access worry most.
Chapter 83 · Part IX

Upkeep Is Scheduled

Every operation needs upkeep. Tools need updating. Dependencies need upgrading. Memory files need pruning. Recipes need refreshing. Credentials need rotating. Backups need testing. Domains, certificates and subscriptions need renewing. None of this is interesting, none of it is urgent until it suddenly is, and all of it is easy to postpone. The only reliable way to get it done is to schedule it.

Unscheduled upkeep happens in one of two ways: never, or in a crisis. The certificate expires and the site goes down. The dependency is three major versions behind and the upgrade is now a week's work instead of an hour's. The credential that should have been rotated is exposed. The backup that was never tested turns out not to restore. Each crisis costs far more than the scheduled upkeep would have, and each arrives at a time chosen by the problem rather than by you.

Scheduled upkeep turns these crises into routine. The cycle is simple. Schedule each upkeep task at a sensible interval. When it comes round, do the update. Then verify that the update worked and nothing broke. Then the next occurrence is already on the schedule. The cycle's starting point, the schedule, is where most of the value lies, because it is what makes the rest happen at all.

Maintenance on the calendar is cheap. Maintenance in an emergency is not.

Build an upkeep calendar. Some tasks are weekly: checking that automated jobs ran, glancing at error logs. Some are monthly: updating dependencies, pruning memory files, reviewing open threads. Some are quarterly: the toolbelt audit, rotating credentials, testing a backup restore, reviewing recipes. Some are annual: renewing domains, reviewing every subscription. Write these down with their intervals and put them on whatever calendar you actually look at.

Agents are excellent at upkeep work, and this is one of the best places to use them. A recipe for dependency updates can ask an agent to upgrade, run the full test suite, and report any failures. A recipe for memory pruning can ask an agent to flag stale entries. A recipe for checking backups can walk through a restore to a test location. Scheduled agent jobs can run some of these automatically. You still need to review the results, but the review is quick when the work is routine and the recipe is good.

The verify step deserves emphasis. Upkeep that is done but not verified is half done. An update that broke something quietly, a rotation that left an old credential still active, a backup that runs but produces empty files: each looks complete until the moment you need it. Every upkeep task should end with a check that it actually achieved its purpose.

Keep the upkeep calendar small and realistic. If it lists fifty tasks, many will be skipped, and the habit will erode. Start with the handful of tasks whose failure would hurt most, schedule those, and add others only once the habit is solid. A short list done reliably is better than a long list done occasionally.

This week, write your upkeep calendar. List the tasks, set their intervals, and put the next occurrence of each on your calendar. Then do one overdue task today. That small act of catching up is the beginning of never having to catch up again.

Upkeep belongs on the calendar Schedule at an interval Update do the task Verify it worked done, not verified = half done THE UPKEEP CALENDAR weekly jobs ran? · error logs monthly deps · prune memory · threads quarterly audit · rotate keys · test restore annual domains · every subscription start with the few whose failure hurts most UNSCHEDULED, IT ARRIVES AS A CRISIS Certificate expires: site down Dependency majors behind: a week Backup untested: no restore Maintenance on the calendar is cheap. In an emergency it is not.
Fig 83 · Upkeep Is Scheduled. Schedule, update, verify: a small calendar turns upkeep crises into routine.
Chapter 84 · Part IX

Permissions as Policy

Agents act on the permissions they are given. Most agent tools now offer a range of modes, from asking before every action to acting freely within certain bounds, along with ways to allow or deny specific commands, files and services. Many operators handle permissions reactively: approving requests one at a time as they pop up, gradually allowing more as they get tired of the prompts. The result is a set of permissions that nobody actually designed, and that may allow far more, or far less, than intended.

A better approach is to treat permissions as policy. Decide once, deliberately, what agents may do without asking, what they must ask about first, and what they may never do. Write that policy down. Configure your tools to enforce it. Then you can stop making the same small decisions dozens of times a day, and you can trust that the boundaries reflect your considered judgement rather than your level of irritation at four in the afternoon.

The policy has three tiers. Allowed without asking: actions that are safe, reversible and routine, such as reading files in the project, running the test suite, editing files within the project's scope. Ask before acting: actions that are significant or harder to reverse, such as installing dependencies, changing configuration, running commands with side effects outside the project, network calls to unfamiliar places. Never allowed: actions that are destructive, irreversible or outside the operation's boundaries, such as deleting outside the project, touching credentials, pushing directly to production, accessing personal files.

Decide your permissions once, when you are calm, so you do not decide them a hundred times when you are busy.

Different projects may need different policies. A throwaway experiment can be generous. A client project with sensitive data should be strict. A production system should be stricter still. Many tools let you set permissions per project, which is exactly what you want: the policy matches the risk.

Review the policy at the toolbelt audit, or when something goes wrong. Look at what you have allowed and ask whether each item is still appropriate. Look at what agents keep asking about and ask whether those requests should be pre-approved or should remain prompts. Look at any incidents and ask whether a permission change would have prevented them. Permissions that are never reviewed tend to drift towards too generous, because each individual approval seems harmless.

Pay particular attention to permissions for unattended work: overnight jobs, scheduled tasks, anything that runs without you watching. These should be the tightest of all, because there is no one to catch a mistake as it happens. A good rule is that unattended agents get only what the specific task needs, and nothing more.

There is a pleasant side effect to a good policy: fewer interruptions. When routine actions are pre-approved and dangerous ones are blocked, the only prompts you see are for genuinely significant actions, and those are the ones worth your attention. The stream of trivial approval requests disappears, and with it the habit of clicking yes without reading.

This week, write your permissions policy for your main project: three short lists. Then check your tool's actual configuration against it and fix any mismatch. Notice how many prompts disappear, and how much more attention you pay to the ones that remain.

Permissions as policy, in three tiers Policy, decided once, when calm Allowed without asking safe · reversible · routine read project files · run tests · edit in scope Ask before acting significant · harder to reverse install deps · change config · unfamiliar network Never allowed destructive · irreversible delete outside · credentials · push to production risk PER PROJECT Experiment generous Client data strict Production strictest Unattended jobs only what it needs Decide once, when calm, not a hundred times when busy.
Fig 84 · Permissions as Policy. Agent permissions as a written policy: allowed, ask first and never, per project.
Chapter 85 · Part IX

Fragile Wins Cost More

Sometimes the quickest way to get something working is a hack. A hard-coded value. A manual step that someone has to remember. A script that works on your machine but nowhere else. An agent workaround that suppresses an error rather than fixing it. The hack gets you a win today, and wins feel good. But fragile wins carry interest, and the interest is paid later, usually at the worst possible moment and usually more than the original win was worth.

The problem with fragile wins is not that they fail. Everything fails eventually. It is that they fail unpredictably and silently. A hard-coded date works until the date passes. A manual step works until you forget it. A suppressed error hides a problem until the problem grows large enough to break through. When they fail, you often do not know they were fragile, because the hack was made months ago, perhaps by an agent, and nobody wrote it down.

Think of solutions on two axes: how quickly they deliver results now, and how fragile they are. A solution that is fast and fragile is tempting and dangerous. A solution that is slow and robust is safe and sometimes overkill. The corner you usually want is fast enough and robust enough: not the quickest possible fix, and not the most bulletproof, but one that will keep working without anyone remembering it exists.

The hack saves an hour today and costs a day on the worst day of the month.

Agents can produce fragile wins very quickly, particularly when asked to make something work under pressure. An agent told to get the tests passing may find the fastest path, which is not always the right one. Watch for the signs in review: hard-coded values that should be configuration, errors caught and ignored, special cases added for a single input, comments saying temporary or fix later. Each of these is a fragile win in the making.

When you do accept a fragile win, and sometimes you should, because a deadline is real or the stakes are low, record it. Write a line in the decision log or the project memory: what the hack is, why it was accepted, and what the proper fix would be. Add a reminder to revisit it. A recorded hack is a known debt with a repayment plan. An unrecorded one is a trap waiting for someone, usually you.

Some fragility comes from the operation itself rather than the work: a step that only you know how to do, a process that depends on your memory, an integration nobody else could repair. These are fragile wins at the level of the machine, and they matter even more for a solo operator, because there is no colleague to fall back on. The next chapters on incidents and runbooks are largely about turning this kind of fragility into something sturdier.

This week, look for fragile wins in one project. Search for hard-coded values, suppressed errors, temporary comments and undocumented manual steps. Record each one you find, and fix the one most likely to cause trouble. It is not exciting work. It is the kind of work that makes next month uneventful, which is the best kind of month to have.

Fast enough, robust enough high fragility low slow SPEED NOW fast The hack tempting, dangerous Bulletproof safe, sometimes overkill Robust enough fast enough, and keeps working fix it properly hard-coded date works until the date passes Warning signs in review hard-coded values errors caught, ignored special case, one input "temporary", "fix later" undocumented manual step Accepting? Record it what, and why accepted the proper fix a date to revisit The hack saves an hour today and costs a day on the worst day of the month. a recorded hack is a debt with a plan; an unrecorded one is a trap
Fig 85 · Fragile Wins Cost More. Solutions placed by speed and fragility, aiming for fast enough and robust enough.
Chapter 86 · Part IX

Mistakes Are Data

Mistakes are inevitable in any operation, and an operation with agents makes a particular variety of them. A brief that was misread. A change that broke something unexpected. A review that missed a problem. A release that went to the wrong place. A message sent with an error in it. Each one is annoying and occasionally costly. Each one is also data, and an operator who treats mistakes as data improves much faster than one who treats them as embarrassments to be fixed and forgotten.

When a mistake happens, there is a fork. One path leads to hiding or forgetting: fix it quietly, move on, hope it does not recur. This is the natural path, because mistakes are uncomfortable and fixing them is satisfying. The other path leads to learning: fix it, then ask why it happened and what would prevent it next time. The second path takes a few more minutes. It is the only one that makes the operation better.

The value lies in the why. Most mistakes have causes that are upstream of the visible error. The agent misread the brief because the brief was ambiguous. The change broke something because there was no test covering that path. The review missed the problem because it happened at the end of a long day. The release went to the wrong place because the deployment process has a confusing step. Fix only the visible error and the upstream cause remains, ready to produce the next mistake.

A mistake fixed is a problem solved. A mistake understood is a class of problems solved.

Collect mistakes somewhere. A section in the notebook, a simple log, or incident notes for the more significant ones, which the next chapter describes. Over time, the collection reveals patterns: the same kind of brief keeps being misread, the same area of code keeps breaking, the same time of day keeps producing errors. Patterns are far more useful than individual mistakes, because they point to systemic fixes that prevent many future errors at once.

Agents make mistakes too, and they deserve the same treatment. When an agent does something wrong, resist the urge to blame the agent and move on. Ask what in the brief, the memory, the permissions or the checks allowed it to happen. Usually there is something, and fixing it improves every future run. An agent's mistake is often your system telling you where it is incomplete.

Your own mistakes deserve the same curiosity, and here a little self-compassion helps. Operators who are hard on themselves tend to hide their mistakes, even from their own notebooks, because writing them down feels like dwelling on failure. But the notebook is not a report card. It is a tool for getting better. A mistake written down calmly and examined is a gift to your future self. A mistake buried is a gift to nobody.

This week, every time something goes wrong, however small, write one line: what happened and why you think it happened. Do not try to fix the causes yet. Just collect. At the weekly review, read the lines and look for a pattern. If you find one, fix its cause. You will have turned a week of annoyances into one lasting improvement, which is a very good trade.

Two paths after a mistake Mistake Fix the error Hide or forget fixed quietly Cause remains ready to repeat the next mistake Ask why a few more minutes Upstream cause ambiguous brief missing test tired review confusing step Log one line what, and why Find the pattern at the weekly review Systemic fix a class of problems A mistake fixed is one problem solved. Understood, it solves a class.
Fig 86 · Mistakes Are Data. After a fix, hiding leaves the cause; asking why leads to a pattern and a lasting fix.
Chapter 87 · Part IX

The Incident Note

When something significant goes wrong, a release that broke a feature, a lost piece of data, a mistake that reached a client, an agent action that should not have happened, write an incident note. Not a lengthy report, not a formal post-mortem, just a short, structured note that captures what happened, why, and what will change. It takes fifteen minutes. It is the most effective way an operator has of making sure the same thing does not happen twice.

The note answers three questions in order. What happened: a factual account, with times if they matter. What was the effect, who noticed, how was it resolved? Keep this part plain and specific. Why it happened: the causes, both the immediate one and the upstream ones. The immediate cause might be the migration deleted rows with empty names. The upstream causes might be the brief did not mention empty names, there was no test for that case, and the change was automerged because it was labelled low risk. What changes: the specific actions you will take to prevent a recurrence. Each action should be concrete and assigned to a time, even if the only person to assign it to is you.

The middle question is where the value lies, and it is worth spending most of the fifteen minutes there. A useful technique is to keep asking why until you reach something you can change in the system rather than in the moment. The agent deleted the rows: why? Because the brief did not say not to: why? Because I did not know empty names existed: why? Because the data has never been profiled: that is something you can fix.

An incident without a note is an incident you have agreed to have again.

The what-changes section should have a small number of real actions, not a long list of good intentions. Add a test. Update the recipe. Change the automerge conditions. Add a line to the project memory. Profile the data. Each one should be done within days, not weeks, and the note should be updated when it is. An incident note whose actions are never completed is a record of a lesson not learned.

Keep incident notes together, in one folder or one file, with a short index. Read them at the quarterly review. Patterns across incidents are often more revealing than any single incident: the same kind of check keeps being missing, the same kind of change keeps going wrong, the same part of the operation keeps being fragile. Those patterns point to the improvements that matter most.

Agents can help write incident notes. Give one the timeline, the relevant diffs and the brief, and ask it to draft the three sections. Its draft will often identify contributing causes you had not considered. But review the draft carefully, especially the why, because the most important causes are often things the agent cannot see: your assumptions, your time pressure, your knowledge of the client.

This week, if anything significant goes wrong, write the note: what happened, why, what changes. If nothing does, write one for the most significant incident of the past month. Then carry out the actions. The note is only the first half. The changes are the point.

Anatomy of an incident note incident-note.md · 15 minutes What happened facts, times, effect who noticed, how resolved Why it happened immediate cause upstream causes: most time here What changes few real actions, dated ASK WHY UNTIL IT IS THE SYSTEM The agent deleted the rows why? The brief did not say not to why? I did not know empty names existed why? The data was never profiled fixable: profile the data, add a test An incident without a note is an incident you have agreed to have again.
Fig 87 · The Incident Note. An incident note in three parts, with a why-ladder down to a fixable cause.
Chapter 88 · Part IX

Blameless, Even Alone

Teams that run incidents well usually adopt a principle called blamelessness: the purpose of reviewing an incident is to understand and improve the system, not to find someone to blame. People who fear blame hide information, and hidden information makes the system worse. It might seem that this principle does not apply to a solo operator, since there is nobody to blame but yourself. In fact it applies with special force, because the person most likely to blame you is you.

Self-blame is corrosive in a specific way. When something goes wrong and your first reaction is I should have caught that, how stupid, you are likely to do one of two things. Either you fix it quickly and move on, avoiding the discomfort of looking closely, which means you never learn the cause. Or you dwell on it, replaying the mistake, which drains energy and attention without improving anything. Neither leads to the calm, curious analysis that actually prevents recurrence.

Blamelessness lives in the overlap between honesty and kindness. Honest, because you must look clearly at what happened, including your own part in it, without minimising or excusing. Kind, because you approach that clear look with the same goodwill you would offer a colleague who made the same mistake: an assumption that they were doing their best with what they knew at the time, and a focus on what would have helped them do better.

Look at the system that let it happen, not the person who was standing nearest.

In practice, this means writing incident notes and mistake logs in a particular tone. Not I carelessly forgot to check the empty case but the empty case was not checked; the brief did not mention it and there was no test. Both are true. The second points at things you can change. The first points only at a feeling. It is not about avoiding responsibility. You are still the operator, and the name on the door is still yours. It is about directing responsibility where it can do some good.

Blamelessness extends to agents, too, in an odd way. When an agent makes a mistake, it is tempting to treat it as the agent's fault and move on, or to lose trust in agents generally. A blameless approach asks what in the system allowed the mistake: the brief, the permissions, the checks, the memory. That question almost always has a useful answer. The agent was wrong almost never does, because you cannot fix an agent's character, only the conditions it works in.

There is a practical benefit beyond better learning: you feel better, and an operator who feels better makes better decisions. Self-blame is a form of stress, and stress narrows attention and degrades judgement. An operator who can look at a mistake calmly, learn from it and move on is in much better shape for the next decision than one who carries the mistake around for days.

This week, reread the last few things you wrote about your own mistakes, whether incident notes, notebook entries or simply what you said to yourself. Notice the tone. If it is harsh, rewrite one of them blamelessly: what happened, what in the system allowed it, what will change. See whether the rewritten version points to a better fix. It usually does.

Blameless, even alone Honest look clearly own your part Kind goodwill as to a peer Blameless Self-blame does one of two rush past: cause never learned dwell: drains, fixes nothing with agents, skip "it was wrong"; ask: brief, permissions, checks REWRITE THE NOTE Blame "I carelessly forgot to check the empty case." → points at a feeling Blameless "The empty case was not checked; the brief omitted it, no test." → points at what to change Look at the system that let it happen, not the person standing nearest.
Fig 88 · Blameless, Even Alone. Blamelessness sits where honesty meets kindness, and rewrites notes towards fixes.
Chapter 89 · Part IX

Guardrails From Scars

The best guardrails in any operation were not designed in advance. They were learned. Each one is the residue of something that went wrong: a scar that became a lesson that became a rule. An operator who systematically turns scars into guardrails ends up with an operation that is resistant to exactly the mistakes it is most prone to, which is a far better protection than any generic set of best practices.

The chain is straightforward. Something goes wrong and leaves a scar: a broken release, a lost afternoon, an embarrassed message to a client. The incident note extracts the lesson: why it happened, what would have prevented it. Then the lesson becomes a guardrail: something concrete in the system that makes the mistake harder or impossible next time. The guardrail is the payoff. Without it, the lesson is just something you hope to remember, and memory, as this book keeps saying, is not something to rely on.

Guardrails come in several forms, from soft to hard. A line in a memory file or recipe, telling agents to avoid something or always check something. A step in a checklist, such as the finish command or the release note. A test that fails if the mistake recurs. A permission that prevents the dangerous action. A check in the release pipeline that blocks the merge. A hook or automated rule that intervenes at the moment of risk. Match the hardness to the severity of the scar.

Every rule in a good operation has a story behind it. Make sure yours do.

Soft guardrails are cheap and easy to add, and they are the right choice for minor mistakes. A line in the memory saying always check that dates are in the user's time zone is enough for a mistake that caused a small confusion once. Hard guardrails are more effort and the right choice for serious or repeated mistakes. If a secret was committed once, a secret scanner that blocks pushes is a hard guardrail that makes it nearly impossible to happen again, and it is worth the setup.

Keep the story with the guardrail. When you add a rule to a memory file, add a brief note of why: added after the March export incident. When you add a check, reference the incident note. This matters for two reasons. It lets you, later, judge whether the guardrail is still needed; if the underlying cause has been fixed, the guardrail may be removable. And it stops the guardrail from looking arbitrary, to you or anyone else, which makes it more likely to be respected.

Guardrails can also accumulate into clutter, like everything else. A memory file with fifty rules, each from a different scar, can become hard for agents to follow and hard for you to maintain. Prune guardrails at the quarterly review just as you prune memory. Remove those whose causes are gone. Promote soft ones that keep being needed into harder ones that enforce themselves. Merge those that overlap.

This week, look at your last three incidents or significant mistakes. For each, check whether a guardrail was added. If not, add one, choosing the hardness to match the severity. Keep the story with it. Scar by scar, the operation takes the shape of its own experience.

From scar to guardrail Scar a broken release Lesson the incident note’s why Guardrail concrete, in the system GUARDRAILS, FROM SOFT TO HARD Memory a line Checklist a step Test catches it Permission blocks it Pipeline no merge Hook intervenes minor scar serious or repeated "dates in the user’s time zone" secret scanner blocks pushes Keep the story with it "added after the March export incident" Prune each quarter remove · promote soft to hard · merge Every rule in a good operation has a story behind it.
Fig 89 · Guardrails From Scars. Scars become lessons, then guardrails whose hardness matches the severity.
Chapter 90 · Part IX

The Runbook Habit

A runbook is a written procedure for doing something: deploying a project, restoring a backup, rotating a credential, onboarding a new client, setting up a fresh machine. Larger organisations keep runbooks because many people need to perform the same tasks consistently. A solo operator needs them for a different reason: the person who will perform the task next time is you, months from now, having forgotten every detail.

The habit is simple. If you have done something twice and expect to do it again, write it down. Not a polished document, just the steps in order, a check after each step to confirm it worked, and how to undo it if something goes wrong. That is the runbook. Keep it in plain text, with the project or beside your recipes.

The trigger, done twice, is important. Writing a runbook after doing something once tends to capture the specifics of that one occasion rather than the general procedure. Writing it after the second time captures what was common to both, which is the procedure. Writing it after the fifth time means you spent the third, fourth and fifth times reconstructing what you had forgotten. Twice is the sweet spot.

The person who will need the runbook is you, on a bad day, having forgotten everything.

The check after each step is what distinguishes a runbook from a list of instructions. Instructions tell you what to do. A runbook also tells you how to know it worked: after deploying, open the health check page and confirm it shows the new version. Without checks, a runbook can be followed perfectly and still produce a broken result, because one step silently failed and nothing told you.

The undo section is what makes a runbook safe to follow under pressure. If a step fails partway through, you need to know whether to continue, retry or go back, and if back, how. Writing this in advance, calmly, is far better than working it out in the moment. For many routine procedures the undo is trivial. For some it is the most important part of the document.

Runbooks and recipes are close cousins, and the line between them is blurry. A recipe is a brief for an agent; a runbook is a procedure, which might be followed by you or by an agent. Many operators find that their runbooks gradually become agent-executable: the steps are clear enough that an agent can follow them, with you reviewing the checks. That is a good direction, as long as the runbook is still readable by a human on the day the agent is not available.

Test runbooks occasionally. A runbook that has not been followed in six months may refer to steps that have changed, tools that have moved or settings that no longer exist. The upkeep calendar is a good place to schedule a periodic dry run of the most important ones: deploy, restore, rotate. Fix whatever is out of date.

This week, find one procedure you have done at least twice and write its runbook: steps, checks, undo. Then follow it once, exactly as written, and fix whatever you find. You will be surprised how many small details your memory had quietly filled in. The runbook now holds them, which means your memory no longer has to.

The runbook habit Done once specific to that day Done twice write it down now Third time on just follow it fifth time? rebuilt thrice Step Check after it Undo 1 Back up the data Backup file is not empty Nothing safe to stop 2 Deploy the release Health page shows new version Redeploy last version 3 Switch on the feature Use it as a user would Switch off instantly plain text, beside the project or the recipes Dry run from the upkeep calendar Fix drift steps, tools, settings Agent can run it a human still can too The person who needs the runbook is you, on a bad day, having forgotten.
Fig 90 · The Runbook Habit. Write the runbook on the second time: steps, a check after each, and an undo.
Part X

Deciding, Not Doing

The operator's real job, stated plainly.

Chapter 91 · Part X

The Decision Queue

Every operator has a backlog of tasks. Fewer realise that they also have a backlog of decisions, and that the second backlog is the one that actually limits the operation. Tasks can be delegated; agents will happily work through a long list of them. Decisions cannot. Every thread waiting on your answer, every project waiting for you to choose a direction, every review waiting for your verdict is a decision in a queue, and the length of that queue is the true measure of how much the operation is waiting on you.

Making the decision queue visible is the first step. Most operators carry it in their heads, which means they cannot see how long it is and cannot prioritise within it. Write it down. One list, with a line for each pending decision: what needs deciding, what it is blocking, and when it needs deciding by. Include the small ones as well as the large, because small decisions left unmade block threads just as effectively as large ones.

Not everything that looks like a decision is one. Agents ask a great many questions, and many of them can be answered by the brief, the memory or a standing policy. Each time you find yourself answering the same question twice, that is a sign it belongs in memory rather than in the queue. Each time you find yourself deciding something that a clear policy would settle, write the policy. The queue should narrow from everything agents ask, through the real decisions that need you, down to the few that need deciding today.

Your backlog of tasks is the agents' problem. Your backlog of decisions is yours.

Work the queue deliberately. In the morning review, look at it and pick the decisions that unblock the most. Decide those first, in your best hours. Decisions that block several threads are worth far more than decisions that block one. Decisions with deadlines come before those without. And decisions that are easy should be made immediately, because an easy decision left in the queue costs as much attention as a hard one, every time you look at it.

Many decisions are slow not because they are hard but because they are uncomfortable: saying no to someone, ending a project, admitting a direction was wrong. These tend to sink to the bottom of the queue and stay there. Notice them. A decision that has been in the queue for more than a week is almost always one of these, and the discomfort of making it is almost always less than the cost of leaving it. The following chapters have more to say about this.

This week, write your decision queue. Every pending decision, with what it blocks and when it is due. Then, each morning, decide the three that unblock the most. At the end of the week, compare the length of the queue with where it started. If it is shorter, the operation will have moved faster, because the thing it was mostly waiting for was you.

The decision queue narrows Everything agents ask "may I use X?" "which option?" "ok to proceed?" "approve this?" "what about Y?" Real decisions only you can make what · blocks · due Decide today the three that unblock the most Asked twice? → put it in memory Policy would settle? → write the policy Order the queue unblocks most first deadlines next easy? decide now a week old? it is the uncomfortable one One line each what needs deciding what it is blocking when it is due Your backlog of tasks is the agents’ problem. Your backlog of decisions is yours.
Fig 91 · The Decision Queue. Agent questions narrow to real decisions, then to the three that unblock the most.
Chapter 92 · Part X

One-Way and Two-Way Doors

Decisions differ in how easily they can be undone. Some are two-way doors: you can walk through, look around, and walk back if you do not like it. Trying a new layout, adjusting a recipe, choosing which project to work on this week, picking a tool for a trial. Others are one-way doors: once through, you cannot easily return. Deleting data, signing a contract, publishing something to a large audience, committing to a public deadline, ending a client relationship.

The distinction matters because the two kinds deserve very different treatment. Two-way doors should be decided fast. The cost of a wrong choice is small, because you can reverse it, and the cost of slow deciding is real, because the operation waits. Gathering more information or deliberating longer rarely improves a two-way decision enough to justify the delay. Make the call, watch what happens, adjust.

One-way doors deserve more care. Here, the cost of a wrong choice may be large and permanent, and some deliberation is worthwhile. Gather the evidence that matters. Sleep on it if you can. Ask an agent to argue the other side. Look for ways to make the decision more reversible, as the chapter on shipping and polish suggested: a trial period, a small release, a backup, a staged commitment.

Most decisions are two-way doors dressed up as one-way ones. Check the hinges.

Think of decisions on two axes: how reversible they are and how high the stakes. Reversible, low-stakes decisions are where you should decide fast and move on. That quadrant contains most of an operator's decisions, which means that most decisions should be fast. The common failure is to treat these as if they were in the irreversible, high-stakes corner, and to deliberate over choices that could simply be tried and revised.

The opposite failure also happens, particularly with agents. Because agents make action so easy, it is possible to walk through a one-way door at the speed of a two-way one. A deletion executed by an agent in seconds is as irreversible as one done by hand. An email sent by a recipe reaches everyone just the same. The permissions policy and the ask-first tier exist largely to slow down one-way doors, so that they get the deliberation they deserve.

A useful habit is to label decisions as you add them to the decision queue: one-way or two-way. The label tells you how much time to spend. Two-way decisions get a minute and a choice. One-way decisions get proper attention, ideally in your best hours, with evidence and a night's sleep where possible. Over time, you will notice that the two-way label applies far more often than your instincts suggested.

This week, label every decision in your queue. Make all the two-way ones today, quickly, without agonising. Give the one-way ones the time they need. Notice how much shorter the queue becomes, and how few of the fast decisions you later regret. The ones you do regret, you can reverse. That was the point.

One-way and two-way doors ONE-WAY DOOR TWO-WAY DOOR high stakes low stakes Deliberate gather key evidence sleep on it argue the other side best hours Try, then adjust small release first watch closely reverse if wrong Check the hinges often really two-way make it reversible: trial, backup, stage Decide fast most decisions a minute, a choice watch, adjust Two-way new layout tweak a recipe this week’s focus a tool trial One-way delete data sign a contract publish widely end a client agents walk one-way doors at two-way speed → the ask-first tier slows them Most decisions are two-way doors dressed up as one-way ones.
Fig 92 · One-Way and Two-Way Doors. Decisions sorted by reversibility and stakes; reversible, low-stakes ones go fast.
Chapter 93 · Part X

Saying No to Good Ideas

Agents are generous with ideas. Ask one to review a project and it will suggest improvements. Ask one to research a market and it will identify opportunities. Ask one to fix a bug and it may well mention three other things worth doing while it is there. Most of these ideas are good. That is precisely the problem. An operator who says yes to every good idea will never finish anything, because there will always be another good idea arriving faster than the previous one can be completed.

The skill, then, is not distinguishing good ideas from bad ones. Bad ideas are easy to decline. The skill is saying no, or not yet, to good ideas, because they do not fit the current priorities, because there is no attention for them this week, or because finishing what is already started is worth more. This is harder than it sounds, because each good idea arrives with its own small glow of possibility, and declining it feels like a loss.

It helps to remember what saying yes actually costs. Every new idea that becomes a project takes a slot on the register, a share of your attention, threads to brief and review, and eventually decisions about its future. The cost is not the agent's time, which is cheap. It is yours, which is not. A yes to a new idea is a quiet no to something else, usually the thing you were supposed to be finishing.

Agents generate options. The operator's job is to prune them.

Saying no does not mean losing the idea. The queue exists precisely for this. Write the idea down in one line, note where it came from, and let it compete at the next weekly or quarterly review with everything else. Many good ideas, given a few weeks in the queue, turn out to be less compelling than they seemed. A few turn out to be even better, and they get started at a time of your choosing, with proper attention, rather than squeezed into a week that was already full.

Some operators find it helpful to keep a simple rule: no new project starts until an existing one finishes or is parked. One in, one out. It forces the comparison that saying no requires: is this new idea better than the weakest thing I am currently running? If yes, swap them. If no, the new idea waits. The rule turns a vague discomfort into a clear decision.

It is also worth telling agents about your appetite for ideas. A line in the memory such as note other improvements at the end of your report; do not implement them keeps the ideas flowing without letting them leak into the work. You get the benefit of the agent's suggestions without the cost of uncontrolled scope.

This week, keep count of the good ideas that come your way, from agents, from reading, from your own head. Say yes to at most one. Queue the rest. At the end of the week, read the queue. Notice which ideas still look exciting and which have faded. The faded ones are the attention you saved by saying no.

Saying no to good ideas Good idea agents or reading Beats the weakest? one in, one out yes Swap them park the weakest no Queue it one line + source Weekly review ideas compete Faded attention you saved Still exciting start it, on your timing A yes costs a register slot your attention threads to brief reviews to do future decisions Memory line: "note improvements at the end of your report; do not implement them" Agents generate options. The operator’s job is to prune them.
Fig 93 · Saying No to Good Ideas. A new idea must beat the weakest thing running, or it waits in the queue.
Chapter 94 · Part X

Deciding on Partial Evidence

Earlier chapters insisted on evidence: evidence over reassurance, evidence in every review, evidence before every ship. That remains true. But there is a trap on the other side, and conscientious operators fall into it often: waiting for complete evidence before deciding. Complete evidence rarely arrives. Most decisions have to be made with what you have, and the skill is knowing when what you have is enough.

Enough is where two things meet: the evidence you have and the time you have. More evidence would always be nice. More time would always be nice. But at some point, the value of more evidence is outweighed by the cost of waiting for it, and that point is where you should decide. It comes sooner for two-way doors than for one-way ones, sooner for low-stakes decisions than high-stakes ones, and sooner than most careful people instinctively feel.

Agents make this trap easier to fall into, because they make gathering evidence so cheap. You can always ask for another analysis, another comparison, another round of research. Each one is quick and each one feels responsible. But each one also delays the decision, and the operation waits. At some point, asking for more research becomes a way of avoiding the discomfort of committing, and the research is no longer serving the decision. It is replacing it.

The question is not whether you know enough to be certain. It is whether you know enough to act.

A useful test is to ask what evidence would change your mind. If you can name it, and it is obtainable in reasonable time, get it. If you cannot name anything that would change your mind, you have already decided and are merely postponing the announcement. If the evidence that would change your mind is unobtainable, or would take longer to get than the decision can wait, decide now with what you have.

Another test is to imagine the decision going wrong and ask whether more evidence would have prevented it. Sometimes yes: a quick check of the data would have shown the problem. Often no: the outcome depended on things that were unknowable in advance. In the second case, waiting would not have helped, and deciding promptly at least gave you more time to notice and adjust.

When you do decide on partial evidence, say so, in the decision log. Decided to proceed with the smaller version; evidence on demand is thin, but waiting another month would cost more. Will revisit after the first two weeks of use. This records the uncertainty honestly and builds in a moment to check whether the decision held up. It also protects you, later, from the false memory that you were more certain than you were.

This week, look at the oldest decision in your queue. Ask what evidence would change your mind. If you can get it today, get it. If not, decide now, write down the uncertainty and set a date to revisit. Notice that the world did not end. It rarely does, and the operation moves again.

When is partial evidence enough? value of more evidence cost of waiting Enough: decide time → comes sooner for two-way doors What would change your mind? nameable, quick get it nothing would announce it too slow to get decide now Log it: "evidence is thin; waiting costs more; revisit in two weeks" Not whether you know enough to be certain. Whether you know enough to act.
Fig 94 · Deciding on Partial Evidence. Decide where the value of more evidence falls below the rising cost of waiting.
Chapter 95 · Part X

The Price of Not Deciding

Not deciding feels like keeping options open. It is not. Not deciding is itself a decision, usually a poor one, made by default rather than by choice. When you leave a decision unmade, the world does not wait. Things drift, circumstances change, and eventually something happens that settles the matter for you, rarely in the way you would have chosen.

The pattern is reliable. You defer a decision, because it is uncomfortable, or you want more information, or it does not seem urgent. While it is deferred, the operation drifts around it. Threads that depend on it stall or proceed on assumptions. Other decisions get made that assume one answer or the other. Time passes and options quietly close. Eventually a default wins: the project dies of neglect, the client makes the choice for you, the deadline arrives and forces whatever is nearest to hand. The decision was made. You just were not the one who made it.

The costs of not deciding are mostly invisible, which is why they are so easy to incur. Nobody sends you a bill for the week a project stalled waiting on you. Nobody points out that three threads proceeded on a wrong assumption because the decision they needed was not there. But these costs are real, and in a one-person operation they fall entirely on you.

If you do not decide, something else will. It will not have your interests at heart.

There is also an emotional price. Undecided matters sit in the mind and generate low-level anxiety. They resurface at odd moments, in the shower, at three in the morning, in the middle of unrelated work. They make the decision queue feel heavier than it is. Many operators find that making a long-deferred decision, even an imperfect one, brings an immediate and disproportionate sense of relief. The relief is the cost they had been paying without noticing.

The remedy is not to decide everything instantly; some decisions genuinely benefit from time. It is to make deferral a decision too. When you choose not to decide something now, write down when you will decide it and what you are waiting for. Deciding on the pricing change on the fifteenth, after the first week of usage data. That is a deliberate deferral, and it is fine. What is not fine is the open-ended I'll think about it, which is how decisions drift into defaults.

The weekly review is a good place to catch drifting decisions. Look at the decision queue for anything that has been there more than a week without a date. Each one is either a deliberate deferral missing its date, or a drifting decision. Give the first kind a date. Make the second kind now.

This week, find the decision you have been avoiding longest. Make it, today, with whatever evidence you have. Write down what you decided and why. Then notice the relief, and the threads that start moving again. That is the price of not deciding, refunded.

The price of not deciding DRIFTING Defer open-ended Drift threads stall Options close time passes Default wins not your choice work proceeds on assumptions neglect, the client or the deadline picks DELIBERATE DEFERRAL Defer with a date and what you wait for "the 15th, after a week of data" Decide on the date your choice, on time Hidden costs, all yours a stalled week · wrong assumptions · 3 a.m. anxiety If you do not decide, something else will.
Fig 95 · The Price of Not Deciding. Open-ended deferral drifts into a default; a dated deferral stays a decision.
Chapter 96 · Part X

Judgement Does Not Delegate

As agents become more capable, the boundary of what they can do keeps moving. Tasks that needed a person last year are routine for an agent this year. It is natural to wonder whether the boundary will eventually reach everything, and whether the operator's role will shrink to nothing. It will not, and the reason is worth understanding clearly, because it tells you where to invest your own development.

Think of the work as a stack. At the bottom is how it gets made: the writing, the coding, the designing, the researching. This layer is increasingly delegable, and agents handle more of it every month. Above it is whether it ships: the review, the quality bar, the release decision. Agents can inform this layer enormously, running checks, comparing options, flagging risks, but the decision has your name on it. Above that is what good means: the standards, the taste, the definition of done for this particular work, for these particular people. And at the top is why it matters at all: which problems are worth solving, which projects are worth running, what the operation is for.

The higher layers resist delegation not because agents are incapable of opinions about them. Agents can produce perfectly sensible opinions about almost anything. They resist delegation because they depend on things only you have: your relationships, your values, your knowledge of your own circumstances, your willingness to answer for the result. An agent can propose why a project matters. It cannot care whether it does, and it cannot be held responsible if it was wrong.

Agents can tell you what is possible. Only you can say what is worth it.

This has practical consequences for how you spend your own learning time. The bottom layer is where agents are improving fastest, and investing heavily in your own skills there yields diminishing returns. The upper layers are where your judgement is irreplaceable, and they reward investment: understanding your users better, developing sharper taste, getting clearer about what you are trying to achieve, learning to decide faster and better. These are the skills that make you more valuable as agents improve, not less.

It also has consequences for how you use agents. Delegate the bottom of the stack freely. Use agents heavily to inform the middle, asking for options, evidence and critiques. Use them sparingly and carefully at the top, as sounding boards rather than deciders. An operator who asks an agent what projects to run, and then simply runs them, has not delegated judgement. They have abandoned it.

None of this is a counsel of suspicion. Agents are extraordinary collaborators at every layer. The point is only that collaboration and delegation are different things. You can collaborate on judgement. You cannot hand it over, because the moment you do, you are no longer the operator. You are a passenger.

This week, look at the decisions you made and ask, for each, which layer of the stack it was in. Notice where you leaned on agents and where you did not. If you find you are delegating the top layers, take them back. If you find you are still doing the bottom layer yourself, let it go. The stack sorts itself once you can see it.

Judgement does not delegate Why it matters what it is for What good means standards, taste, done Whether it ships review, quality bar How it gets made writing, coding, research THE AGENT’S PART sounding board only examples and critique inform heavily delegate freely invest your learning ↑ Agents can tell you what is possible. Only you can say what is worth it. collaborate on judgement; hand it over and you are a passenger
Fig 96 · Judgement Does Not Delegate. Four layers of work, from how it is made to why it matters, and the agent’s part.
Chapter 97 · Part X

Teach What You Run

There is an old observation that you do not really understand something until you can teach it. For an operator, this has a practical edge. Writing down how your operation works, clearly enough that someone else could follow it, is one of the best ways to discover what you actually do, what you merely think you do and where the gaps are. Teaching what you run is a way of running it better.

The obvious form is documentation for a collaborator. If you ever bring someone into your operation, even briefly, you will need to explain how it works: the rhythm, the register, the recipes, the review standards, the release process. Writing that explanation forces a kind of clarity that daily practice does not. You discover habits you had never articulated, rules you were applying inconsistently and steps that only made sense because you knew the history.

But you do not need a collaborator to benefit. Writing for an imagined newcomer works almost as well. So does writing for agents, which is, in a sense, what memory files and recipes already are: teaching documents for a very fast, very literal student who forgets everything overnight. An operator who writes good memory files is already teaching, and can extend the habit to the whole operation.

Explaining your system to someone else is the fastest way to find out what it is.

Teaching takes several forms, gathered around the same centre. Writing it: putting the operation into words, in a handbook of your own. Explaining it: describing it aloud to someone, which reveals different gaps from writing. Sharing it: publishing parts of it, which invites questions you would never have asked yourself. And refining it: using what you learned from writing, explaining and sharing to improve the operation itself. Each feeds the others.

Sharing in particular is underrated by solo operators, who tend to think their methods are too idiosyncratic to interest anyone. Usually they are wrong. Other people running similar operations face similar problems and are often glad to see how someone else solved them. And the act of preparing something for others to read raises your own standard: you will not publish a description of your release process without first making sure it is a release process you would be proud of.

There is a further benefit for a solo operator. A written account of how your operation works is insurance. If you are ill, on holiday or simply away for a while, the account lets you, or anyone helping you, pick up the threads. It is the operation's own runbook, at the highest level, and like any runbook it is only useful if it was written before it was needed.

This week, write one page explaining how your operating day works, for an imagined newcomer: the review, the sessions, the close, the register, the way you brief and review. Be specific. Then read it back and mark everything that surprised you, contradicted what you actually do or felt hard to explain. Those marks are your next improvements. You set out to teach, and you ended up learning, which is how it usually goes.

Teach what you run Your operation Write it habits never said Explain it gaps writing hides Share it raises the bar Refine it the operation TEACH IT TO a collaborator a newcomer your agents memory files already teach insurance, if you are away Explaining your system to someone else is the fastest way to find out what it is.
Fig 97 · Teach What You Run. Writing, explaining and sharing your operation feed back into refining it.
Chapter 98 · Part X

The Operator Ahead

Where is this role going? Any honest answer begins with uncertainty. The tools are changing quickly, the capabilities of agents are expanding in ways that are hard to forecast, and anyone who claims to know exactly what a solo operator's day will look like in a few years is guessing with confidence. But the direction is visible, and the direction suggests that the core of this book will matter more, not less.

The trend so far is towards longer leashes. Agents work for longer without check-ins, take on larger pieces of work, coordinate with each other, run in the background and on schedules, and handle more of the routine judgement that used to require a person. Each step moves the operator further from the doing and further towards the edges of the work: setting the intent at the beginning and exercising judgement at the end.

That shape is the protocol at the heart of the operator's job, and it is already visible today. The operator states intent: what is wanted and why, with a definition of done. The agents return work and proof: the result and the evidence that it meets the intent. And the operator exercises judgement: whether this is good, whether it ships, what comes next. As agents improve, the middle of that exchange gets longer and more capable. The beginning and the end stay with the operator, and become a larger share of what the operator does.

As the agents do more, the operator does less, and what the operator does matters more.

What does this mean for the skills worth building? Clear intent: the ability to say precisely what you want, which is the brief, the definition of done, the constraints. Sound judgement: the ability to evaluate work and decide what to do with it, which is the review, the taste, the release decision. And good rhythm: the ability to structure your own time and attention so that intent and judgement are exercised at their best, which is the operating day, the weekly review, the energy budget. Every part of this book is about one of those three.

It also means the operator's job becomes, in some ways, more human. When the doing is delegated, what remains is the part that depends on being a particular person with particular relationships, values and responsibilities. Understanding what a client actually needs. Deciding what is worth building. Taking responsibility for the result. These are not technical skills, and they do not become obsolete when the technology improves. They become the job.

There will be new tools, new practices and new names for things, and some of the specifics in this book will date. Treat the specifics as examples and the principles as the substance. The principles, clear briefs, evidence over reassurance, small releases, honest records, sustainable rhythm, ownership of outcomes, are older than agents and will outlast any particular version of them.

This week, ask yourself which of the three, intent, judgement or rhythm, is your weakest. Pick one practice from this book that strengthens it and adopt it for a month. The future of the role will reward the operator who got better at the edges while the middle was being automated. Start now. The middle is already moving.

The shape of the job ahead Operator Agents intent: what, why, done longer runs larger pieces coordinating on schedules ↓ the middle grows work and proof Judgement the edges stay yours SKILLS WORTH BUILDING Clear intent brief · done · limits Sound judgement review · taste · ship Good rhythm day · week · energy As the agents do more, what the operator does matters more.
Fig 98 · The Operator Ahead. Intent goes out, work and proof come back; the middle grows, the edges stay yours.
Chapter 99 · Part X

The Handbook as Habit

A handbook is not meant to be read once. It is meant to be used: opened when a particular problem comes up, consulted when a habit has slipped, reread when something is not working and you cannot quite tell why. This one is no exception. Its hundred chapters are not a course to complete but a set of practices to adopt, a few at a time, and to return to as your operation changes.

The cycle that turns a handbook into a habit has three steps. Read a chapter, or a part, when it is relevant. Practise what it suggests for a week or two, deliberately, in your actual work. Then revise: decide whether the practice works for you as written, needs adapting, or should be dropped. Then read again when the next problem arrives. The practising is the step that matters. Reading without practising is entertainment. Practising without revising is rigidity. The cycle needs all three.

Do not try to adopt everything at once. An operator who starts the morning review, the close, the register, the decision log, the release notes, the incident notes and the weekly review all in the same week will keep none of them. Pick one or two practices, the ones that address your biggest current problem, and give them a month. When they have become automatic, add another. Habits compound, but only if they survive long enough to become habits.

Read it once for the ideas. Then use it for the practice.

Adapt freely. This book describes practices that work for many operators, but your operation is your own. Perhaps your review works better at lunchtime. Perhaps your register works better as a table. Perhaps today's three should be today's two. The principles matter more than the specifics: decide on purpose, brief clearly, review with evidence, ship small, write things down, rest. How you implement them is yours to work out, and the revise step is where you do it.

Keep your own handbook alongside this one. As you adapt practices, write down your versions: your morning review checklist, your brief template, your release note format, your upkeep calendar. Over time, your own handbook will become more useful to you than this one, because it will fit your operation exactly. That is the goal. This book is a starting point, and a good starting point is one you eventually outgrow.

Return to it at the quarterly review. Skim the parts. Ask which practices you have adopted, which you have let slip, and which you have never tried. Choose one or two to work on next quarter. This keeps the cycle running at the longest timescale, and it means that every quarter your operation is a little more deliberate than the one before.

This week, pick the one chapter from this book that addresses your biggest current problem. Practise it for two weeks. Then revise: keep, adapt or drop. Then pick the next. In a year, you will have built an operating system that fits you. You will also, very likely, have stopped needing to read about it, which is the best thing a handbook can hope for.

Turning the handbook into habit Read when relevant Practise 1–2 weeks, real Revise keep, adapt, drop one or two practices at a time next problem Read, never practise = entertainment Practise, never revise = rigidity Your own handbook checklists, templates, your versions Read it once for the ideas. Then use it for the practice.
Fig 99 · The Handbook as Habit. Read, practise and revise in a loop, one or two practices at a time.
Chapter 100 · Part X

The Job Is Deciding

Here is the thesis of this book, stated plainly: the operator's real job is deciding, not doing. Agents do the doing, more of it every month, faster and more capably than any one person could. What they cannot do is decide what is worth doing, what good looks like, whether the result is good enough to carry your name, and what should happen next. Those decisions are the job. Everything else in this book is machinery for making them well.

Delegate the doing, because the doing is now delegable. Drafting, building, researching, testing, summarising, tidying: hand them over with a clear brief, sensible permissions and a definition of done. Holding on to them out of habit or pride is not diligence. It is spending your scarcest resource, attention, on the most plentiful one, effort.

Review the work, because review is where your judgement enters it. Read the diff, check the evidence, apply the cheapest checks first and the deepest where the risk demands. Ask whether the shape is right before polishing the detail. Ask whether it is not merely correct but yours. Review is not a lesser activity than making. It is the activity that makes the made thing trustworthy.

And own the decision, because ownership is what makes the whole arrangement work. Ship it or send it back. Start the project or say no. Park it or end it. Decide on the evidence you have, quickly where the door swings both ways and carefully where it does not. Write down why. Then live with the result, learn from it if it goes wrong, and decide again. That is what an operator is.

The agents supply the effort. You supply the decisions. That is the whole job.

The rest of this book is in service of that. The operating day exists to give deciding its best hours. Briefs exist to carry decisions into the work. Memory exists so that decisions do not have to be made twice. The register and the decision queue exist so that you can see what needs deciding. Release discipline exists so that each decision to ship is small and reversible. The weekly and quarterly reviews exist so that you decide the direction rather than drift into it. Rest exists so that the decider is in a fit state to decide. Even the incident notes are about deciding: deciding what to change so the mistake does not repeat.

This is good news, although it may not feel like it at first. Deciding is harder than doing in some ways, less tangible and less immediately satisfying. But it is also more interesting, more human and more durable. The doing will keep getting automated. The deciding will keep needing someone. If you get good at it, you will be more valuable each year, not less.

So tomorrow morning, before anything else, open a blank page and write down the three decisions that would most move your operation forward. Then make them. Not the tasks, the decisions. Let the agents handle the rest. You will find, as many operators have, that the work gets easier and the job gets more interesting. That is not a contradiction. It is the job, finally seen clearly. It always was deciding. The agents simply took away everything that was hiding it.

The job is deciding Delegate the doing brief, permissions, done Review the work diff, evidence, taste Own the decision ship, park, end, say no Deciding the whole job Operating day its best hours Briefs carry it into work Memory never decide twice Register, queue see what waits Release discipline small, reversible Weekly, quarterly direction, not drift Incident notes what to change Rest a fit decider tomorrow: write the three decisions that would move things most, then make them The agents supply the effort. You supply the decisions.
Fig 100 · The Job Is Deciding. Delegate, review, own: every part of the operation exists to serve deciding.
The Operator’s Handbook · First Edition, October 2026
100 chapters · 10 parts · one hundred diagrams
by Mat Siems · MS Books, No. 15 · 2026