Work Bigger Than One Conversation
This Part
In Part 1 I said that every conversation has a ceiling, and I promised that we would get to it at the end, and now is the time to talk about that.
Everything we have said so far was about work that finishes inside a single conversation: summarize one paper, write one text for a patient, check one claim against its source. But a lot of real work is not shaped like that. Say you want to write ten educational texts for your practice’s website, or prepare an hour-long presentation, or review thirty papers. That work takes several days, and several days means several separate conversations.
First, what actually happens at the ceiling
You start the conversation and it goes well. You hand over papers, you make decisions, you revise. And then somewhere, with no signal at all, the quality drops.
These are the signs: the model forgets a rule you set at the start. It contradicts a decision the two of you made an hour ago, as if nothing had happened. It rewrites something it had already written. Or worse, it writes one paper’s finding under another paper’s name.
And none of it arrives with a warning. The tone is the same confident tone as the first hour. You are the one who has to notice that this conversation is no longer the conversation you started.
The practical rule is simple: one conversation, one job. When the job changes, change the conversation too. And if a job is so large that it does not fit in one conversation at all, break it into pieces.
How to break it up
Each piece has to be a finished thing in itself, not half a job. That means it has a beginning, a body and an ending, and that something comes out of it that is usable on its own. A ten-part series means ten conversations, each one a complete part you could publish that same day. A review of thirty papers means, say, five conversations, six papers each with its own summary, plus one final conversation that puts those five summaries side by side.
The order matters too. First, in one conversation, lay out the plan for the whole job and build one piece all the way through so you have a pattern to work from — the same thing we called a mockup in the previous part. Then build each following piece in its own conversation.
Lengthwise, not crosswise
Let me add one small thing in parentheses before explaining how to break the work up, because it is an important idea and method, and then I will go on with how the work is divided.
Suppose you have ten papers and you want each of them summarized in one conversation. There are two ways to say it. The first: “Read these ten papers and summarize each one.” The second: “Read the first paper, write its summary and save it, then move on to the second.” I always say the second.
The reason is that when ten papers sit in front of the model at once, it does not give all of them the same attention. The ones in the middle get seen less, and worse, the outputs stick to each other: the third paper’s finding gets written under the seventh paper’s name. The same mixing I mentioned above, except from the very beginning.
With the second method each paper gets the model’s full attention, not a tenth of it. After each paper a finished output is saved that no longer depends on the model’s memory; if things fall apart at the eighth paper, the previous seven are intact. And if there is an error somewhere, you know which paper it belongs to.
You lose exactly one thing: comparison. When the model reads the fourth paper, it no longer has the first one in front of it. The answer to that is not to go back to the first method. The answer is one final pass — not over the papers themselves, but over the summaries: hand over the ten short summaries at once and ask it to compare them now. Ten summaries fit comfortably in one conversation; ten papers do not.
So whether you break a job across several conversations or hand the papers over one at a time inside a single conversation, the logic is the same: every piece separate, every piece with a saved output, and one final pass to put them side by side.
Another way of breaking the discussion up is into separate conversations, so that we do not hit the ceiling inside one of them.
But this very separation creates a new problem, and that problem is bigger than the ceiling problem.
The new problem: the next conversation knows nothing
When you start the third part in a fresh conversation, the model does not know what the first and second parts were. It does not know which tone you chose, how you write the terminology, which paper you set aside and why. It starts from zero.
The bad solution is to paste in the text of all the previous parts again. Which means you have filled half the ceiling from the very start and you reach the same problem sooner.
The solution I use myself is something else, and I took its name from a familiar job. When your shift ends and you hand a patient over to the next colleague, you do not recount the entire history; you tell them what they need in order to continue. In English this handing over is called a handoff. It is the same thing here, except the next colleague is the next conversation.
What a handoff document is
At the end of every conversation, before closing it, I ask the model to write a document for the next conversation. Not a summary of what was said, but the things the next person needs in order to continue. Usually these five:
The goal of the whole job in one paragraph. The decisions that were made, along with their reasons, because a decision without a reason gets reopened in the next conversation. The style rules and terminology that have been settled. The current state: what is done and what is left. And the things that explicitly must not be done again.
Then I read it myself. Do not skip this step. The document was written by the model, and as we saw in Part 2 and Part 3, it may have dropped a decision. But its more common error is this: it records something that was only its own suggestion as if it were your decision. If you do not catch it, in the next conversation it gets carried out as a rule of yours.
I ask the model to hand the document over in Markdown, the format we talked about in Part 2. A simple file you can put wherever you want.
Where the document lives
You have a few options here and you choose depending on the job.
The simplest: paste the document at the start of the next conversation and say continue from here. For a job that runs three or four conversations, that is enough.
If the job is longer: put the document in the project knowledge you built in Part 2, so every new conversation has it on its own. This is where the first two parts of the chapter connect to each other.
And if the job lives somewhere outside the chat — the files of your practice’s website, say, or the folder where all your educational texts are — put the document right there beside the work itself. That is what I do for my own site; the handoff document stays next to the site’s files, and every tool that works on those files reads it first.
Another use that lowers your cost
The handoff document is not only for passing work from one conversation to the next. It also works for passing work from one model to another.
The more expensive, slower models are better at thinking and deciding, and the cheaper, faster ones are better at carrying out work that has already been made clear. So I do the job in two stages: with the more expensive model I lay out the plan, make the decisions, and at the end ask it to write a precise handoff document. Then I give that document to the cheaper model and tell it to execute. The expensive model writes the structure of the ten educational texts and the rules for all of them, say, and the cheap model builds the ten one by one.
This works because the hard part of the job — the thinking and the deciding — has been done once and now sits in the document. What is left is following instructions, and you do not need the expensive model for that.
A real example: this book
This book was written exactly this way. Each chapter was several conversations, and at the end of each one a handoff document was written: what we decided, what we set aside for later chapters, how we write the terminology. The next chapter started from that document.
And the same error I mentioned above happened here too. Several times the document recorded something I had not said and the model had concluded on its own. Because I read the document, I caught them. If I had not read it, they would have turned into rules of the book.
The chapter in summary
It was five parts, and if you put them one after another, a way of working comes out:
Know which tool you are working with and why. Tie the model to your own sources. Run the output through the sieve so your judgement collects on a few specific points. Before building, open the problem. And break large work so that each piece fits in one conversation, and from each piece to the next, a document passes from hand to hand.
None of this makes the model more correct. All of it makes its errors less able to stay hidden, and lets you see them sooner. The same first principle of the book: a companion to thought, not a substitute for judgement.
What the next chapter is
So far the book has left one thing out, and left it out deliberately: the patient.
In every example in this chapter there was a paper, a file, an educational text. But when you put a patient’s record into the project knowledge, or show the model a radiograph, or send a case summary out for a consultation, that information leaves your practice and reaches the server of a company you do not know. And if one day a text you wrote with a model carries an error that harms a patient, the model does not answer for it. You do.
The final chapter is about exactly those two things: what you must not give the model, and who is responsible when the model’s output reaches a patient. And there we come back to something I pointed at several times in this chapter and passed over: imaginary guarantees. A model that corrects itself and leaves you feeling its work has been checked; two models that agree and leave you feeling it must therefore be right; and a number the model gives you about its own confidence. For instance, when you tell the model to answer only if it is 80 percent sure, and it says it is 92 percent sure. All three do the same thing: they create the feeling that someone other than you is responsible. There we will see that no one is.