Plan Expensive, Execute Cheap

One very capable, very expensive thinker was doing two jobs that are not the same job. So the expensive model plans, twice, and the cheap model builds.

On 13 July I watched a session eat my entire five hour usage allowance in fifteen minutes. What I wrote at the time was not "the model is bad." It was:

"it burns up my entire 5 hour usage in 15 minutes, so our method is wrong"

The method. That distinction matters, because the fix was not a better model or a bigger allowance. It was noticing that I had one very capable, very expensive thinker doing two jobs that are not the same job.

It was thinking about the problem, and then it was also grinding through the implementation of every ticket the thinking produced. That is paying an architect by the hour to hang your doors. They will do it, and they will do it well, and it is not what you are paying them for.

So the work got split by the kind of thinking it needs.

The expensive model plans. Twice.

The first pass writes the user stories into Linear. Epics, sub-tickets, one shippable story per card. This pass is about scope and sequence and nothing else.

The second pass goes back over every card and adds the coding notes. How this one should be approached, what it touches, what was decided and why, what to leave alone. Discussion in the comments where there is a real choice to make.

Measure twice, cut once. I learned that building houses, where I was the cutter, which is the end of the job at which being wrong is expensive. Plan twice, execute once is the same rule moved indoors.

Two passes rather than one, because scope and approach are different questions and asking them together produces cards that are either well bounded and vague or detailed and wrongly bounded. Separating them costs one extra pass and saves the day where you find out at the end.

Then the cheap model builds.

By the time a card is picked up there is very little left to decide. That is the entire point. The expensive thinking is already sitting in the ticket. What remains is closer to transcription than invention, and transcription does not need the best reasoner available. Not because the cheaper model is worse, but because you are no longer asking it the hard question.

The saving is real and it is not marginal. But the reason I keep doing it is not the token bill.

It is that this is the thing that separates a method from vibe coding. Vibe coding goes after whatever comes to mind next, and it feels productive, and it produces a house where somebody put the roofing on before the frame was up. Not because there is a rule against roofing first. Because there is nothing to nail it to.

I built houses for a living once. You do not discover the frame while you are working. You mark it out, and then you cut, because the offcut does not go back on.

The question worth asking about your own setup is not which model is best. It is which parts of your work are actually thinking, and whether you are paying thinking rates for the rest of it.

Next: the two hours I lost to half a shell command.