Blog

Fewer, bigger jumps

When you work with coding agents, tokens are cheap and your attention is not. Ask for a leap that lands near the goal, close the gap in shrinking corrections, and spend the time it buys on the next leap.

the codecast team7 min read

Most advice about prompting coding agents is about getting each answer right. Be specific, give context, ask for one thing at a time. That advice optimizes the wrong number. It treats the model's output as the scarce thing, when the scarce thing is you: the minutes you spend reading what came back, deciding what is wrong with it, and writing the next prompt.

Call each of those cycles a jump. Every feature is a distance from where the code is to where you want it, and every jump costs a slice of your attention whether it moves you an inch or a mile. Tokens are cheap and getting cheaper every quarter. Your attention is fixed at a few good hours a day. So the thing to minimize is the number of jumps, and the way to do that is to make the first one enormous.

Small hops
Each prompt asks for one safe step. You read, judge and re-prompt every time.
what you meantstartyour turns16
One leap, then corrections
One prompt asks for the whole thing. The agent iterates on its own and lands near; you steer the rest.
what you meantstartthe agent loops on its own:plan, build, test, look, fixlands close,wrong in ways you can seecorrectionsshrinkyour turns4
Same start, same target. The left takes sixteen of your turns; the right takes four.

Two ways across

The familiar way is small hops. You ask for one safe, well specified step, check it, ask for the next. Each hop lands roughly where you aimed, which feels like control. But a feature of any size is fifteen or twenty of them, and every one waits on you. Worse, you are doing the planning, the sequencing and most of the verification in your head, which is exactly the work the agent could have done.

The other way is one leap. You describe the destination, what done looks like and how to check it, and then you tell the agent to keep going on its own: plan, build, run it, look at it, find what is wrong, fix it, and go around again. In the figure, those coils inside the first orange arc are that loop. They cost tokens and wall clock time. They cost you nothing.

The leap will not land on the target. It will land somewhere near it, built on a few guesses you would have made differently. That is fine, and it is the part people underrate.

Why missing is useful

A working artifact that is wrong in specific ways is the cheapest thing in the world to steer. Before the leap you would have had to specify everything up front, including dozens of choices you had no opinion on until you saw one. After it, you only have to name the handful you disagree with, and you can name them by pointing.

Here is what that looks like in practice, from a session in our own repository this week. One prompt asked for a page that tells the story of everything the team is shipping. The agent worked through more than 750 messages on that prompt alone, building a per-day timeline, wiring it to real data, screenshotting it and fixing what it saw. Then came the second real prompt: the correction.

Changes page 1,187 messagesmsg 1 · the leap“a high level natural language beautifullyrendered timeline of everything that is changing…”msg 760 · the correction“a single timeline - not one page per day…just try to pare it down to the core”Docs audit 300 messages“do an audit and update our docs…, add images,screenshots, and vector drawings” then “remove the stray/leftovers”a prompt with direction in ita bare “continue” after a pauseagent messages
Two real sessions in the codecast repository, read on October 4. Orange is every prompt that carried direction.

The correction is a big jump of its own: a single timeline instead of one page per day, plainer copy, fewer features on the surface. It moves a long way, but it moves from a place that was already close, so it lands much closer still. The jumps after it get smaller: a spacing fix, a label, a tweak to the default zoom. Each one is cheaper to write than the one before, because the remaining error is smaller and easier to see. The docs audit below it is the same shape at a smaller size: one prompt to rewrite the docs with diagrams and screenshots, about 270 messages of work, and a three word correction.

What a leap prompt asks for

A leap prompt is not a longer hop prompt. It asks for different things. It names where you are going rather than the next step. It says how the agent will know it has arrived, in terms it can check itself: tests, a browser, a screenshot. It explicitly invites the agent to spend more: plan first, fan work out to subagents, run a workflow that builds and reviews in parallel, and keep iterating on polish after the thing works. And it says when to stop and ask, which should be rarely.

a hop

Add a date column to the changes table.

Fine on its own. The trouble is the next nineteen like it, each waiting on you.

a leap
the destination

Build a changes page: one timeline of everything the team is shipping, drawn from sessions and commits, that a person can understand at a glance and zoom into.

how to know it's done

Done means it runs on real data, you have opened it in the browser and screenshotted every zoom level, the grouping logic has tests, and you have reread your own diff once.

spend compute

Plan first. Fan the independent parts out to subagents. Once it works, keep iterating on the copy and the polish until you would show it to a customer.

the only interrupt

Work on your own until then. Stop only for a choice that changes what the product does.

The hop is a fine prompt. It is just one of twenty.

The most important line is the definition of done. Without one, an agent stops at the first plausible result, and you become the test suite. With one, the agent does the looking for you: it opens the page, notices the overlapping labels, and fixes them before you ever see them. Every problem it catches on its own is a jump you did not have to make.

This deliberately spends more compute per prompt, often a lot more. A leap can run for an hour and burn through more tokens than a day of hops. That is the trade, and it is a good one. You are buying back the only resource in the loop that does not scale.

Fill the time the leap buys

A leap that runs for forty minutes is forty minutes in which you are not needed. The mistake is to spend them watching. The point of making each session need you less is that you can run several, and spend your attention moving between them: write the leap for one, launch it, write the leap for the next, then come back to the first when it lands and send its correction.

0 min20 min40 min60 min80 min100 minSmall hops, one sessionagentyou40% busythe rest waitingBig leaps, five sessionssession Asession Bsession Csession Dsession Eyou83% busyfive features
Schematic, not measured. Hops keep you waiting on one agent; leaps keep five agents waiting on you, briefly.

Throughput here means the share of your attention spent on decisions rather than waiting. With small hops it is low no matter how fast the model is, because every hop ends in a wait. With leaps it can approach all of it, and the corrections shrink as each session converges, so the later rounds go by quickly. Five sessions moving in big jumps finish more in an afternoon than one session moving in small ones finishes in a week.

Running like this needs a place to see who is waiting on you. That is what codecast's inbox is for: every session on every machine, sorted by who acts next, so the next jump you make is always the one that is ready for it. Long leaps can be given to cast spawn --subagent workers, and a session can be woken later with a trigger instead of a reminder in your head.

The rule

Before you send a prompt, ask whether it could carry the next three prompts too. Usually it can: say where you are going, say how to check, say keep going. Expect the result to be wrong, and be glad it is wrong somewhere you can point at. Then make the big correction, then the smaller ones, and spend the gaps on the next leap.

Count your jumps, not your tokens.

The session traces are real: message counts and prompt positions from two codecast sessions in this repository, read on 2026-10-04, with the quoted prompts trimmed where marked. Bare "continue" messages after a pause are drawn separately because they carry no direction. The jump paths and the attention schedule are illustrations, not measurements.