Ralph Loop Claude Code

The Ralph loop, also called the Ralph Wiggum technique, is a pattern for running Claude Code autonomously: a bash while true loop that feeds the same prompt file to a fresh agent session on every iteration, against a plan file and a set of specs, and lets it grind until the plan is done. It was named and popularized by Geoffrey Huntley, who describes it plainly:

“In its purest form, Ralph is a Bash loop.”

The whole mechanism fits in one line:

while :; do cat PROMPT.md | claude-code ; done

Each pass reads a prompt file that points at @fix_plan.md (a prioritized todo list), @specs/* (the requirements), and an agent guide with build and test instructions. The agent picks the next item, works it, and exits. The loop starts over with a brand-new session.

Why the loop works

The reason this is worth taking seriously, not dismissing as a hack, is the one design choice it gets right by construction: every iteration starts with a clean context window. Nothing from iteration one is still sitting in context by iteration fifty. Huntley’s own reasoning for this is direct. He frames the whole technique around the ceiling you are working under, approximately 170k of context window source , and treats every token spent re-reading old work as a token not spent on the current item.

That single property, a fresh session per unit of work instead of one long accumulating conversation, is the most important cost and quality lever any agent harness has. Cost tracks accumulated context times the turns still to come, so a session that resets on every iteration never pays the tax of the previous forty iterations’ history. See Claude Code token usage for the mechanism: the client resends the full conversation on every request, so a loop that never lets that conversation grow past one task is already ahead of a single marathon session doing the same work. Shiploop is built on the identical insight, just applied to a managed queue instead of a bare loop: see the shiploop section below.

Where an unmanaged loop costs you

The loop itself has no opinion about any of the following. They are not flaws in the idea, they are things nobody wired into fifteen lines of bash, and they show up the moment you run one for more than an afternoon.

No gate on repeat failure. A loop that cannot complete a task attempts it again on the next pass, at full price, forever. Nothing in while :; do ... ; done counts consecutive failures on the same item or stops feeding it back in. Huntley’s own writeup names this class of problem directly, including agents that assume unimplemented code is already done, or ship placeholder implementations unless the prompt explicitly forbids it. The fix is a failure-streak counter: after N failed attempts on the same item, stop or escalate instead of re-running it at the same tier forever.

No dependency ordering. A flat todo list has no notion that item eight needs item three finished first. The loop discovers that the hard way: it burns an iteration attempting item eight, fails or produces something wrong, and only finds out why on a later pass, if at all.

No tier routing. Every iteration runs the same model at the same cost, whether the item is a one-line fix or a genuinely hard piece of design work. There’s no signal in the loop that distinguishes them. What that actually costs is measurable: across shiploop’s own fifty-day run, one maintainer’s 8 shiploop PROOF.md -repo workspace (a single case study, not a controlled experiment) resolved 281 shiploop PROOF.md tasks out of 379 shiploop PROOF.md filed, at a mean of $4.49 shiploop PROOF.md and median $3.03 shiploop PROOF.md per resolved task (n=32, measured from total_cost_usd, skewed toward harder opus-tier tasks so closer to an upper bound than a floor average). Running every one of those at a single fixed tier, instead of routing the easy ones cheap, is the difference a loop with no tier logic leaves on the table.

No isolation. Every iteration works in the same checkout. Two overlapping items touching the same files collide, and nothing in the loop notices until the build breaks or a diff gets overwritten. See Claude Code worktrees for the isolation model a loop like this is missing: one working tree per unit of work, so two things in flight can never step on each other’s uncommitted state.

Nothing learned between runs. Huntley’s own limitation list includes broken codebases that need a manual git reset, and non-deterministic search behavior that trips agents up repeatedly. A bare loop hits the same gotcha on iteration sixty that it hit on iteration six, because nothing promotes what iteration six learned into what iteration sixty reads. The fix is mechanical: write the lesson into a file the next session actually loads, once, instead of re-discovering it forever.

Plain loop vs. governed backlog

PropertyPlain Ralph loopGoverned backlog
Context growth per unit of workFlat (fresh session each time)Flat (fresh session each time)
Repeat failureRetries forever at full costFailure-streak counter escalates or stops
Work orderingFlat list, no dependenciesDependency gate before dispatch
Model routingOne tier for everythingCheap floor, escalates once on evidence
IsolationShared checkoutOne worktree per item
Learning carryoverRediscovered every runPromoted into a file the next run reads
Setup and installNone: about fifteen lines of bashA queue, a dispatcher, and workspace scaffolding

That last row is real, not a concession. A while true loop needs no install, no config file, and no daemon. It is the honest baseline every more elaborate harness is competing against, and for a lot of work it’s the right amount of machinery.

Run the plain loop if

Run the bash loop as-is when it’s one repo, one task you can specify tightly in a plan file, and you’re going to be watching the run rather than walking away from it. That’s the case the loop was built for and it’s a genuinely good fit: no setup cost, no queue to maintain, and a human close enough to notice if it goes sideways.

Where shiploop adds the missing pieces

shiploop is the governed version of the same insight: work items are named tasks, not a flat todo list, dispatched to headless Claude Code workers (claude -p) at a cheap model floor (GOVERN_WORKER_MODEL, default sonnet) with a single escalation on evidence of failure (GOVERN_WORKER_ESCALATION_MODEL, default opus), stamped so a task can’t buy the ceiling twice. Every task gets its own worktree (worktree:new -- <name>), so parallel work never collides on a shared checkout. And when a worker hits a gotcha, the lesson gets promoted into a file the next session’s subagents and dispatched workers actually read, instead of getting rediscovered on the next pass through the loop.

None of this replaces the loop’s central bet, that fresh context beats accumulated context. It’s the same bet, with the parts a bare while true doesn’t do wired in around it. For the full picture on where automated task resolution can go wrong in other ways, see shiploop vs. the caveman approach. For a broader map of what’s on this site, see the guides index.

Last updated 2026-09-04.