What changed in each shiploop release
This page mirrors the most recent entries of the hub CHANGELOG.md. The unreleased
section below is reproduced in full; the two released versions after it are condensed
for length, with detail cut rather than reworded. For the complete, unabridged history
back to the project’s first release, read
CHANGELOG.md on GitHub.
# Changelog
## Unreleased
### Fixed
**Parallel worktrees could exhaust a machine's RAM, and finished worktrees were never reclaimed.**
Two independent leaks, both surfacing as "my laptop is unusable and my disk is full" rather than as
any test failure.
- **Concurrent dependency installs had no global cap.** A project bootstrap hook installs deps for
several sub-repos at once and backgrounds each one; the obvious throttle, a `wait` at the end of
the bootstrap, is per-worktree and cannot see another worktree's installs. Under
`GOVERN_PARALLEL_DEFAULT` tasks in flight the real ceiling was (tasks x repos x ~1 GB per
install). Measured on a 24 GB laptop: 13 concurrent installs, ~14 GB resident, 3.7 GB swap,
unusable for minutes. Neither layer was individually wrong, the governor's concurrency knob
reasons about API spend, the bootstrap's `&` reasons about one worktree, and nothing anywhere was
denominated in machine memory.
New `scripts/lib/install-semaphore.sh`: a cross-process, package-manager-agnostic cap (`mkdir`
atomic, works on stock bash 3.2). Default 4 via `WORKTREE_INSTALL_PARALLEL`. A slot orphaned by a
killed install is reclaimed by pid liveness; a slot wait degrades to running uncapped and loud
after `WORKTREE_INSTALL_WAIT_TIMEOUT` rather than deadlocking every session on the machine.
`worktree/new.sh` documents the seam so a project bootstrap opts in with two lines.
Note for anyone measuring this: `ps` RSS does not show the spike, it reported a 364 MB maximum
while the OS process monitor showed 1.18 GB for the same processes, because RSS excludes
compressed and swapped pages.
- **New `<pm> run worktree:reap`** reclaims disk from worktrees nothing needs. Worktrees accumulate
from two directions: `run-loop.sh` preserves parked/failed tasks' worktrees on purpose and never
ages them out, and hand-made worktrees have no cleanup path at all. Measured: 26 GB across 12
worktrees.
Dry-run by default. A worktree is reaped only when it is idle AND clean AND merged across every
repo *including the meta worktree root*, and any check that cannot be answered blocks the reap.
That last part is load-bearing: of four worktrees with zero commits off `origin/main`, one had a
live session working in it and another had uncommitted files, deleting on the merge check alone
would have destroyed both. `--strip` drops regenerable build output from kept-but-idle worktrees
instead of removing them.
### Added
**Fleet visibility, the governor's live state, on four surfaces.** Governor workers are detached
`claude -p` processes whose pid lives only in a bash array inside `run-loop.sh`; structured state was
written only at completion, so while a run was in flight *nothing on disk said "running"* and no
surface could show it. (Claude's own subagent panel is not an option: it renders Task-tool children
of the session, and cannot be injected into from outside.)
One append-only event log is now the single source of truth, and everything folds it.
- **`GOVERN_EVENTS` (default `0`, off)**, `scripts/govern/lib/events.sh` adds `govern::event`, which
appends one JSON object per line to `governor/events.jsonl`. Types: `run_started`,
`driver_spawned`, `worker_spawned`, `worker_escalated`, `worker_done`, `ticket_parked`,
`driver_reaped`, `run_done`. Always-on fields are `ts`, `run_id` (the TokenJam run id, so events
join against the OTel attributes workers are already tagged with), and `type`. **The emitter can
never abort a run**, the whole body is a guarded group with an explicit `return 0`, so an
unwritable log, a full disk, or a malformed key is swallowed silently under `set -euo pipefail`.
Nothing about a run changes at `0`, and existing installs are unaffected until they opt in.
- **`npm run govern:status`** (`scripts/govern/status.sh`), one-shot reader, text by default,
`--json` for machines. Folds the log **last-event-wins per (run_id, ticket)**, which is what makes
a retry (spawn, done, spawn) read as active where a spawned-minus-done count would not. Verifies
every claimed-live worker with `kill -0` and reaps the phantoms a killed driver leaves behind,
appending a synthetic `status:"stale"` row so the log self-heals. No jq dependency, no model call,
no lock, runnable from inside a Claude session, from CI, or over SSH. `--no-reap` and `--all-runs`
included.
- **Plugin monitor** (`monitors/monitors.json` + `tools/fleet-monitor.sh`), the in-session channel.
Every stdout line becomes a notification in the driver's context, which is the resource shiploop
exists to conserve, so it emits **state transitions only** (never a raw tail), dedupes repeated
states, caps itself at `GOVERN_MONITOR_MAX_PER_MIN` (6) lines a minute with the overflow collapsed
into one line, attaches at the *end* of the log so history is never replayed, and prints absolutely
nothing when there is no event log, which is nearly every session. `GOVERN_MONITOR=0` disables it.
- **`/shiploop:statusline`** + `scripts/govern/statusline-{segment,chain,install}.sh`, an opt-in
statusline segment, silent when no fleet is running. **It chains, it does not replace.**
`statusLine.command` is a single string, so an installer that writes its own value destroys a
user's ccusage or custom HUD; instead the installer records the *entire* previous `statusLine`
object verbatim to `~/.claude/shiploop-statusline.json`, wraps it (stdin is read once and replayed
to the original, whose output comes first), and uninstall restores the recording byte for byte,
including removing the key entirely when there was none. It refuses to re-record over an existing
recording, and refuses to touch a malformed `settings.json`. `refreshInterval` defaults to 5s
(yours wins if you had one); without it the elapsed time freezes while the session is idle. Never
installed by `scaffold.sh`, `/shiploop:setup`, or `/shiploop:update`.
- Five new tests, covering the never-abort contract, the fold, stale reaping, the monitor's rate
limit and silence, and the statusline chain/restore. `GOVERN_EVENTS=0` is exported from
`test/assert.sh` for the whole suite, per the standing dispatch-path-mechanism rule.
## 1.17.2, 2026-08-03
### Fixed
Documentation drift: six places where a shipped doc described a command, env-var default, or
slash command that the code stopped honouring some releases ago. No behavior changes, but four
of the six ship into every scaffolded workspace, and two of those land in an always-loaded
`CLAUDE.md`, so a model paid for the wrong fact on every turn.
- **The seed `CLAUDE.md` advertised `npm run status`, which no longer exists.** The Commands line
also omitted `sync` and `tail`, which *are* shipped, wrong in both directions.
- **`GOVERN_BATCH_MAX` was documented as `default 1 = off` in two places.** The real default has
been `2` since batching was re-keyed onto measured file overlap. All four references now agree:
default `2`, `1` disables.
- **`GOVERN_RUN_MAX_WORKERS` does not exist.** The knob that actually bounds a run's task count
is `GOVERN_MAX_TICKETS`.
- **Three references to commands retired earlier.** The trigger was cut deliberately; the loop was
not.
- **Command tables under-documented the shipped script set.** Both tables now match the real
package.json script list exactly.
## 1.17.1, 2026-08-03
### Fixed
- **New knobs never reached any workspace.** They were registered in the config checker, which only
*reports* what is set, but not in the template `/shiploop:update` actually distributes, so the
append-only merge had nothing to append.
The sharp edge was `GOVERN_WORKER_ESCALATION_MODEL`. Undeclared, it fell back to the script
default, while the workspace template still pinned the worker floor to the same value. Floor and
ceiling both resolved to the same tier, which re-collapses the two knobs the split exists to
separate and makes escalation a same-tier re-bet, precisely the failure the split was built to
prevent, shipped by the release that built it.
**Existing fleets needed one manual edit** to lower an already-pinned floor, since the append-only
merge can add a new ceiling but cannot change a value already set. Fresh scaffolds got the correct
pair with no action.
Full release history, including everything before these three entries, is on GitHub. See the homepage for what shiploop is and how its cost claims are measured.
Last updated 2026-09-03.