Caveman Claude Code: What It Saves and What It Doesn't
Caveman is a skill for AI coding agents, Claude Code among more than thirty others, that makes the assistant’s replies shorter. Install the skill, the agent starts answering in clipped, caveman-style prose instead of full paragraphs. Should you use it? If your agent’s verbosity is the thing that bugs you in a solo interactive session, yes, it is a genuinely good one-command fix. It does not touch the size of your conversation history, your tool output, or your thinking tokens, and its own README says so plainly. That distinction is the whole story below.
What caveman actually does
Caveman ships two things under one name. The small one is a skill: a rule file that changes how the model talks. Install it with:
npx skills add JuliusBrussee/caveman
The repository also carries a .claude-plugin/marketplace.json declaring a marketplace named
caveman with a single plugin of the same name, so inside Claude Code it installs the same way any
other plugin marketplace does:
/plugin marketplace add JuliusBrussee/caveman
/plugin install caveman@caveman
Once installed, /caveman lite, full, ultra, and wenyan variants set how terse the agent gets, and /caveman off restores normal replies. Code, commands, file paths, and error messages are left untouched. Only the surrounding prose gets cut.
The bigger piece is a local proxy, installed with npm install -g @caveman-ai/cli && caveman setup --install, that sits between the agent and the provider and compresses what the agent reads, not what it writes: JSON, logs, diffs, and search results get squeezed before they reach the model, with the original bytes recoverable on disk. This is a different mechanism from the skill, and it changes your setup, not just your prompt.
The honest caveat, in the project’s own words
Caveman’s README reports a headline output-token saving of 65% caveman README, 'The numbers' for the skill, averaged across ten coding prompts run through the real Claude API. It also states, right below that table, why you should not apply that number to your invoice:
“Input and reasoning tokens don’t change”
The same section adds that the skill’s own rules cost roughly one to one and a half thousand input tokens on every turn, so whole-session savings land below the headline figure, and can go negative on work that was already terse. The README says as much itself: shorter replies do not guarantee a smaller bill on every task.
That caveat is the correct one to lead with, because in a real Claude Code session, output is usually not where the money goes. The full conversation is resent on every request, and every tool result, log tail, and file read joins that history for every remaining turn. See Anthropic’s own guidance on managing Claude Code costs for the mechanism. A tool that shortens replies is working on a slice of the bill that keeps shrinking relative to everything else piling up on the input side.
Caveman’s own proxy exists for exactly that reason. Its README reports a separate, more modest, real-workload figure for the input side: 33.2% caveman README, 'Proxy: reading less' lower input tokens across a pinned fifty-four run Claude Code benchmark, with per-case results ranging from a fifty-five percent cut down to one case that came in nine point nine percent worse because there was nothing compressible to squeeze. That is the more useful comparison to shiploop, because it targets the same half of the bill.
Caveman vs shiploop
| What it changes | Side of the bill | What it costs you | Right pick when | |
|---|---|---|---|---|
| Caveman (skill) | Shortens the model’s prose | Output tokens only | Terser, clipped replies; its own rules add extra input tokens every turn (about one to one and a half thousand, per the README) | You want less throat-clearing in an interactive session, zero workflow change |
| Caveman (proxy) | Compresses logs, diffs, JSON, search results before the model reads them | Input tokens (what’s read) | A local process in the loop; one case in its own benchmark got worse, not better | You read large noisy tool output constantly and want it squeezed automatically |
| shiploop | Trims the tool-schema block a headless worker loads, routes work to a cheaper model tier, keeps each task in its own disposable session | Input tokens (what accumulates and what’s resent) | Only pays off for delegated, task-shaped work; adds a workspace/governor structure | You run many small, repeatable coding tasks and want the accumulation itself capped |
shiploop’s own measurement is on the same accumulation problem: a real headless worker’s first-turn request was 164,795 bytes shiploop PROOF.md, worker request payload measurement, n=1 , over half of it ( 85,260 bytes shiploop PROOF.md, worker request payload measurement, n=1 ) just the tool-schema block. Trimming that block to the tools a headless -p worker can actually call cut it to 28,417 bytes shiploop PROOF.md, worker request payload measurement, n=1 , taking the whole request down 34.5% shiploop PROOF.md, worker request payload measurement, n=1 . That is a single measurement on one CLI version, not a controlled study, and it is called out as such in the source.
Use caveman if
You are working solo, interactively, and the thing that annoys you is watching the model write three paragraphs where two sentences would do. That is exactly what the skill fixes, in one command, with no change to how you work. If you also read a lot of dense tool output (long test failures, big diffs, sprawling logs) and want that squeezed automatically, the proxy is the piece for that, and it is a separate, larger install with a local process to keep running.
Use shiploop if
Your problem is not one session getting verbose, it’s many similar tasks each resending an ever-growing tool-schema block and letting context pile up turn over turn with nothing discarded. shiploop cuts what a headless worker loads before the first request, routes routine work to a cheaper model tier by default, and closes each task’s session so nothing from it carries into the next one. That is a different failure mode than a chatty reply, and a different fix.
They compose
Nothing here is exclusive. Running the caveman skill for shorter interactive replies and shiploop for how your delegated work is structured and dispatched addresses two different terms in the same bill, and neither one substitutes for the other. If you already run one and are curious what’s driving the other half of your spend, start with what tokens you’re actually paying for and where Claude Code cost goes in practice. If your usage is dominated by delegated work specifically, how subagents affect your context and cost is the more relevant next read.
Last updated 2026-09-04.