Claude Code Token Usage

Claude Code token usage climbs because the client resends the entire conversation on every request, not because any single message is large. The bill for a session is not the size of one prompt, it is the size of the accumulated context multiplied by the number of turns still to come. Every tool call is its own request carrying the growing history plus that call’s results, and prompt caching changes the rate the history is billed at, it does not make re-reading it free. That’s the mechanical answer to “where do my claude code tokens go”: mostly the same conversation, read again and again, every request growing a little heavier than the last.

Once that clicks, the rest of claude code token cost follows from it: anything that adds bytes to the conversation (verbose tool output, a long CLAUDE.md, deferred MCP tool schemas expanding into context) taxes every future request in the session, not just the one that produced it. Anything that removes turns (clearing, delegating, compacting) is a different kind of lever from anything that trims bytes, and they don’t substitute for each other.

Why usage climbs in a long session

Anthropic’s own list of what drives cost up in a running session, from Manage costs effectively:

  • Long context: the full conversation gets resent every request.
  • Cache misses after a break longer than the prompt cache lifetime.
  • Scheduled tasks firing while you’re idle.
  • Cross-session messages delivered as a new turn.
  • Goal check-ins (capped at three per goal since v2.1.246; disable with CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0).
  • Active agent teammates.
  • /compact itself, which is a large request in its own right.

The prompt cache is the piece people misread as “free reuse.” It is not free, it is cheaper: a cache hit means the history is billed at the cached read rate instead of full input price, and the cache expires ( one hour source on a subscription, dropping to five minutes source once you’re drawing on usage credits; five minutes by default on an API key or cloud provider). Step away longer than that and the next request is a full-price cache miss on everything that came before.

How to actually see your usage

Run /usage in a session. Its Session block reports Total cost, Total duration (API), Total duration (wall), Total code changes, and Usage by model. As of v2.1.251, it also prints a Prompt cache (main) line: request count, the share of input tokens served from cache, and a miss count. A request counts as a miss when it re-processed more than 5% source and at least 2,000 tokens source of what it could have read from cache instead. That line covers the main conversation only, not any subagents you spawned during the session, so a heavy subagent won’t show up in it.

/usage also shows a plan usage breakdown (Pro, Max, Team, Enterprise) that attributes recent usage to skills, subagents, plugins, and individual MCP servers. Press d or w to switch between the last day (d) or the last week (w). This is computed entirely from local session history on the machine you’re running on, so it won’t reconcile across machines or team members.

Beyond /usage:

/context   # what's currently consuming context space
/status    # remaining allocation
/insights  # writes ~/.claude/usage-data/report.html (up to 200 new sessions per run)

For a fleet or a team, none of the above aggregates across people or machines. The OpenTelemetry export is the option that streams per-user token and cost metrics into your own stack in near real time; it’s the only path here built for cross-machine visibility.

One accounting note: cost figures from /usage are computed locally from token counts at list price, unless a modelPricing managed setting is in effect. They’re estimates, not an invoice. Totals used to accumulate across /clear; since v2.1.211 they reset when /clear starts a new session.

Where the bytes actually go: a measured breakdown

The shiploop project instrumented one real headless worker request end to end (scripts/govern/measure-prefix.sh, Claude Code CLI 2.1.220 shiploop PROOF.md , --model opus) to see the actual composition of a first-turn request before any conversation has even accumulated. This is a single measurement (n=1, one CLI version, one worker spawn), not a controlled benchmark, so treat the shape as informative and the exact bytes as a snapshot that will drift across CLI versions.

ComponentBytesShare of request
Tool schemas 85,260 shiploop PROOF.md 51.7% shiploop PROOF.md
Messages 71,981 shiploop PROOF.md 43.7% shiploop PROOF.md
System prompt 7,093 shiploop PROOF.md 4.3% shiploop PROOF.md
Other 461 shiploop PROOF.md 0.3% shiploop PROOF.md
Total 164,795 shiploop PROOF.md n/a

Tool schemas alone were the largest single component, ahead of the actual conversation. Broken down further, the top five schemas accounted for 46,127 shiploop PROOF.md of those 85,260 shiploop PROOF.md bytes:

ToolBytesShare of tool blockShare of whole request
Workflow 21,525 shiploop PROOF.md 25.2% shiploop PROOF.md 13.1% shiploop PROOF.md
DesignSync 8,978 shiploop PROOF.md 10.5% shiploop PROOF.md n/a
Monitor 7,767 shiploop PROOF.md 9.1% shiploop PROOF.md n/a
EnterWorktree 4,027 shiploop PROOF.md 4.7% shiploop PROOF.md n/a
ScheduleWakeup 3,838 shiploop PROOF.md 4.5% shiploop PROOF.md n/a

Every one of those five tools was flagged in the source measurement as unusable by a headless -p worker in the first place: a worker with no interactive user structurally cannot call Workflow (it requires explicit user opt-in), and the other four have no path to fire outside an interactive session either. The schema was billed on every request regardless.

What actually reduces token usage

Anthropic’s own reduce-usage list, from the same costs doc, roughly ordered by how much a given change actually moves the number rather than by how often it’s mentioned:

  • Delegate verbose work to subagents. This is the one lever that removes bytes from the accumulation entirely rather than shrinking them: a subagent reads the noisy output in its own context, and only its conclusion returns to your conversation. Everything below this either trims what already accumulated or reduces how many turns happen; delegation is the only one that stops bytes from entering the running total in the first place. See Claude Code subagents for how the scoping works.
  • /clear between unrelated tasks. /rename first if you’ll want it back, /resume to return. This caps how many turns of unrelated history keep getting resent.
  • Hooks that preprocess data before Claude sees it. The doc’s own example: a PreToolUse hook matching Bash that rewrites a test command to grep for failures only, cutting “tens of thousands of tokens to hundreds.” See Claude Code hooks.
  • Choose the right model, /model mid-session, or model: haiku in subagent config. Model choice doesn’t change how many bytes get resent, but it changes the price of resending them. More on tiering at Claude Code model selection.
  • Prefer CLI tools over MCP servers where both exist (gh, aws, gcloud). MCP tool definitions are deferred by default, but each connected server still adds to what can get loaded; /mcp disables ones you aren’t using.
  • Move instructions from CLAUDE.md into skills. The doc’s target: keep CLAUDE.md itself under two hundred lines, since it’s resent in full every turn same as the rest of the conversation.
  • /compact Focus on... with a specific instruction, or a custom compact instruction in CLAUDE.md, rather than a bare /compact. Note /compact itself is a large request.
  • Lower reasoning effort with /effort or MAX_THINKING_TOKENS. Adaptive-reasoning models ignore a nonzero budget cap, so use effort levels on those instead of the token env var.
  • Agent teams: keep teammates on Sonnet, keep the team small, keep spawn prompts focused, and shut teammates down when done. Agent teams running in plan mode use roughly 7x source more tokens than a standard session, so this only pays off if you actually need the parallelism.
  • Plan mode (Shift+Tab) before implementation, Escape to course-correct instead of letting a wrong direction run, /rewind to a checkpoint instead of re-deriving context by hand.

How shiploop cuts this for headless workers

shiploop dispatches tickets to headless Claude Code workers (claude -p) and applies three of the levers above by construction rather than leaving them to each worker’s judgment. Each ticket is just a named piece of work, shiploop’s own term for the unit it queues and hands off, the same word you’ll see in queue/tickets.md and in config like GOVERN_MAX_TICKETS.

First, it trims the tool schema block. A headless -p worker gets a --tools allow-list instead of the CLI’s full default set, dropping tools like Workflow and Monitor that a headless worker can’t call anyway. Measured on the same worker spawn as the table above, that trim took the request from 164,795 shiploop PROOF.md bytes to 107,985 shiploop PROOF.md bytes, a 34.5% shiploop PROOF.md reduction in total request size, driven by a 66.7% shiploop PROOF.md cut to the tool schema block specifically. Same caveat as above applies: n=1, one CLI version, one worker spawn on --model opus, not a benchmark.

Second, it routes each ticket to a cheap model floor (Sonnet by default, configurable) and escalates to a higher tier only once per ticket, on evidence of an actual budget or judgment failure, rather than running every ticket at the most capable (and most expensive) model by default.

Third, each ticket runs in its own disposable session and worktree. Because the mechanism above is that cost tracks accumulated context times remaining turns, a worker that starts fresh per ticket never carries forward the history of every other ticket it touched. That’s a structural version of /clear between tasks, applied automatically instead of relying on a human to remember to run it.

None of this changes the physics: the model still resends what’s in context on every request. It changes what’s in context, and how much of it there is to resend.

Last updated 2026-09-04.