Claude Code Headless Mode
Headless mode is claude -p "<prompt>": a single non-interactive run that prints a result and exits. There is no TUI, no session to attach to, and nothing left running once the process exits. It is for CI jobs, dispatchers that fan work out to many Claude Code processes, batch scripts, and hooks, anywhere a program needs to invoke Claude and read back an answer. It is not a background version of the interactive app: there is no live view to reconnect to, and a long-running task inside it either finishes before the process exits or is gone.
The flag combines with most other CLI options. Add --bare and Claude Code skips auto-discovery of hooks, skills, custom commands, subagents, plugins, MCP servers, auto memory, and CLAUDE.md, so a script gets the same result on every machine instead of picking up whatever a project’s .claude/settings.json happens to load. Anthropic’s own docs call --bare “the recommended mode for scripted and SDK calls” and say it will become the default for -p in a future release. Without --bare, -p still runs the hooks and MCP servers configured in the working directory, with no trust dialog and no per-server approval prompt, so know what you are pointing it at.
The invocation
claude -p "Run the test suite and fix any failures" \
--allowedTools "Bash,Read,Edit" \
--output-format json \
--permission-mode acceptEdits
Claude Code exits 0 on success and non-zero on failure, so a script can branch on the exit status without parsing output at all. An invalid flag is reported to stderr before the run even starts; a failure inside the run (bad auth, a dead API key) comes back as the result text on stdout instead, which is why exit code alone is not enough, see the parsing section below.
Output format controls what comes back:
text(default): the plain response, nothing else.json: one JSON object with the response inresult, plussession_idand cost metadata. Good for a script that just needs the final answer.stream-json: newline-delimited JSON, one event per line, as the run happens. Use this whenever you want per-turn visibility rather than one blob at the end, a dispatcher watching for a stuck tool call, a UI streaming tokens, or a supervisor that needs to see asystem/initevent fail before the first turn even runs. Pair it with--verbose.
Model and permissions. --model sonnet (or opus, haiku, fable, or a full model id) picks the model; --effort sets reasoning effort separately. -p starts in Manual permission mode on every plan, so pick one explicitly: --permission-mode acceptEdits for file edits without prompts, --permission-mode bypassPermissions (or --dangerously-skip-permissions) to skip prompts entirely, or --permission-mode auto to have a classifier review actions instead of a human. For a scheduled job where nobody is watching, add --permission-prompts none (Claude Code v2.1.259+) so anything that would have prompted is denied outright rather than hanging.
Tool allowlists. --allowedTools "Bash,Read,Edit" lets those tools run without a permission prompt; --disallowedTools denies specific tools or patterns; --tools restricts which tools are even present in the session. Bare tool names or Bash(git diff *)-style scoped rules both work, per the permission rule syntax.
Working directory. --add-dir ../apps ../lib adds directories Claude may read and edit beyond the invocation’s cwd. In bare mode, an added directory is a partial exception: its .claude/skills/ folder still loads, but its .claude/commands/ and .claude/agents/ folders do not.
cat build-error.txt | claude -p 'concisely explain the root cause of this build error' > output.txt
Piping stdin works like any other CLI tool and is capped at 10MB; past that, Claude Code exits with an error rather than truncating silently. Full flag list: CLI reference.
Parsing the result
With --output-format json, the single JSON object on stdout carries result (the response text), session_id, and cost fields including total_cost_usd, per Anthropic’s own docs on running Claude Code programmatically. Those cost figures are client-side estimates, not an invoice.
With stream-json, the last line of the stream is the type: result event, and it is the authoritative record of what the run cost, not the streamed deltas along the way. A dispatcher reading a worker’s outcome should do the same: take that final result event and pull .message.usage.input_tokens, .output_tokens, .cache_read_input_tokens, and .cache_creation_input_tokens for token accounting, and .total_cost_usd for spend. Reading only the streamed deltas and never the final event is the most common way to under-count a run’s real cost.
Exit code and the result event answer different questions and you need both. Exit code tells you whether the process ran to completion. The result payload tells you what the model actually concluded. A worker that hangs waiting on a permission prompt, or dies mid-tool-call, may never produce a result event at all, so a dispatcher has to treat “no result event before the process exits” as its own failure category, distinct from “result event says it failed.” The pattern that works is requiring the agent to emit a JSON object with a status field (a workable set: resolved, failed, parked, timeout) and separates that from three other outcomes it never conflates: an infra death (auth or transport failure, not the task’s fault), a hard timeout, and a plain no-report exit with nothing parseable. Collapsing all four into one bucket is how a dispatcher ends up retrying the wrong kind of failure.
The failure mode that costs the most to debug
The dangerous case in headless mode is not a crash, it is a flag that silently does nothing. If you pass a flag your installed claude binary does not recognize yet, that is not always a loud parse error, especially inside a wrapper script that only checks the exit code and a couple of expected fields. In a fleet where a dispatcher and its workers do not all run the same Claude Code build, a flag that exists on the machine you tested on can degrade into a generic, unclassified failure on a machine running an older binary, and that failure looks identical to a real task failure in your logs.
The fix is to never assume the CLI version you tested against is the CLI version that runs in production. Probe once, cache the result, and gate the flag behind it:
if [[ -z "${_CLAUDE_HELP_CACHE:-}" ]]; then
_CLAUDE_HELP_CACHE="$(claude -p --help 2>&1)"
fi
if grep -q -- '--exclude-dynamic-system-prompt-sections' <<<"$_CLAUDE_HELP_CACHE"; then
extra_flag="--exclude-dynamic-system-prompt-sections"
else
extra_flag=""
fi
Never gate on a version-string compare; a --help probe answers the only question that matters, does this binary understand this flag, and survives version-string formats changing under you. Ship a kill switch alongside every new mechanism, an env var that turns it off with no code change, and default new mechanisms off until you’ve watched them run clean at least once. Encoded as a rule: every new invocation flag is capability-probed via --help, cached, wired behind an env var, and defaults off, because the fleet’s CLI version is never guaranteed to be yours.
Tool budget: the schema tax you pay before the first turn
A headless -p run still sends the full tool schema block to the model on the first turn unless you trim it, even for tools the run structurally cannot use. An interactive-only tool like a workflow launcher or a live-monitor tool needs a human in the loop or a capability a one-shot process doesn’t have, but its schema is sent anyway if you don’t explicitly restrict tools with --tools or --disallowedTools.
shiploop measured this directly on one real worker spawn (Claude Code CLI 2.1.220, model opus, single measurement, not a controlled benchmark): before trimming, the tool schema block was 85,260 bytes shiploop PROOF.md, worker request payload split , or 51.7% shiploop PROOF.md, worker request payload split of the entire first-turn request, larger than the system prompt and the conversation messages combined. The five biggest single schemas accounted for 46,127 bytes shiploop PROOF.md, worker request payload split of that block, and every one of them was a tool a headless worker cannot call at all: a workflow-launch tool, a design-sync tool, a live monitor, a worktree-entry tool, and a wakeup scheduler. Restricting the tool list to what a headless worker can actually use cut the schema block to 28,417 bytes shiploop PROOF.md, worker request payload split (down 66.7% shiploop PROOF.md, worker request payload split ) and the total request size to 107,985 bytes shiploop PROOF.md, worker request payload split , a 34.5% shiploop PROOF.md, worker request payload split drop overall. That is one measurement on one CLI version, so treat the exact byte counts as directional, but the shape of the finding, unusable tools still cost tokens on every single turn until you restrict them, holds regardless of version. See reducing Claude Code token usage for the rest of the token-cost picture and how MCP tool listings interact with the same problem.
Restrict with --tools (an explicit allowlist) or --disallowedTools "mcp__*" style deny patterns; see Claude Code permissions for the full rule syntax and how allow and deny rules interact with permission modes. Worth noting as a live example of the previous section: --tools is not listed in the published CLI reference, which itself warns that claude --help does not list every flag either. Probe for it rather than assuming it either way, which is exactly the discipline the failure mode above demands.
shiploop’s dispatcher
shiploop is a headless dispatcher for Claude Code. Every piece of work you hand it becomes a ticket, shiploop’s own term for a unit of dispatched work, the one used throughout its queue and config. It runs one disposable claude -p session per work ticket, each in its own git worktree, reads the final result event for cost and status, and escalates a ticket to a stronger model exactly once if the floor-tier attempt fails, instead of retrying blind. It restricts each worker’s tool list to what a headless process can use, probes new CLI flags behind --help before depending on them, and treats a missing report as its own failure category rather than lumping it in with a task the model genuinely failed. None of that is specific to shiploop, it is the same discipline any headless Claude Code deployment needs once more than one machine or CLI version is involved.
For related patterns, see how to pick a model for headless workers, when to reach for subagents instead of a fresh -p process, and the full guide index.
Last updated 2026-09-04.