Claude Code Subagents
A Claude Code subagent is a Markdown file with YAML frontmatter, stored in .claude/agents/ (project) or ~/.claude/agents/ (user), that defines a named helper Claude can delegate a task to. It runs in its own isolated context window, does the work there, and returns only a summary back to the conversation that spawned it. Nothing else transfers back: not its tool calls, not the files it read, not its intermediate reasoning. The parent pays for the summary, not the work.
That last sentence is the reason subagents matter more than the “organize your workflow” framing most writing gives them. Every other token-reduction technique (shorter tool output, capped results, /compact) shrinks what is already in your context. Delegation is different: it keeps bytes out of your context in the first place. A codebase sweep that would dump forty thousand tokens of file contents into your window instead happens in a subagent’s window, and you never pay for that transcript again, not now and not on every later turn it would otherwise sit in.
What a subagent file looks like
Per Anthropic’s subagent docs, only two frontmatter fields are required: name and description. Everything else is optional.
---
name: test-runner
description: Runs the test suite and reports failures with file:line references
tools: Bash, Read, Grep
model: haiku
---
You are a test-runner subagent. Run the project's test command, then report
only the failing tests with their file and line number. Do not paste full
test output back to the caller.
The frontmatter block configures the subagent; the Markdown body below it becomes the subagent’s system prompt. The full documented field set:
| Field | Required | What it does |
|---|---|---|
name | Yes | Unique identifier in lowercase letters and hyphens; used as the agent type for hooks and CLI references |
description | Yes | Tells Claude when to delegate to this subagent; kept short since every subagent’s description shares one combined token budget |
tools | No | Allowlist of tools the subagent may call; inherits every available tool if omitted |
disallowedTools | No | Denylist applied before the tools allowlist resolves |
model | No | Which model tier runs the subagent: sonnet, opus, haiku, fable, a full model ID, or inherit to match the caller |
permissionMode | No | default, acceptEdits, auto, dontAsk, bypassPermissions, plan, or manual; a parent session already running bypassPermissions or acceptEdits overrides it |
maxTurns | No | Stops the subagent after a fixed number of agentic turns and marks its output partial |
skills | No | Preloads named skill content directly into the subagent’s context at startup |
mcpServers | No | MCP servers available to this subagent, by reference or inline definition; ignored for plugin subagents |
hooks | No | Lifecycle hooks that run only while this subagent is active |
memory | No | Enables persistent memory across sessions: user, project, or local, each backed by its own directory |
background | No | Keeps the subagent running detached from the caller even when Claude would otherwise run it in the foreground |
effort | No | Effort level while this subagent runs: low, medium, high, xhigh, or max, depending on what the model supports |
isolation | No | worktree runs the subagent in its own temporary git worktree, cleaned up automatically if it makes no changes |
color | No | Display color for the subagent in the task list and transcript |
initialPrompt | No | Text auto-submitted as the first turn when this subagent runs as the main session agent |
experimental | No | Experimental options, currently just a cacheTtl key, to set the prompt cache lifetime |
Definitions can live in several places: .claude/agents/ for the current project, ~/.claude/agents/ for every project on the machine, a plugin’s own agents/ directory, the --agents CLI flag as JSON for a single session, or managed settings deployed organization-wide. Both .claude/agents/ and ~/.claude/agents/ are scanned recursively, so subfolders are fine for organizing many subagents. Project subagents belong in version control; user subagents are personal and apply across every project you open.
Three of these fields carry outsized weight for cost and safety. model and effort together set how much compute a subagent burns per turn, and picking the wrong tier either wastes budget on a trivial lookup or starves a genuinely hard synthesis task of the reasoning it needs; see Claude Code model selection for how to size that call. isolation: worktree is the difference between a subagent editing code safely on its own branch and two agents silently colliding in the same working tree; see the guide to Claude Code worktrees. permissionMode decides whether a subagent can edit files or run commands without asking, which matters most exactly when it’s also been handed a broad tool allowlist; see the guide to Claude Code permissions.
At startup, a subagent’s context holds its own system prompt (not Claude Code’s default one), the task description from the caller, CLAUDE.md files from the directory hierarchy, a git status snapshot, and any preloaded skills. What does not carry over: the parent conversation’s history, its auto memory, and previously invoked skills. This is why a fresh subagent needs to be told things the parent already knows, not just pointed at “the current task.”
You invoke one by naming it (“use the test-runner subagent to check this”), by @-mentioning it to force delegation, or by setting it as a session default via claude --agent test-runner or "agent": "test-runner" in .claude/settings.json. Full mechanics are in the subagent documentation.
When delegation pays, and when it doesn’t
The rule is simple: delegate work whose output you want but whose transcript you don’t. Delegate:
- Codebase sweeps across many files
- Log trawling and error investigation
- Verbose builds and test runs
- Multi-file “where is X handled” investigation
Do not delegate work whose context you already hold. If you just read the three files that answer a question, answering it yourself is free; spawning a subagent to re-read those same three files buys you a cold start where it rediscovers what’s already sitting in your window, then pays again to summarize it back to you. Delegate execution, not understanding you already have.
Sizing the model to the job
A subagent’s model field, or the CLAUDE_CODE_SUBAGENT_MODEL environment variable, controls which tier it runs on. Anthropic’s own guidance: use haiku for simple read-only exploration, sonnet (the default) for balanced work, opus for complex reasoning, and inherit to match the caller’s model. To force every subagent onto one tier regardless of its own model field, set both CLAUDE_CODE_SUBAGENT_MODEL and CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1.
In practice this maps to a three-tier ladder:
- Cheap tier (haiku): mechanical extraction, lookups, “does this string exist,” single-file reads.
- Mid tier (sonnet): search, investigation, standard edits.
- Top tier (opus): judgment-heavy synthesis, architectural calls, anything where being wrong is expensive.
A fan-out of many similar children (ten files each needing the same grep-and-summarize pass) is never top-tier work, even if the codebase overall is complex. Size each child to its own task, not to the project’s difficulty.
The stronger version of this rule is escalate-once, never downgrade: start every task at the cheap floor for its class, and if it fails, retry one tier up. Never start expensive and drop down on success, because you can’t know in advance which cheap attempts would have worked, and a failed cheap attempt is informative (you now know the floor wasn’t enough) in a way a successful expensive one never confirms was necessary. This only holds if failures die cheaply and fast; a task where a wrong cheap attempt is costly to detect or undo should start higher.
Agent teams are a different tool
Agent teams are not subagents with a different name. A subagent returns a result to its caller and the caller manages all the work; a team of teammates message each other directly, share a task list, and coordinate themselves. Anthropic’s own comparison: subagents are for “focused tasks where only the result matters,” teams are for “complex work requiring discussion and collaboration.” See the agent teams documentation for the full architecture.
Teams are disabled by default. Enable them with:
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
The cost difference is not subtle. Anthropic states that agent teams use approximately 7x more tokens than standard sessions when teammates run in plan mode Anthropic, Manage costs effectively (code.claude.com/docs/en/costs) . Each teammate is a full separate Claude Code instance with its own context window; token usage scales with how many are active, not with how much work actually needs doing. Anthropic’s own cost-control advice: use Sonnet for teammates rather than Opus, keep teams small, keep spawn prompts focused, and shut teammates down when their work is done rather than leaving them idle.
If a task doesn’t need teammates challenging each other’s findings or claiming work off a shared list, it doesn’t need a team. Most delegation is a subagent’s job.
Failure modes nobody writes about
The docs describe the happy path. Real operating experience surfaces four problems that show up regardless of model tier.
Subagents fabricate with confidence. A subagent will invent a quote, describe file contents it never opened, or state an API detail it never verified, and it will say all of this in the same tone it uses for things it actually checked. Never relay a subagent’s number, quote, or file claim as fact without spending one grep or Read to verify it yourself first. Treat what comes back as a lead, not a source.
A subagent does not load a nested repo’s CLAUDE.md. If you delegate work into a sub-repo that has its own rules file, and the subagent operates one directory up or down from where that file lives, it never sees the gotcha someone already wrote down. It re-hits the same bug, re-debugs it from scratch, and burns exactly the tokens the CLAUDE.md file existed to save. Name the sub-repo’s rules file explicitly in the delegation prompt, and hand over any rule the parent’s own CI enforces, since the subagent won’t inherit that either.
Two children building two halves of one contract each pass their own tests. If you fan out “build the API client” to one subagent and “build the server handler” to another, each one writes tests against its own understanding of the shape between them, both pass, and the seam is still wrong: a null where the other expected a zero, a string where the other expected an enum. Write the wire shape down first, in one place, and hand both subagents the identical text. Testing each half alone never catches a seam bug; only testing the placement of the actual value across the boundary does.
One working tree per repo, per agent. Pointing two subagents (or a subagent and yourself) at the same git checkout means the second one’s branch switch happens underneath the first one’s uncommitted work, and one of them loses changes silently. See the guide to Claude Code worktrees for the isolation pattern; the short version is that a delegated agent doing code work needs its own worktree, not a shared one.
Shiploop’s model
Every piece of work shiploop dispatches is called a ticket, its own term for a unit of work, the same word used in queue/tickets.md and config vars like GOVERN_MAX_TICKETS.
Shiploop is a self-improving Claude Code harness that dispatches one worker per queued ticket. Each worker starts at a fixed model floor and can escalate exactly once per ticket on failure, never the reverse: the floor defaults to
sonnetviaGOVERN_WORKER_MODEL, and an escalated attempt runs atopusviaGOVERN_WORKER_ESCALATION_MODEL, stamped so a second retry doesn’t re-buy the same ceiling. Every dispatched model is also clamped so a worker can never run above the spawning session’s own model family.Workers get a trimmed tool allowlist rather than the full default set, and each worker’s prompt is explicit that a subagent’s claim is a lead to verify, not a fact to relay, matching the failure mode above. When a worker touches a sub-repo, it is handed that sub-repo’s own rules directly in its prompt rather than relying on it to load them on its own.
Last updated 2026-09-04.