Changing the Claude Code Model, and Which One to Pick
Type /model in any Claude Code session and pick a model from the picker, or type /model sonnet (or opus, haiku, fable) to switch immediately without opening it. That change is mid-session only unless you press Enter in the picker, which also saves it as your default. For a non-interactive run, pass --model on the command line instead. The rest of this page is about the harder question underneath the mechanical one: which model, and when switching actually helps versus when it does nothing at all.
The command, every way you’d need it
Mid-session, interactively:
/model
/model sonnet
/model opus
Opening the bare /model picker shows an effort-level slider alongside the model list and current pricing when you’re on the Anthropic API. Pressing Enter switches the model and saves it as your default; pressing s switches it for the current session only, without touching the default.
Setting a default before you ever start a session, in settings.json:
{
"model": "sonnet"
}
At startup, for one run:
claude --model opus
Environment variable, which several people ask about because it’s the one that survives across every terminal and script without editing a config file:
export ANTHROPIC_DEFAULT_MODEL=sonnet
And in a subagent’s own frontmatter, which is a separate knob from your interactive session’s model:
---
model: haiku
---
model: inherit in that same field makes the subagent use whatever model the parent session is currently on, instead of pinning its own.
Which model for which job
Anthropic’s own guidance is that Sonnet handles most coding tasks, and Opus should be reserved for complex architectural decisions and multi-step reasoning, per the model configuration reference. In practice that maps to a simple table:
| Job | Tier | Why |
|---|---|---|
| Grep a symbol, extract a value, reformat a file, apply a known fix | haiku | Mechanical. No judgment call in the loop, so the cheapest tier does it exactly as well. |
| Investigate a bug, search across files, write a standard feature or fix | sonnet | This is the daily-driver tier: enough reasoning for real code work without paying for headroom you don’t need. |
| Design an architecture, resolve a genuine tradeoff, do the final review before something ships | opus | Judgment and synthesis, where a wrong call is expensive to unwind later. |
That table is a starting allocation, not a hard rule you set once. The move that actually saves money is deciding tier per task rather than per person or per project.
Start cheap, escalate once
The counterintuitive part, and the one that’s easy to get backwards: don’t start on the expensive model “to be safe.” Start at the cheapest tier that could plausibly do the job, and escalate to the next tier up only if that attempt fails. The reason this works out cheaper on average, not just morally cheaper, is that a failed attempt on a cheap model dies early: it hits a wall, produces an obviously wrong or incomplete result, and stops burning tokens well before a full successful run would. A wrong guess at the floor tier costs a fraction of what a correct answer at the top tier costs, so the expected cost of “try cheap, escalate once” beats “always start expensive” even when the cheap tier fails some of the time.
Each of those tasks, once named and queued for dispatch, is what shiploop calls a ticket, its internal term for a single unit of work. We measured this directly running a dispatch harness across dozens of real tickets: mean cost per resolved ticket was $4.49 shiploop PROOF.md , median $3.03 shiploop PROOF.md , over a sample skewed toward harder, opus-tier work, which makes it closer to an upper bound than a floor-tier average. The caveat that matters more than the number: this is one workspace’s operational log, not a controlled experiment, and it doesn’t isolate escalation savings from everything else that changed. What it does show is that resolved-ticket cost stays in single digits per unit of work even when the sample leans expensive, which only happens if most work is getting done well below the top tier.
Escalate on genuine failure, not on a hunch that the bigger model would be safer. Re-running the exact same task at a higher tier the moment a cheap attempt succeeds is pure waste; the floor tier earns its keep by actually trying first.
Model limits: when switching helps and when it doesn’t
Claude Code’s hit-a-limit messages come in two different shapes, and only one of them is fixed by changing model.
“You’ve hit your Opus limit” (or “You’ve hit your Sonnet limit”) is model-specific. Switching to a different model family with /model keeps you working immediately, per Anthropic’s cost documentation.
“You’ve hit your session limit” or “You’ve hit your weekly limit” is a seat-based window shared across every model on the account. It resets on a rolling schedule the message itself states, and switching model with /model will not restore access, because the limit isn’t attached to the model at all: it’s attached to the seat. Full mechanics, including what actually consumes that shared window fastest, are in our usage limits guide.
The second dial: effort and thinking budget
Model choice controls capability. Effort controls how hard that model works on a given turn, and it’s the dial worth reaching for before jumping to a bigger model.
/effort
/effort high
/effort auto
/effort opens an interactive slider, or set a level directly: low for short latency-sensitive turns, medium when you want lower token spend on routine work, high as the balanced default, and xhigh or max for turns that genuinely need deeper reasoning, with diminishing returns past that point. --effort at startup and CLAUDE_CODE_EFFORT_LEVEL as an environment variable set the same thing outside the interactive picker.
Thinking budget is the underlying token spend effort is drawing from, and it bills as ordinary output tokens, so a high effort level or a large fixed budget isn’t free just because it’s “just reasoning.” You can turn thinking off or cap it directly:
export MAX_THINKING_TOKENS=0
export MAX_THINKING_TOKENS=8000
The catch: Claude Code’s newer models use adaptive reasoning and ignore a nonzero MAX_THINKING_TOKENS budget entirely, per the model configuration reference. On those models, effort level is the actual control; the fixed-budget env var only does something on older models that still use a fixed thinking budget. Check which behavior your current model has before assuming a MAX_THINKING_TOKENS change did anything.
The shiploop angle
shiploop is a dispatch harness that routes every ticket to a model tier automatically instead of a person deciding by hand each time, escalating exactly once per ticket on genuine failure and holding a repeat failure at the floor tier rather than re-buying the ceiling. The floor is GOVERN_WORKER_MODEL (default sonnet), the escalation target is GOVERN_WORKER_ESCALATION_MODEL (default opus), and a ticket’s own Model: frontmatter field overrides both when a human has already made the call. A model ceiling rail (GOVERN_MODEL_CEILING, on by default) also caps what any dispatched worker can buy at the spawning session’s own model family, so a session can’t accidentally authorize spend above what a person is actually running at.
For more on how that dispatch works end to end, see our subagents guide, and for the token-level mechanics that make cheap-tier attempts cheap in the first place, see our token usage guide.
Last updated 2026-09-04.