A11: Model Selection Guide
Related chapters: Ch1 Installation & First Session, Ch15 Production Workflow Patterns
Choosing the right Claude model for a given task is one of the highest-leverage decisions you can make as a Claude Code user. The Claude model family spans a wide range of capability, speed, and cost trade-offs. Using Opus to rename a variable is like hiring a senior architect to move a desk; using Haiku for a complex multi-file refactor will produce mediocre results and waste your time on rework. This appendix provides the decision frameworks you need to make that choice confidently.
The Claude Model Family (2025-2026)
Claude currently ships in three tiers. Each tier represents a different balance of intelligence, speed, and cost.
Intelligence ──────────────────────────────────────────────── ►
┌────────────┐ ┌────────────┐ ┌────────────────────┐
│ Haiku 4.5 │ │ Sonnet 4.6 │ │ Opus 4.7 │
│ │ │ │ │ │
│ Fastest │ │ Balanced │ │ Most capable │
│ Cheapest │ │ Mid-tier │ │ Deepest reasoning │
│ 200k ctx │ │ 200k ctx │ │ 1M context │
└────────────┘ └────────────┘ └────────────────────┘
◄ ────────────────────────────────────────────────── Speed/CostOpus 4.7 (claude-opus-4-7-20250219)
Characteristics:
- Most powerful reasoning and code generation in the Claude family
- 1,000,000 token context window (1M) -- 5x larger than Sonnet/Haiku
- Extended thinking capable: can use structured internal reasoning on hard problems
- Best-in-class performance on complex multi-step tasks
- Strongest instruction following and nuance comprehension
Best use cases:
- Complex architectural refactors spanning 20+ files
- Debugging subtle concurrency or state-management bugs
- Security audits and vulnerability analysis
- Designing new system architectures from requirements
- Code review where correctness is critical (payments, auth, data integrity)
- Tasks requiring deep understanding of large codebases
- Extended reasoning tasks (algorithm design, performance optimization)
Cost profile (approximate, as of early 2026):
| Metric | Value |
|---|---|
| Input tokens | $15 / 1M tokens |
| Output tokens | $75 / 1M tokens |
| Cache read | $1.50 / 1M tokens |
| Cache write | $18.75 / 1M tokens |
| Context window | 1,000,000 tokens |
When NOT to use Opus: Simple file edits, boilerplate generation, formatting changes, running tests and reporting results. These tasks waste Opus's capacity (and your budget).
Sonnet 4.6 (claude-sonnet-4-6-20250514)
Characteristics:
- Strong balance of capability and speed
- 200,000 token context window
- Excellent code generation for standard patterns
- Good instruction following with fewer edge-case failures than Haiku
- Supports extended thinking (though benefits less than Opus from it)
- Fast enough for interactive development workflows
Best use cases:
- Day-to-day feature implementation
- Standard refactoring (rename, extract method, move file)
- Writing tests for existing code
- Generating boilerplate and scaffolding
- Documentation writing and updates
- Most pull request reviews
- Interactive pair-programming sessions
Cost profile (approximate):
| Metric | Value |
|---|---|
| Input tokens | $3 / 1M tokens |
| Output tokens | $15 / 1M tokens |
| Cache read | $0.30 / 1M tokens |
| Cache write | $3.75 / 1M tokens |
| Context window | 200,000 tokens |
The default choice: If you are unsure which model to use, Sonnet is almost always the right starting point. It handles 80% of development tasks well.
Haiku 4.5 (claude-haiku-4-5-20241022)
Characteristics:
- Fastest response times in the Claude family
- 200,000 token context window
- Lowest cost per token
- Good at well-defined, narrow tasks
- Weaker on tasks requiring nuanced judgment or complex reasoning
- Best suited for high-volume, low-complexity operations
Best use cases:
- Triage and classification (e.g., categorizing issues, routing requests)
- Simple code transformations (formatting, import sorting)
- Generating commit messages, PR descriptions, changelogs
- Quick lookups and summaries
- Running as a sub-agent for high-volume parallel tasks
- Linting-style checks across many files
Cost profile (approximate):
| Metric | Value |
|---|---|
| Input tokens | $0.80 / 1M tokens |
| Output tokens | $4 / 1M tokens |
| Cache read | $0.08 / 1M tokens |
| Cache write | $1 / 1M tokens |
| Context window | 200,000 tokens |
When NOT to use Haiku: Anything requiring multi-step reasoning, subtle judgment calls, or complex code generation. Haiku will produce an answer, but it may be wrong in ways that cost more to debug than using a better model upfront.
Switching Models in Claude Code
You can change the model used by Claude Code in several ways:
# Set model via environment variable
export ANTHROPIC_MODEL=claude-sonnet-4-6-20250514
claude
# Set model via CLI flag
claude --model claude-opus-4-7-20250219
# Set model in settings.json (persistent)
# In .claude/settings.json:
{
"model": "claude-sonnet-4-6-20250514"
}Within an active session, you can also switch models using the /model command:
> /model claude-opus-4-7-20250219
Model changed to claude-opus-4-7-20250219Model Comparison at a Glance
| Dimension | Haiku 4.5 | Sonnet 4.6 | Opus 4.7 |
|---|---|---|---|
| Reasoning depth | Basic | Strong | Exceptional |
| Code generation | Simple patterns | Full-stack development | Architect-level |
| Instruction following | Good for clear tasks | Very good | Best-in-class |
| Speed (time to first token) | Fastest (~200ms) | Medium (~500ms) | Slowest (~1-3s) |
| Speed (tokens/sec) | ~150 tok/s | ~90 tok/s | ~40 tok/s |
| Context window | 200K | 200K | 1M |
| Extended thinking | No | Yes | Yes (strongest benefit) |
| Input cost (per 1M tokens) | $0.80 | $3 | $15 |
| Output cost (per 1M tokens) | $4 | $15 | $75 |
| Typical task cost | $0.01-0.05 | $0.05-0.30 | $0.30-2.00 |
Extended Thinking: When It Matters
Extended thinking allows Claude to use internal chain-of-thought reasoning before responding. This benefits complex tasks disproportionately:
# Enable extended thinking with a budget
claude --model claude-opus-4-7-20250219 \
-p "Design the database schema for a multi-tenant SaaS billing system.
Consider edge cases around proration, currency conversion, and audit trails."Opus with extended thinking excels at tasks where:
- Multiple constraints must be balanced simultaneously
- The solution requires considering non-obvious edge cases
- Correctness matters more than speed (security audits, data migrations)
Sonnet with extended thinking is useful for:
- Moderately complex implementation tasks
- Code review where reasoning about interactions matters
- Planning multi-step changes
Haiku does not support extended thinking. For tasks requiring it, upgrade to Sonnet or Opus.
See A09 Extended Thinking for a detailed treatment.
Model Routing Strategies
The most cost-effective teams do not pick one model and use it for everything. They route different tasks to different models. This is called model routing or model cascading.
Strategy 1: Triage with Haiku, Implement with Sonnet, Review with Opus
This is the most common multi-model workflow:
Haiku 4.5 Sonnet 4.6 Opus 4.7
┌───────────┐ ┌────────────┐ ┌──────────┐
New task ───► │ Classify │ ───────► │ Implement │ ──────► │ Review │
│ Estimate │ │ Generate │ │ Verify │
│ Route │ │ Test │ │ Approve │
└───────────┘ └────────────┘ └──────────┘
~$0.01 ~$0.15 ~$0.50
per task per task per taskExample workflow:
- Haiku reads a GitHub issue and classifies it as
bug/feature/docs, estimates complexity (S/M/L), and routes it - Sonnet implements the solution, writes tests, generates the PR
- Opus reviews the diff for correctness, security, and architectural consistency
Strategy 2: The oh-my-claudecode Pattern
The open-source oh-my-claudecode framework implements smart model routing by intercepting task descriptions and selecting the appropriate model:
# oh-my-claudecode configuration example
# ~/.config/oh-my-claudecode/routing.yaml
rules:
- match: "simple edit|rename|typo|format"
model: claude-haiku-4-5-20241022
reason: "Simple changes don't need heavy reasoning"
- match: "implement|feature|refactor|test"
model: claude-sonnet-4-6-20250514
reason: "Standard development work"
- match: "security|architecture|review|debug complex|performance"
model: claude-opus-4-7-20250219
reason: "Tasks requiring deep analysis"
- default:
model: claude-sonnet-4-6-20250514The key insight: model selection should be automatic, not manual. If developers must remember to switch models, they will default to one model and waste money or capability.
Strategy 3: Escalation on Failure
Start with a cheaper model. If it fails or produces low-quality output, escalate:
# Pseudocode for escalation pattern
attempt_with_sonnet() {
result=$(claude --model claude-sonnet-4-6-20250514 -p "$task")
if tests_pass "$result"; then
echo "$result"
else
# Escalate to Opus
claude --model claude-opus-4-7-20250219 -p "$task (previous attempt failed: $result)"
fi
}This works well because most tasks succeed with Sonnet, and you only pay the Opus premium when necessary.
Benchmarks That Matter for Coding Tasks
Generic LLM benchmarks (MMLU, HellaSwag, etc.) do not predict coding performance well. Here are the benchmarks that actually correlate with Claude Code effectiveness:
| Benchmark | What It Measures | Relevance to Claude Code |
|---|---|---|
| SWE-bench Verified | Can the model resolve real GitHub issues? | Direct correlation -- this IS the Claude Code use case |
| HumanEval+ / MBPP+ | Can the model write correct functions? | Measures raw code generation quality |
| Aider polyglot | Multi-language editing accuracy | Tests the Edit/Write workflow specifically |
| Terminal-bench | Can the model use bash tools effectively? | Tests the Bash tool-calling loop |
| GPQA Diamond | Graduate-level reasoning | Correlates with debugging complex issues |
What to look for: When comparing models, focus on SWE-bench Verified scores. A model that scores 70% on SWE-bench will resolve about 70% of real-world coding tasks correctly on the first try, while a model at 50% will require more iteration.
As of early 2026, approximate SWE-bench Verified performance:
| Model | SWE-bench Verified |
|---|---|
| Opus 4.7 | ~72% |
| Sonnet 4.6 | ~65% |
| Haiku 4.5 | ~45% |
The gap between Sonnet and Opus is smaller than the gap between Haiku and Sonnet. This is why Sonnet is the default -- it covers most tasks while being 5x cheaper than Opus.
When to Upgrade vs Downgrade Mid-Session
Signs You Should Upgrade to Opus
- Claude is going in circles, trying the same approach repeatedly
- The task involves subtle interactions between multiple systems
- You've already spent significant tokens on failed Sonnet attempts (sunk cost is real -- the rework cost exceeds the Opus premium)
- The task requires understanding a very large codebase (Opus's 1M context is decisive here)
- You need high-confidence correctness (security, financial calculations, data migrations)
Signs You Should Downgrade to Haiku
- You are running many small, independent tasks in parallel
- The task is well-defined with a clear template (e.g., "add this field to these 20 files")
- You are generating content that will be reviewed by a human anyway
- Speed matters more than perfection (rapid prototyping, exploratory coding)
The Cost of Wrong Model Choice
Scenario: Implement a feature that touches 8 files
With Haiku (wrong choice):
- 3 attempts, each partially correct 3 x $0.05 = $0.15
- Manual debugging and fixing 1 hour of your time
- Total: $0.15 + opportunity cost of 1 hour
With Sonnet (right choice):
- 1 attempt, correct $0.15
- Quick review and merge 5 minutes
- Total: $0.15 + 5 minutes
With Opus (overkill):
- 1 attempt, correct with extra polish $0.60
- Quick review and merge 5 minutes
- Total: $0.60 + 5 minutesIn this example, Sonnet and Opus both produce the right result, but Sonnet is 4x cheaper. Haiku looks cheapest per attempt but is the most expensive overall when rework is included. The cheapest model is the one that gets it right on the first try.
The 1M Context Advantage
Opus 4.7's 1,000,000 token context window is a qualitative, not just quantitative, difference. With 1M tokens, you can fit:
- ~25,000 lines of code (an entire medium-sized project)
- A full conversation history spanning hundreds of turns
- Multiple large files simultaneously for cross-reference analysis
This matters for tasks like:
- Cross-codebase refactoring: Opus can hold the entire dependency graph in context
- Long debugging sessions: No compaction needed, so no context loss
- Architecture reviews: Read every file in the project without summarizing
When you use Sonnet or Haiku (200k context), the /compact command becomes essential for long sessions. With Opus, you can often complete entire features without ever compacting.
Decision Framework Quick Reference
Is this task simple and well-defined?
├── YES ──► Does it need to run in high volume?
│ ├── YES ──► Haiku 4.5
│ └── NO ──► Sonnet 4.6
└── NO ──► Is deep reasoning or large context critical?
├── YES ──► Opus 4.7
└── NO ──► Sonnet 4.6For most developers, the default should be Sonnet, with Opus reserved for the hard problems and Haiku reserved for high-volume automation.
Key Takeaways
- Sonnet 4.6 is the default choice for 80% of development tasks. Start here unless you have a specific reason not to.
- Opus 4.7 is for hard problems: complex debugging, security review, architectural refactoring, and anything requiring the 1M context window.
- Haiku 4.5 is for volume: triage, classification, simple transformations, and sub-agent tasks where speed matters more than depth.
- Automate model routing rather than relying on developers to choose manually. Tools like oh-my-claudecode make this practical.
- The cheapest model is the one that succeeds on the first attempt. Rework from a weaker model often costs more than using a stronger model upfront.
- SWE-bench Verified is the benchmark that matters most for predicting Claude Code effectiveness.
See also: A07 Token Economics for understanding the cost mechanics behind model pricing, and A10 Prompt Caching for reducing costs regardless of which model you choose.