Skip to content

A11: Model Selection Guide ​

Related chapters: Ch1 Installation & First Session, Ch15 Production Workflow Patterns

Choosing the right Claude model for a given task is one of the highest-leverage decisions you can make as a Claude Code user. The Claude model family spans a wide range of capability, speed, and cost trade-offs. Using Opus to rename a variable is like hiring a senior architect to move a desk; using Haiku for a complex multi-file refactor will produce mediocre results and waste your time on rework. This appendix provides the decision frameworks you need to make that choice confidently.

The Claude Model Family (2025-2026) ​

Claude currently ships in three tiers. Each tier represents a different balance of intelligence, speed, and cost.

Intelligence ──────────────────────────────────────────────── ►
                                                              
  ┌────────────┐     ┌────────────┐     ┌────────────────────┐
  │  Haiku 4.5 │     │ Sonnet 4.6 │     │    Opus 4.7        │
  │            │     │            │     │                    │
  │  Fastest   │     │  Balanced  │     │  Most capable      │
  │  Cheapest  │     │  Mid-tier  │     │  Deepest reasoning │
  │  200k ctx  │     │  200k ctx  │     │  1M context        │
  └────────────┘     └────────────┘     └────────────────────┘
                                                              
◄ ────────────────────────────────────────────────── Speed/Cost

Opus 4.7 (claude-opus-4-7-20250219) ​

Characteristics:

  • Most powerful reasoning and code generation in the Claude family
  • 1,000,000 token context window (1M) -- 5x larger than Sonnet/Haiku
  • Extended thinking capable: can use structured internal reasoning on hard problems
  • Best-in-class performance on complex multi-step tasks
  • Strongest instruction following and nuance comprehension

Best use cases:

  • Complex architectural refactors spanning 20+ files
  • Debugging subtle concurrency or state-management bugs
  • Security audits and vulnerability analysis
  • Designing new system architectures from requirements
  • Code review where correctness is critical (payments, auth, data integrity)
  • Tasks requiring deep understanding of large codebases
  • Extended reasoning tasks (algorithm design, performance optimization)

Cost profile (approximate, as of early 2026):

MetricValue
Input tokens$15 / 1M tokens
Output tokens$75 / 1M tokens
Cache read$1.50 / 1M tokens
Cache write$18.75 / 1M tokens
Context window1,000,000 tokens

When NOT to use Opus: Simple file edits, boilerplate generation, formatting changes, running tests and reporting results. These tasks waste Opus's capacity (and your budget).

Sonnet 4.6 (claude-sonnet-4-6-20250514) ​

Characteristics:

  • Strong balance of capability and speed
  • 200,000 token context window
  • Excellent code generation for standard patterns
  • Good instruction following with fewer edge-case failures than Haiku
  • Supports extended thinking (though benefits less than Opus from it)
  • Fast enough for interactive development workflows

Best use cases:

  • Day-to-day feature implementation
  • Standard refactoring (rename, extract method, move file)
  • Writing tests for existing code
  • Generating boilerplate and scaffolding
  • Documentation writing and updates
  • Most pull request reviews
  • Interactive pair-programming sessions

Cost profile (approximate):

MetricValue
Input tokens$3 / 1M tokens
Output tokens$15 / 1M tokens
Cache read$0.30 / 1M tokens
Cache write$3.75 / 1M tokens
Context window200,000 tokens

The default choice: If you are unsure which model to use, Sonnet is almost always the right starting point. It handles 80% of development tasks well.

Haiku 4.5 (claude-haiku-4-5-20241022) ​

Characteristics:

  • Fastest response times in the Claude family
  • 200,000 token context window
  • Lowest cost per token
  • Good at well-defined, narrow tasks
  • Weaker on tasks requiring nuanced judgment or complex reasoning
  • Best suited for high-volume, low-complexity operations

Best use cases:

  • Triage and classification (e.g., categorizing issues, routing requests)
  • Simple code transformations (formatting, import sorting)
  • Generating commit messages, PR descriptions, changelogs
  • Quick lookups and summaries
  • Running as a sub-agent for high-volume parallel tasks
  • Linting-style checks across many files

Cost profile (approximate):

MetricValue
Input tokens$0.80 / 1M tokens
Output tokens$4 / 1M tokens
Cache read$0.08 / 1M tokens
Cache write$1 / 1M tokens
Context window200,000 tokens

When NOT to use Haiku: Anything requiring multi-step reasoning, subtle judgment calls, or complex code generation. Haiku will produce an answer, but it may be wrong in ways that cost more to debug than using a better model upfront.

Switching Models in Claude Code ​

You can change the model used by Claude Code in several ways:

bash
# Set model via environment variable
export ANTHROPIC_MODEL=claude-sonnet-4-6-20250514
claude

# Set model via CLI flag
claude --model claude-opus-4-7-20250219

# Set model in settings.json (persistent)
# In .claude/settings.json:
{
  "model": "claude-sonnet-4-6-20250514"
}

Within an active session, you can also switch models using the /model command:

> /model claude-opus-4-7-20250219
Model changed to claude-opus-4-7-20250219

Model Comparison at a Glance ​

DimensionHaiku 4.5Sonnet 4.6Opus 4.7
Reasoning depthBasicStrongExceptional
Code generationSimple patternsFull-stack developmentArchitect-level
Instruction followingGood for clear tasksVery goodBest-in-class
Speed (time to first token)Fastest (~200ms)Medium (~500ms)Slowest (~1-3s)
Speed (tokens/sec)~150 tok/s~90 tok/s~40 tok/s
Context window200K200K1M
Extended thinkingNoYesYes (strongest benefit)
Input cost (per 1M tokens)$0.80$3$15
Output cost (per 1M tokens)$4$15$75
Typical task cost$0.01-0.05$0.05-0.30$0.30-2.00

Extended Thinking: When It Matters ​

Extended thinking allows Claude to use internal chain-of-thought reasoning before responding. This benefits complex tasks disproportionately:

bash
# Enable extended thinking with a budget
claude --model claude-opus-4-7-20250219 \
  -p "Design the database schema for a multi-tenant SaaS billing system. 
  Consider edge cases around proration, currency conversion, and audit trails."

Opus with extended thinking excels at tasks where:

  • Multiple constraints must be balanced simultaneously
  • The solution requires considering non-obvious edge cases
  • Correctness matters more than speed (security audits, data migrations)

Sonnet with extended thinking is useful for:

  • Moderately complex implementation tasks
  • Code review where reasoning about interactions matters
  • Planning multi-step changes

Haiku does not support extended thinking. For tasks requiring it, upgrade to Sonnet or Opus.

See A09 Extended Thinking for a detailed treatment.

Model Routing Strategies ​

The most cost-effective teams do not pick one model and use it for everything. They route different tasks to different models. This is called model routing or model cascading.

Strategy 1: Triage with Haiku, Implement with Sonnet, Review with Opus ​

This is the most common multi-model workflow:

                  Haiku 4.5              Sonnet 4.6             Opus 4.7
                ┌───────────┐          ┌────────────┐         ┌──────────┐
  New task ───► │  Classify  │ ───────► │ Implement  │ ──────► │  Review  │
                │  Estimate  │          │  Generate   │         │  Verify  │
                │  Route     │          │  Test       │         │  Approve │
                └───────────┘          └────────────┘         └──────────┘
                   ~$0.01                  ~$0.15                ~$0.50
                  per task               per task              per task

Example workflow:

  1. Haiku reads a GitHub issue and classifies it as bug/feature/docs, estimates complexity (S/M/L), and routes it
  2. Sonnet implements the solution, writes tests, generates the PR
  3. Opus reviews the diff for correctness, security, and architectural consistency

Strategy 2: The oh-my-claudecode Pattern ​

The open-source oh-my-claudecode framework implements smart model routing by intercepting task descriptions and selecting the appropriate model:

bash
# oh-my-claudecode configuration example
# ~/.config/oh-my-claudecode/routing.yaml
rules:
  - match: "simple edit|rename|typo|format"
    model: claude-haiku-4-5-20241022
    reason: "Simple changes don't need heavy reasoning"

  - match: "implement|feature|refactor|test"
    model: claude-sonnet-4-6-20250514
    reason: "Standard development work"

  - match: "security|architecture|review|debug complex|performance"
    model: claude-opus-4-7-20250219
    reason: "Tasks requiring deep analysis"

  - default:
    model: claude-sonnet-4-6-20250514

The key insight: model selection should be automatic, not manual. If developers must remember to switch models, they will default to one model and waste money or capability.

Strategy 3: Escalation on Failure ​

Start with a cheaper model. If it fails or produces low-quality output, escalate:

bash
# Pseudocode for escalation pattern
attempt_with_sonnet() {
  result=$(claude --model claude-sonnet-4-6-20250514 -p "$task")
  if tests_pass "$result"; then
    echo "$result"
  else
    # Escalate to Opus
    claude --model claude-opus-4-7-20250219 -p "$task (previous attempt failed: $result)"
  fi
}

This works well because most tasks succeed with Sonnet, and you only pay the Opus premium when necessary.

Benchmarks That Matter for Coding Tasks ​

Generic LLM benchmarks (MMLU, HellaSwag, etc.) do not predict coding performance well. Here are the benchmarks that actually correlate with Claude Code effectiveness:

BenchmarkWhat It MeasuresRelevance to Claude Code
SWE-bench VerifiedCan the model resolve real GitHub issues?Direct correlation -- this IS the Claude Code use case
HumanEval+ / MBPP+Can the model write correct functions?Measures raw code generation quality
Aider polyglotMulti-language editing accuracyTests the Edit/Write workflow specifically
Terminal-benchCan the model use bash tools effectively?Tests the Bash tool-calling loop
GPQA DiamondGraduate-level reasoningCorrelates with debugging complex issues

What to look for: When comparing models, focus on SWE-bench Verified scores. A model that scores 70% on SWE-bench will resolve about 70% of real-world coding tasks correctly on the first try, while a model at 50% will require more iteration.

As of early 2026, approximate SWE-bench Verified performance:

ModelSWE-bench Verified
Opus 4.7~72%
Sonnet 4.6~65%
Haiku 4.5~45%

The gap between Sonnet and Opus is smaller than the gap between Haiku and Sonnet. This is why Sonnet is the default -- it covers most tasks while being 5x cheaper than Opus.

When to Upgrade vs Downgrade Mid-Session ​

Signs You Should Upgrade to Opus ​

  • Claude is going in circles, trying the same approach repeatedly
  • The task involves subtle interactions between multiple systems
  • You've already spent significant tokens on failed Sonnet attempts (sunk cost is real -- the rework cost exceeds the Opus premium)
  • The task requires understanding a very large codebase (Opus's 1M context is decisive here)
  • You need high-confidence correctness (security, financial calculations, data migrations)

Signs You Should Downgrade to Haiku ​

  • You are running many small, independent tasks in parallel
  • The task is well-defined with a clear template (e.g., "add this field to these 20 files")
  • You are generating content that will be reviewed by a human anyway
  • Speed matters more than perfection (rapid prototyping, exploratory coding)

The Cost of Wrong Model Choice ​

Scenario: Implement a feature that touches 8 files

With Haiku (wrong choice):
  - 3 attempts, each partially correct       3 x $0.05 = $0.15
  - Manual debugging and fixing               1 hour of your time
  - Total: $0.15 + opportunity cost of 1 hour

With Sonnet (right choice):
  - 1 attempt, correct                       $0.15
  - Quick review and merge                    5 minutes
  - Total: $0.15 + 5 minutes

With Opus (overkill):
  - 1 attempt, correct with extra polish      $0.60
  - Quick review and merge                    5 minutes
  - Total: $0.60 + 5 minutes

In this example, Sonnet and Opus both produce the right result, but Sonnet is 4x cheaper. Haiku looks cheapest per attempt but is the most expensive overall when rework is included. The cheapest model is the one that gets it right on the first try.

The 1M Context Advantage ​

Opus 4.7's 1,000,000 token context window is a qualitative, not just quantitative, difference. With 1M tokens, you can fit:

  • ~25,000 lines of code (an entire medium-sized project)
  • A full conversation history spanning hundreds of turns
  • Multiple large files simultaneously for cross-reference analysis

This matters for tasks like:

  • Cross-codebase refactoring: Opus can hold the entire dependency graph in context
  • Long debugging sessions: No compaction needed, so no context loss
  • Architecture reviews: Read every file in the project without summarizing

When you use Sonnet or Haiku (200k context), the /compact command becomes essential for long sessions. With Opus, you can often complete entire features without ever compacting.

Decision Framework Quick Reference ​

Is this task simple and well-defined?
  ├── YES ──► Does it need to run in high volume?
  │             ├── YES ──► Haiku 4.5
  │             └── NO  ──► Sonnet 4.6
  └── NO  ──► Is deep reasoning or large context critical?
                ├── YES ──► Opus 4.7
                └── NO  ──► Sonnet 4.6

For most developers, the default should be Sonnet, with Opus reserved for the hard problems and Haiku reserved for high-volume automation.

Key Takeaways ​

  1. Sonnet 4.6 is the default choice for 80% of development tasks. Start here unless you have a specific reason not to.
  2. Opus 4.7 is for hard problems: complex debugging, security review, architectural refactoring, and anything requiring the 1M context window.
  3. Haiku 4.5 is for volume: triage, classification, simple transformations, and sub-agent tasks where speed matters more than depth.
  4. Automate model routing rather than relying on developers to choose manually. Tools like oh-my-claudecode make this practical.
  5. The cheapest model is the one that succeeds on the first attempt. Rework from a weaker model often costs more than using a stronger model upfront.
  6. SWE-bench Verified is the benchmark that matters most for predicting Claude Code effectiveness.

See also: A07 Token Economics for understanding the cost mechanics behind model pricing, and A10 Prompt Caching for reducing costs regardless of which model you choose.

Released under MIT License