Skip to content

A09: Extended Thinking ​

Related chapters: Ch8 Subagents — Specialized Agents for Complex Work, Ch13 Agent SDK

Extended thinking is Claude's ability to reason through a problem step by step in a dedicated "thinking" phase before producing a visible response. For complex tasks --- architectural decisions, multi-step debugging, intricate refactoring --- thinking can dramatically improve result quality. But it comes with trade-offs in latency, cost, and behavior that you need to understand.

What Extended Thinking Is ​

The Thinking Token Stream ​

When extended thinking is enabled, Claude's response has two parts:

┌──────────────────────────────────────────────┐
│  Thinking tokens (internal reasoning)         │
│  - Visible to developers via API              │
│  - NOT visible to the end user by default     │
│  - Billed as output tokens                    │
│  - Cannot be interrupted                      │
│                                               │
│  "Let me analyze the dependency graph...      │
│   Module A imports from B and C.              │
│   B also imports from C, so there's a         │
│   shared dependency. The circular import      │
│   must be between A and B because..."         │
├──────────────────────────────────────────────┤
│  Response tokens (visible output)             │
│  - What the user sees                         │
│  - Informed by the thinking phase             │
│                                               │
│  "The circular dependency is between          │
│   src/auth/middleware.ts and                   │
│   src/auth/session.ts. Here's the fix..."     │
└──────────────────────────────────────────────┘

The thinking tokens let Claude "work through" a problem before committing to an answer. This is analogous to how a programmer might sketch on a whiteboard before writing code --- the sketch is not the deliverable, but it improves the deliverable.

How It Works Internally ​

Without extended thinking, Claude must generate its response left-to-right, one token at a time, with no opportunity to backtrack or reconsider. The first word of the response constrains everything that follows.

With extended thinking, Claude gets a dedicated space to:

  1. Decompose the problem into sub-problems
  2. Explore multiple approaches before committing
  3. Evaluate trade-offs between alternatives
  4. Self-correct by catching errors in its own reasoning
  5. Plan the structure of its response

This produces measurably better results on tasks that require multi-step reasoning, but adds no value for tasks that are straightforward.

When to Use Extended Thinking ​

Complex Reasoning Tasks ​

Extended thinking excels when the task requires holding multiple constraints in mind simultaneously:

Prompt (thinking is valuable):
"Refactor the payment processing module to support multiple 
payment providers (Stripe, PayPal, Square) without breaking 
the existing Stripe integration. The module handles 
subscriptions, one-time payments, and refunds. Each provider 
has different API patterns for these operations."

Why thinking helps:
- Multiple constraints (don't break existing, support 3 providers)
- Multiple operations (subscriptions, one-time, refunds)
- Design patterns to evaluate (strategy, adapter, factory)
- Backward compatibility to maintain

Multi-Step Planning ​

When Claude needs to determine the order of operations:

Prompt (thinking is valuable):
"Migrate our database from MySQL to PostgreSQL. We have 47 
tables, custom stored procedures, and a read replica. Plan 
the migration and execute the schema conversion."

Why thinking helps:
- Dependency ordering (foreign keys, views, stored procs)
- Risk assessment (what can break, rollback strategy)
- Platform-specific syntax differences
- Sequencing decisions (schema first, then data, then procs)

Architecture Decisions ​

When evaluating trade-offs between approaches:

Prompt (thinking is valuable):
"We need to add real-time notifications. Our stack is 
Next.js + Express + PostgreSQL. Evaluate whether we should 
use WebSockets, Server-Sent Events, or a third-party service 
like Pusher. Consider our current infrastructure."

Why thinking helps:
- Multiple options to compare
- Infrastructure constraints to evaluate
- Cost/complexity trade-offs
- Long-term maintenance implications

Debugging Complex Issues ​

When the root cause is not obvious:

Prompt (thinking is valuable):
"Users intermittently get 500 errors on the /api/checkout 
endpoint. The logs show 'connection reset by peer' but only 
during peak hours. The error rate is ~2% of requests. 
Investigate."

Why thinking helps:
- Multiple possible causes to consider
- Need to reason about concurrency, load, timing
- Requires correlating symptoms with infrastructure

When NOT to Use Extended Thinking ​

Simple, Direct Tasks ​

Prompt (thinking adds no value):
"Add a created_at timestamp column to the users table migration."

Without thinking: Claude writes the migration correctly.
With thinking: Claude spends 500 tokens "thinking" about it,
  then writes the same migration. You paid 5x for the thinking
  tokens with no quality improvement.

High-Throughput Scenarios ​

When you are running many requests in parallel (CI/CD reviews, batch processing), thinking adds latency to every request:

Without thinking: 50 PR reviews × 3 seconds each = 150 seconds
With thinking:    50 PR reviews × 8 seconds each = 400 seconds

Cost difference: 50 reviews × ~2,000 extra thinking tokens = 
  ~100K tokens × $15/1M = $1.50 extra (Sonnet) for no benefit

Formatting and Mechanical Tasks ​

Tasks where thinking is wasted:
- "Convert this YAML to JSON"
- "Rename the variable 'x' to 'userCount'"
- "Add semicolons to these lines"
- "Sort these imports alphabetically"

When Context Already Contains the Answer ​

If the answer is directly stated in the provided context (e.g., a file Claude just read), thinking adds latency without improving accuracy.

Thinking Budget and Cost Implications ​

How the Budget Works ​

Extended thinking has a configurable budget measured in tokens. The budget sets the maximum number of thinking tokens Claude can use per response:

Budget SettingMax Thinking TokensTypical Use Case
Low / minimal~1,024Light reasoning, minor decisions
Medium~8,000Standard development tasks
High~32,000Complex architecture, deep debugging
Maximum~128,000+Extremely complex multi-faceted problems

Claude does not always use the full budget. If the problem is simple, it may think for only a few hundred tokens even with a high budget. The budget is a ceiling, not a target.

Cost Impact ​

Thinking tokens are billed as output tokens at the standard output rate:

Example: Complex refactoring with Sonnet

Without thinking:
  Input:  20,000 tokens × $3.00/1M  = $0.060
  Output:  3,000 tokens × $15.00/1M = $0.045
  Total: $0.105

With thinking (medium budget, 5,000 thinking tokens used):
  Input:  20,000 tokens × $3.00/1M  = $0.060
  Output:  8,000 tokens × $15.00/1M = $0.120  (3K response + 5K thinking)
  Total: $0.180  (71% more expensive)

With thinking (high budget, 20,000 thinking tokens used):
  Input:  20,000 tokens × $3.00/1M  = $0.060
  Output: 23,000 tokens × $15.00/1M = $0.345  (3K response + 20K thinking)
  Total: $0.405  (286% more expensive)

The cost increase can be substantial. Use high thinking budgets only when the task complexity justifies it.

Latency Impact ​

Thinking tokens must be generated before the visible response begins. This means the user sees a longer pause before any output appears:

Without thinking:
  Time to first visible token: ~0.5-1 second

With thinking (medium):
  Thinking phase: ~3-5 seconds (generating ~5,000 tokens)
  Time to first visible token: ~4-6 seconds

With thinking (high):
  Thinking phase: ~10-20 seconds (generating ~20,000 tokens)
  Time to first visible token: ~11-21 seconds

In Claude Code, the thinking phase is shown as an animated indicator so you know the model is working, not stalled.

How Extended Thinking Interacts with Tool Calling ​

This interaction is one of the most important and least intuitive aspects of extended thinking.

The Think-Act-Think Cycle ​

In Claude Code, extended thinking operates within the agent loop:

Turn 1:
  [Think] "I need to understand the current implementation..."
  [Act]   Call Read tool: src/auth/middleware.ts
  
Turn 2 (after tool result):
  [Think] "The middleware uses JWT validation. The bug is likely
           in the token refresh logic. Let me check the refresh
           handler..."
  [Act]   Call Read tool: src/auth/refresh.ts

Turn 3 (after tool result):
  [Think] "Found it. The refresh handler doesn't check token
           revocation. I need to add a check against the
           revocation list. Let me also check if there's an
           existing revocation service..."
  [Act]   Call Grep tool: "revocation" in src/

Turn 4 (after tool result):
  [Think] "There's a RevocationService in src/services/auth.ts.
           I'll import it and add the check. The fix should go
           in the refresh handler, right after the token
           validation step."
  [Act]   Call Edit tool: modify src/auth/refresh.ts

Each turn in the loop can have its own thinking phase. The thinking budget applies per turn, not across the entire interaction.

When Claude Thinks vs When It Acts ​

Claude uses thinking when it needs to reason about what to do next. It acts (calls tools) when it knows what information it needs. In practice:

  • First turn: Claude often thinks extensively to plan its approach
  • Middle turns: Thinking is shorter, focused on interpreting tool results
  • Final turn: Thinking may be longer again as Claude synthesizes findings

Extended Thinking with Subagents ​

When Claude spawns subagents (see Ch8), each subagent has its own thinking budget. The orchestrator agent uses thinking to plan the delegation, and each subagent uses thinking for its specialized task:

Orchestrator:
  [Think] "This is a complex feature. I'll delegate:
           - Subagent 1: backend API implementation
           - Subagent 2: frontend UI components
           - Subagent 3: test coverage"
  [Act]   Spawn subagents with specific instructions

Subagent 1 (backend):
  [Think] "I need to create the API endpoint, the service
           layer, and the database migration..."
  [Act]   Read/Write/Edit tools

Subagent 2 (frontend):
  [Think] "I need to create the React components and
           connect them to the API..."
  [Act]   Read/Write/Edit tools

Total thinking tokens can be significant in multi-agent scenarios. Monitor costs accordingly.

Practical Examples: With vs Without Thinking ​

Example 1: Finding a Race Condition ​

Without thinking (Claude may miss the issue):

User: "Users occasionally see stale data after updating their profile.
       The issue is intermittent and happens more under load."

Claude (no thinking): "I'll look at the profile update endpoint."
→ Reads the endpoint, sees a cache invalidation call, 
  suggests adding a cache TTL. Misses the actual race condition 
  between the write and the cache invalidation.

With thinking (Claude reasons through the concurrency):

Claude (thinking): "Intermittent + under load suggests a race 
  condition. The sequence is likely:
  1. Request A writes to DB
  2. Request B reads from cache (stale)
  3. Request A invalidates cache (too late for B)
  
  I should check if the cache invalidation is synchronous 
  and whether there's a read-after-write consistency issue..."

→ Identifies the actual race condition and proposes 
  a cache-aside pattern with write-through semantics.

Example 2: API Design Decision ​

Without thinking (first approach, may not be optimal):

User: "Design an API for our notification system that supports 
       email, SMS, push, and in-app notifications."

Claude (no thinking): Immediately starts writing a REST API 
  with /notifications endpoint. Produces a workable but 
  tightly-coupled design.

With thinking (evaluates alternatives):

Claude (thinking): "Let me consider the options:
  - Single endpoint with type field: simple but mixes concerns
  - Separate endpoints per channel: explicit but lots of duplication
  - Event-driven with channel adapters: cleanest separation
  
  The event-driven approach is best because:
  - Channels have different delivery semantics
  - Users want per-channel preferences
  - We'll likely add channels in the future..."

→ Produces an event-driven design with channel adapters,
  preference management, and retry logic per channel.

Configuration in Claude Code and Agent SDK ​

Claude Code ​

In Claude Code, extended thinking is controlled by the model's capabilities. When using a thinking-capable model, Claude Code enables thinking automatically. You can influence thinking behavior through your prompts:

# Encourage deeper thinking:
"Think carefully about the trade-offs before implementing."
"Consider at least 3 approaches before picking one."
"Analyze the performance implications before making changes."

# Discourage unnecessary thinking:
"This is a straightforward change — just do it."
"Quick fix: add the missing null check in src/utils.ts line 42."

Agent SDK (Python) ​

python
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=16000,
    thinking={
        "type": "enabled",
        "budget_tokens": 10000  # max thinking tokens
    },
    messages=[{
        "role": "user",
        "content": "Design a caching strategy for our microservices..."
    }]
)

# Access thinking and response separately
for block in response.content:
    if block.type == "thinking":
        print(f"Thinking ({len(block.thinking)} chars):")
        print(block.thinking[:200] + "...")
    elif block.type == "text":
        print(f"\nResponse:")
        print(block.text)

Agent SDK (TypeScript) ​

typescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const response = await client.messages.create({
  model: "claude-sonnet-4-20250514",
  max_tokens: 16000,
  thinking: {
    type: "enabled",
    budget_tokens: 10000,
  },
  messages: [{
    role: "user",
    content: "Design a caching strategy for our microservices...",
  }],
});

// Access thinking and response separately
for (const block of response.content) {
  if (block.type === "thinking") {
    console.log(`Thinking (${block.thinking.length} chars):`);
    console.log(block.thinking.slice(0, 200) + "...");
  } else if (block.type === "text") {
    console.log(`\nResponse:`);
    console.log(block.text);
  }
}

Streaming with Extended Thinking ​

When streaming, thinking tokens arrive before response tokens. You can display a "thinking..." indicator while thinking tokens are being generated:

python
with client.messages.stream(
    model="claude-sonnet-4-20250514",
    max_tokens=16000,
    thinking={"type": "enabled", "budget_tokens": 10000},
    messages=[{"role": "user", "content": "..."}]
) as stream:
    current_type = None
    for event in stream:
        if hasattr(event, "type"):
            if event.type == "content_block_start":
                block = event.content_block
                if block.type == "thinking":
                    print("[Thinking...]", end="", flush=True)
                    current_type = "thinking"
                elif block.type == "text":
                    print("\n[Response]")
                    current_type = "text"
            elif event.type == "content_block_delta":
                if current_type == "text":
                    print(event.delta.text, end="", flush=True)

Decision Framework ​

Use this framework to decide whether to enable extended thinking:

Is the task mechanically simple?
  (rename, format, simple edit)
  → YES: Skip thinking. No benefit, just cost.
  → NO: Continue...

Does the task require evaluating trade-offs?
  (architecture, design, multiple valid approaches)
  → YES: Enable thinking (medium-high budget).
  → NO: Continue...

Does the task require multi-step reasoning?
  (debugging intermittent issues, complex refactoring)
  → YES: Enable thinking (medium budget).
  → NO: Continue...

Is this a high-throughput scenario?
  (CI/CD pipeline, batch processing)
  → YES: Disable thinking. Latency and cost matter more.
  → NO: Enable thinking (low-medium budget).

Key Takeaways ​

  1. Extended thinking lets Claude reason before responding, producing better results on complex tasks at the cost of higher latency and token usage.
  2. Use thinking for architecture decisions, complex debugging, and multi-step planning. Skip it for simple edits, formatting, and high-throughput workflows.
  3. Thinking tokens are billed as output tokens (the expensive kind). A high thinking budget can triple the cost of a request.
  4. In Claude Code's agent loop, thinking happens per turn, not once. Each tool-call cycle can include its own thinking phase.
  5. The budget is a ceiling, not a target. Claude uses only what it needs. Set a reasonable maximum and let the model self-regulate.
  6. Prompt wording influences thinking depth. "Think carefully about trade-offs" encourages deeper reasoning; "Quick fix" discourages unnecessary deliberation.

See also: Ch8 Subagents for how thinking works in multi-agent scenarios, and A07 Token Economics for cost management strategies.

Released under MIT License