Skip to content

A09: Extended Thinking(扩展思考) ​

相关章节:第 8 章 子 Agent——用于复杂工作的专业 Agent、第 13 章 Agent SDK

Extended Thinking(扩展思考)是 Claude 在产生可见回复之前,通过专门的"思考"阶段逐步推理问题的能力。对于复杂任务——架构决策、多步调试、精细重构——思考可以显著提升结果质量。但它在延迟、成本和行为方面存在需要理解的权衡。

什么是 Extended Thinking ​

思考 Token 流 ​

当启用 Extended Thinking 时,Claude 的回复包含两部分:

┌──────────────────────────────────────────────┐
│  Thinking tokens (internal reasoning)         │
│  - Visible to developers via API              │
│  - NOT visible to the end user by default     │
│  - Billed as output tokens                    │
│  - Cannot be interrupted                      │
│                                               │
│  "Let me analyze the dependency graph...      │
│   Module A imports from B and C.              │
│   B also imports from C, so there's a         │
│   shared dependency. The circular import      │
│   must be between A and B because..."         │
├──────────────────────────────────────────────┤
│  Response tokens (visible output)             │
│  - What the user sees                         │
│  - Informed by the thinking phase             │
│                                               │
│  "The circular dependency is between          │
│   src/auth/middleware.ts and                   │
│   src/auth/session.ts. Here's the fix..."     │
└──────────────────────────────────────────────┘

思考 Token 让 Claude 在给出答案之前先"推演"问题。这类似于程序员在写代码之前先在白板上画草图——草图不是交付物,但它提升了交付物的质量。

内部工作原理 ​

在没有 Extended Thinking 的情况下,Claude 必须从左到右、逐个 Token 地生成回复,没有回溯或重新考虑的机会。回复的第一个词会约束后续的所有内容。

启用 Extended Thinking 后,Claude 获得专门的空间来:

  1. 分解问题为子问题
  2. 探索多种方案后再做选择
  3. 评估各方案之间的权衡
  4. 自我纠正,发现自身推理中的错误
  5. 规划回复的结构

这在需要多步推理的任务上产生可衡量的更好结果,但对简单直接的任务没有附加价值。

何时使用 Extended Thinking ​

复杂推理任务 ​

当任务需要同时考虑多个约束时,Extended Thinking 表现出色:

Prompt (thinking is valuable):
"Refactor the payment processing module to support multiple 
payment providers (Stripe, PayPal, Square) without breaking 
the existing Stripe integration. The module handles 
subscriptions, one-time payments, and refunds. Each provider 
has different API patterns for these operations."

Why thinking helps:
- Multiple constraints (don't break existing, support 3 providers)
- Multiple operations (subscriptions, one-time, refunds)
- Design patterns to evaluate (strategy, adapter, factory)
- Backward compatibility to maintain

多步规划 ​

当 Claude 需要确定操作顺序时:

Prompt (thinking is valuable):
"Migrate our database from MySQL to PostgreSQL. We have 47 
tables, custom stored procedures, and a read replica. Plan 
the migration and execute the schema conversion."

Why thinking helps:
- Dependency ordering (foreign keys, views, stored procs)
- Risk assessment (what can break, rollback strategy)
- Platform-specific syntax differences
- Sequencing decisions (schema first, then data, then procs)

架构决策 ​

当需要评估不同方案之间的权衡时:

Prompt (thinking is valuable):
"We need to add real-time notifications. Our stack is 
Next.js + Express + PostgreSQL. Evaluate whether we should 
use WebSockets, Server-Sent Events, or a third-party service 
like Pusher. Consider our current infrastructure."

Why thinking helps:
- Multiple options to compare
- Infrastructure constraints to evaluate
- Cost/complexity trade-offs
- Long-term maintenance implications

调试复杂问题 ​

当根本原因不明显时:

Prompt (thinking is valuable):
"Users intermittently get 500 errors on the /api/checkout 
endpoint. The logs show 'connection reset by peer' but only 
during peak hours. The error rate is ~2% of requests. 
Investigate."

Why thinking helps:
- Multiple possible causes to consider
- Need to reason about concurrency, load, timing
- Requires correlating symptoms with infrastructure

何时不使用 Extended Thinking ​

简单直接的任务 ​

Prompt (thinking adds no value):
"Add a created_at timestamp column to the users table migration."

Without thinking: Claude writes the migration correctly.
With thinking: Claude spends 500 tokens "thinking" about it,
  then writes the same migration. You paid 5x for the thinking
  tokens with no quality improvement.

高吞吐量场景 ​

当你并行运行大量请求时(CI/CD 审查、批处理),思考会为每个请求增加延迟:

Without thinking: 50 PR reviews × 3 seconds each = 150 seconds
With thinking:    50 PR reviews × 8 seconds each = 400 seconds

Cost difference: 50 reviews × ~2,000 extra thinking tokens = 
  ~100K tokens × $15/1M = $1.50 extra (Sonnet) for no benefit

格式化和机械性任务 ​

Tasks where thinking is wasted:
- "Convert this YAML to JSON"
- "Rename the variable 'x' to 'userCount'"
- "Add semicolons to these lines"
- "Sort these imports alphabetically"

当上下文中已包含答案时 ​

如果答案直接在提供的上下文中(例如,Claude 刚读取的文件),思考只会增加延迟而不会提高准确性。

思考预算与成本影响 ​

预算的工作方式 ​

Extended Thinking 有一个以 Token 为单位的可配置预算。预算设定了每次回复中 Claude 可以使用的最大思考 Token 数:

预算设置最大思考 Token典型使用场景
Low / minimal~1,024轻量推理,简单决策
Medium~8,000标准开发任务
High~32,000复杂架构设计,深度调试
Maximum~128,000+极其复杂的多面性问题

Claude 不会总是用完全部预算。如果问题简单,即使预算很高,它可能只思考几百个 Token。预算是上限,不是目标。

成本影响 ​

思考 Token 作为输出 Token 按标准输出费率计费:

Example: Complex refactoring with Sonnet

Without thinking:
  Input:  20,000 tokens × $3.00/1M  = $0.060
  Output:  3,000 tokens × $15.00/1M = $0.045
  Total: $0.105

With thinking (medium budget, 5,000 thinking tokens used):
  Input:  20,000 tokens × $3.00/1M  = $0.060
  Output:  8,000 tokens × $15.00/1M = $0.120  (3K response + 5K thinking)
  Total: $0.180  (71% more expensive)

With thinking (high budget, 20,000 thinking tokens used):
  Input:  20,000 tokens × $3.00/1M  = $0.060
  Output: 23,000 tokens × $15.00/1M = $0.345  (3K response + 20K thinking)
  Total: $0.405  (286% more expensive)

成本增加可能很大。仅在任务复杂度值得时才使用高思考预算。

延迟影响 ​

思考 Token 必须在可见回复开始之前生成。这意味着用户会看到更长的暂停才出现输出:

Without thinking:
  Time to first visible token: ~0.5-1 second

With thinking (medium):
  Thinking phase: ~3-5 seconds (generating ~5,000 tokens)
  Time to first visible token: ~4-6 seconds

With thinking (high):
  Thinking phase: ~10-20 seconds (generating ~20,000 tokens)
  Time to first visible token: ~11-21 seconds

在 Claude Code 中,思考阶段会显示动画指示器,让你知道模型正在工作,而不是卡住了。

Extended Thinking 与工具调用的交互 ​

这种交互是 Extended Thinking 最重要且最不直观的方面之一。

思考-行动-思考循环 ​

在 Claude Code 中,Extended Thinking 在 Agent 循环内运行:

Turn 1:
  [Think] "I need to understand the current implementation..."
  [Act]   Call Read tool: src/auth/middleware.ts
  
Turn 2 (after tool result):
  [Think] "The middleware uses JWT validation. The bug is likely
           in the token refresh logic. Let me check the refresh
           handler..."
  [Act]   Call Read tool: src/auth/refresh.ts

Turn 3 (after tool result):
  [Think] "Found it. The refresh handler doesn't check token
           revocation. I need to add a check against the
           revocation list. Let me also check if there's an
           existing revocation service..."
  [Act]   Call Grep tool: "revocation" in src/

Turn 4 (after tool result):
  [Think] "There's a RevocationService in src/services/auth.ts.
           I'll import it and add the check. The fix should go
           in the refresh handler, right after the token
           validation step."
  [Act]   Call Edit tool: modify src/auth/refresh.ts

循环中的每个回合都可以有自己的思考阶段。思考预算是按回合应用的,而不是跨整个交互。

Claude 何时思考 vs 何时行动 ​

Claude 在需要推理下一步该做什么时使用思考。它在知道需要什么信息时行动(调用工具)。在实践中:

  • 第一个回合:Claude 通常会大量思考来规划方法
  • 中间回合:思考较短,侧重于解读工具结果
  • 最后一个回合:思考可能再次变长,因为 Claude 在综合发现结果

Extended Thinking 与子 Agent ​

当 Claude 生成子 Agent(参见第 8 章)时,每个子 Agent 都有自己的思考预算。编排 Agent 使用思考来规划委派,每个子 Agent 使用思考来完成其专业化任务:

Orchestrator:
  [Think] "This is a complex feature. I'll delegate:
           - Subagent 1: backend API implementation
           - Subagent 2: frontend UI components
           - Subagent 3: test coverage"
  [Act]   Spawn subagents with specific instructions

Subagent 1 (backend):
  [Think] "I need to create the API endpoint, the service
           layer, and the database migration..."
  [Act]   Read/Write/Edit tools

Subagent 2 (frontend):
  [Think] "I need to create the React components and
           connect them to the API..."
  [Act]   Read/Write/Edit tools

在多 Agent 场景中,总思考 Token 可能很可观。请相应地监控成本。

实际示例:有思考 vs 无思考 ​

示例 1:发现竞态条件(Race Condition) ​

无思考(Claude 可能遗漏问题):

User: "Users occasionally see stale data after updating their profile.
       The issue is intermittent and happens more under load."

Claude (no thinking): "I'll look at the profile update endpoint."
→ Reads the endpoint, sees a cache invalidation call, 
  suggests adding a cache TTL. Misses the actual race condition 
  between the write and the cache invalidation.

有思考(Claude 推理并发性问题):

Claude (thinking): "Intermittent + under load suggests a race 
  condition. The sequence is likely:
  1. Request A writes to DB
  2. Request B reads from cache (stale)
  3. Request A invalidates cache (too late for B)
  
  I should check if the cache invalidation is synchronous 
  and whether there's a read-after-write consistency issue..."

→ Identifies the actual race condition and proposes 
  a cache-aside pattern with write-through semantics.

示例 2:API 设计决策 ​

无思考(直接采用第一个方案,可能不是最优的):

User: "Design an API for our notification system that supports 
       email, SMS, push, and in-app notifications."

Claude (no thinking): Immediately starts writing a REST API 
  with /notifications endpoint. Produces a workable but 
  tightly-coupled design.

有思考(评估多种方案):

Claude (thinking): "Let me consider the options:
  - Single endpoint with type field: simple but mixes concerns
  - Separate endpoints per channel: explicit but lots of duplication
  - Event-driven with channel adapters: cleanest separation
  
  The event-driven approach is best because:
  - Channels have different delivery semantics
  - Users want per-channel preferences
  - We'll likely add channels in the future..."

→ Produces an event-driven design with channel adapters,
  preference management, and retry logic per channel.

在 Claude Code 和 Agent SDK 中的配置 ​

Claude Code ​

在 Claude Code 中,Extended Thinking 由模型的能力控制。使用支持思考的模型时,Claude Code 会自动启用思考。你可以通过提示措辞影响思考行为:

# Encourage deeper thinking:
"Think carefully about the trade-offs before implementing."
"Consider at least 3 approaches before picking one."
"Analyze the performance implications before making changes."

# Discourage unnecessary thinking:
"This is a straightforward change — just do it."
"Quick fix: add the missing null check in src/utils.ts line 42."

Agent SDK(Python) ​

python
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=16000,
    thinking={
        "type": "enabled",
        "budget_tokens": 10000  # max thinking tokens
    },
    messages=[{
        "role": "user",
        "content": "Design a caching strategy for our microservices..."
    }]
)

# Access thinking and response separately
for block in response.content:
    if block.type == "thinking":
        print(f"Thinking ({len(block.thinking)} chars):")
        print(block.thinking[:200] + "...")
    elif block.type == "text":
        print(f"\nResponse:")
        print(block.text)

Agent SDK(TypeScript) ​

typescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const response = await client.messages.create({
  model: "claude-sonnet-4-20250514",
  max_tokens: 16000,
  thinking: {
    type: "enabled",
    budget_tokens: 10000,
  },
  messages: [{
    role: "user",
    content: "Design a caching strategy for our microservices...",
  }],
});

// Access thinking and response separately
for (const block of response.content) {
  if (block.type === "thinking") {
    console.log(`Thinking (${block.thinking.length} chars):`);
    console.log(block.thinking.slice(0, 200) + "...");
  } else if (block.type === "text") {
    console.log(`\nResponse:`);
    console.log(block.text);
  }
}

流式传输与 Extended Thinking ​

流式传输时,思考 Token 在回复 Token 之前到达。你可以在生成思考 Token 时显示"thinking..."指示器:

python
with client.messages.stream(
    model="claude-sonnet-4-20250514",
    max_tokens=16000,
    thinking={"type": "enabled", "budget_tokens": 10000},
    messages=[{"role": "user", "content": "..."}]
) as stream:
    current_type = None
    for event in stream:
        if hasattr(event, "type"):
            if event.type == "content_block_start":
                block = event.content_block
                if block.type == "thinking":
                    print("[Thinking...]", end="", flush=True)
                    current_type = "thinking"
                elif block.type == "text":
                    print("\n[Response]")
                    current_type = "text"
            elif event.type == "content_block_delta":
                if current_type == "text":
                    print(event.delta.text, end="", flush=True)

决策框架 ​

使用以下框架来决定是否启用 Extended Thinking:

Is the task mechanically simple?
  (rename, format, simple edit)
  → YES: Skip thinking. No benefit, just cost.
  → NO: Continue...

Does the task require evaluating trade-offs?
  (architecture, design, multiple valid approaches)
  → YES: Enable thinking (medium-high budget).
  → NO: Continue...

Does the task require multi-step reasoning?
  (debugging intermittent issues, complex refactoring)
  → YES: Enable thinking (medium budget).
  → NO: Continue...

Is this a high-throughput scenario?
  (CI/CD pipeline, batch processing)
  → YES: Disable thinking. Latency and cost matter more.
  → NO: Enable thinking (low-medium budget).

核心要点 ​

  1. Extended Thinking 让 Claude 在回复前进行推理,在复杂任务上产生更好的结果,但代价是更高的延迟和 Token 用量。
  2. 将思考用于架构决策、复杂调试和多步规划。对于简单编辑、格式化和高吞吐量工作流则跳过。
  3. 思考 Token 作为输出 Token 计费(较贵的那种)。高思考预算可能使请求成本增加两倍。
  4. 在 Claude Code 的 Agent 循环中,思考按回合发生,而不是一次性的。每个工具调用周期都可以包含自己的思考阶段。
  5. 预算是上限,不是目标。Claude 只使用需要的量。设置合理的最大值,让模型自行调节。
  6. 提示措辞影响思考深度。"Think carefully about trade-offs" 鼓励更深入的推理;"Quick fix" 抑制不必要的深思。

参见:第 8 章 子 Agent 了解思考在多 Agent 场景中的工作方式,以及 A07 Token 经济学 了解成本管理策略。

基于 MIT 许可发布