A09: Extended Thinking(扩展思考)
Extended Thinking(扩展思考)是 Claude 在产生可见回复之前,通过专门的"思考"阶段逐步推理问题的能力。对于复杂任务——架构决策、多步调试、精细重构——思考可以显著提升结果质量。但它在延迟、成本和行为方面存在需要理解的权衡。
什么是 Extended Thinking
思考 Token 流
当启用 Extended Thinking 时,Claude 的回复包含两部分:
┌──────────────────────────────────────────────┐
│ Thinking tokens (internal reasoning) │
│ - Visible to developers via API │
│ - NOT visible to the end user by default │
│ - Billed as output tokens │
│ - Cannot be interrupted │
│ │
│ "Let me analyze the dependency graph... │
│ Module A imports from B and C. │
│ B also imports from C, so there's a │
│ shared dependency. The circular import │
│ must be between A and B because..." │
├──────────────────────────────────────────────┤
│ Response tokens (visible output) │
│ - What the user sees │
│ - Informed by the thinking phase │
│ │
│ "The circular dependency is between │
│ src/auth/middleware.ts and │
│ src/auth/session.ts. Here's the fix..." │
└──────────────────────────────────────────────┘思考 Token 让 Claude 在给出答案之前先"推演"问题。这类似于程序员在写代码之前先在白板上画草图——草图不是交付物,但它提升了交付物的质量。
内部工作原理
在没有 Extended Thinking 的情况下,Claude 必须从左到右、逐个 Token 地生成回复,没有回溯或重新考虑的机会。回复的第一个词会约束后续的所有内容。
启用 Extended Thinking 后,Claude 获得专门的空间来:
- 分解问题为子问题
- 探索多种方案后再做选择
- 评估各方案之间的权衡
- 自我纠正,发现自身推理中的错误
- 规划回复的结构
这在需要多步推理的任务上产生可衡量的更好结果,但对简单直接的任务没有附加价值。
何时使用 Extended Thinking
复杂推理任务
当任务需要同时考虑多个约束时,Extended Thinking 表现出色:
Prompt (thinking is valuable):
"Refactor the payment processing module to support multiple
payment providers (Stripe, PayPal, Square) without breaking
the existing Stripe integration. The module handles
subscriptions, one-time payments, and refunds. Each provider
has different API patterns for these operations."
Why thinking helps:
- Multiple constraints (don't break existing, support 3 providers)
- Multiple operations (subscriptions, one-time, refunds)
- Design patterns to evaluate (strategy, adapter, factory)
- Backward compatibility to maintain多步规划
当 Claude 需要确定操作顺序时:
Prompt (thinking is valuable):
"Migrate our database from MySQL to PostgreSQL. We have 47
tables, custom stored procedures, and a read replica. Plan
the migration and execute the schema conversion."
Why thinking helps:
- Dependency ordering (foreign keys, views, stored procs)
- Risk assessment (what can break, rollback strategy)
- Platform-specific syntax differences
- Sequencing decisions (schema first, then data, then procs)架构决策
当需要评估不同方案之间的权衡时:
Prompt (thinking is valuable):
"We need to add real-time notifications. Our stack is
Next.js + Express + PostgreSQL. Evaluate whether we should
use WebSockets, Server-Sent Events, or a third-party service
like Pusher. Consider our current infrastructure."
Why thinking helps:
- Multiple options to compare
- Infrastructure constraints to evaluate
- Cost/complexity trade-offs
- Long-term maintenance implications调试复杂问题
当根本原因不明显时:
Prompt (thinking is valuable):
"Users intermittently get 500 errors on the /api/checkout
endpoint. The logs show 'connection reset by peer' but only
during peak hours. The error rate is ~2% of requests.
Investigate."
Why thinking helps:
- Multiple possible causes to consider
- Need to reason about concurrency, load, timing
- Requires correlating symptoms with infrastructure何时不使用 Extended Thinking
简单直接的任务
Prompt (thinking adds no value):
"Add a created_at timestamp column to the users table migration."
Without thinking: Claude writes the migration correctly.
With thinking: Claude spends 500 tokens "thinking" about it,
then writes the same migration. You paid 5x for the thinking
tokens with no quality improvement.高吞吐量场景
当你并行运行大量请求时(CI/CD 审查、批处理),思考会为每个请求增加延迟:
Without thinking: 50 PR reviews × 3 seconds each = 150 seconds
With thinking: 50 PR reviews × 8 seconds each = 400 seconds
Cost difference: 50 reviews × ~2,000 extra thinking tokens =
~100K tokens × $15/1M = $1.50 extra (Sonnet) for no benefit格式化和机械性任务
Tasks where thinking is wasted:
- "Convert this YAML to JSON"
- "Rename the variable 'x' to 'userCount'"
- "Add semicolons to these lines"
- "Sort these imports alphabetically"当上下文中已包含答案时
如果答案直接在提供的上下文中(例如,Claude 刚读取的文件),思考只会增加延迟而不会提高准确性。
思考预算与成本影响
预算的工作方式
Extended Thinking 有一个以 Token 为单位的可配置预算。预算设定了每次回复中 Claude 可以使用的最大思考 Token 数:
| 预算设置 | 最大思考 Token | 典型使用场景 |
|---|---|---|
| Low / minimal | ~1,024 | 轻量推理,简单决策 |
| Medium | ~8,000 | 标准开发任务 |
| High | ~32,000 | 复杂架构设计,深度调试 |
| Maximum | ~128,000+ | 极其复杂的多面性问题 |
Claude 不会总是用完全部预算。如果问题简单,即使预算很高,它可能只思考几百个 Token。预算是上限,不是目标。
成本影响
思考 Token 作为输出 Token 按标准输出费率计费:
Example: Complex refactoring with Sonnet
Without thinking:
Input: 20,000 tokens × $3.00/1M = $0.060
Output: 3,000 tokens × $15.00/1M = $0.045
Total: $0.105
With thinking (medium budget, 5,000 thinking tokens used):
Input: 20,000 tokens × $3.00/1M = $0.060
Output: 8,000 tokens × $15.00/1M = $0.120 (3K response + 5K thinking)
Total: $0.180 (71% more expensive)
With thinking (high budget, 20,000 thinking tokens used):
Input: 20,000 tokens × $3.00/1M = $0.060
Output: 23,000 tokens × $15.00/1M = $0.345 (3K response + 20K thinking)
Total: $0.405 (286% more expensive)成本增加可能很大。仅在任务复杂度值得时才使用高思考预算。
延迟影响
思考 Token 必须在可见回复开始之前生成。这意味着用户会看到更长的暂停才出现输出:
Without thinking:
Time to first visible token: ~0.5-1 second
With thinking (medium):
Thinking phase: ~3-5 seconds (generating ~5,000 tokens)
Time to first visible token: ~4-6 seconds
With thinking (high):
Thinking phase: ~10-20 seconds (generating ~20,000 tokens)
Time to first visible token: ~11-21 seconds在 Claude Code 中,思考阶段会显示动画指示器,让你知道模型正在工作,而不是卡住了。
Extended Thinking 与工具调用的交互
这种交互是 Extended Thinking 最重要且最不直观的方面之一。
思考-行动-思考循环
在 Claude Code 中,Extended Thinking 在 Agent 循环内运行:
Turn 1:
[Think] "I need to understand the current implementation..."
[Act] Call Read tool: src/auth/middleware.ts
Turn 2 (after tool result):
[Think] "The middleware uses JWT validation. The bug is likely
in the token refresh logic. Let me check the refresh
handler..."
[Act] Call Read tool: src/auth/refresh.ts
Turn 3 (after tool result):
[Think] "Found it. The refresh handler doesn't check token
revocation. I need to add a check against the
revocation list. Let me also check if there's an
existing revocation service..."
[Act] Call Grep tool: "revocation" in src/
Turn 4 (after tool result):
[Think] "There's a RevocationService in src/services/auth.ts.
I'll import it and add the check. The fix should go
in the refresh handler, right after the token
validation step."
[Act] Call Edit tool: modify src/auth/refresh.ts循环中的每个回合都可以有自己的思考阶段。思考预算是按回合应用的,而不是跨整个交互。
Claude 何时思考 vs 何时行动
Claude 在需要推理下一步该做什么时使用思考。它在知道需要什么信息时行动(调用工具)。在实践中:
- 第一个回合:Claude 通常会大量思考来规划方法
- 中间回合:思考较短,侧重于解读工具结果
- 最后一个回合:思考可能再次变长,因为 Claude 在综合发现结果
Extended Thinking 与子 Agent
当 Claude 生成子 Agent(参见第 8 章)时,每个子 Agent 都有自己的思考预算。编排 Agent 使用思考来规划委派,每个子 Agent 使用思考来完成其专业化任务:
Orchestrator:
[Think] "This is a complex feature. I'll delegate:
- Subagent 1: backend API implementation
- Subagent 2: frontend UI components
- Subagent 3: test coverage"
[Act] Spawn subagents with specific instructions
Subagent 1 (backend):
[Think] "I need to create the API endpoint, the service
layer, and the database migration..."
[Act] Read/Write/Edit tools
Subagent 2 (frontend):
[Think] "I need to create the React components and
connect them to the API..."
[Act] Read/Write/Edit tools在多 Agent 场景中,总思考 Token 可能很可观。请相应地监控成本。
实际示例:有思考 vs 无思考
示例 1:发现竞态条件(Race Condition)
无思考(Claude 可能遗漏问题):
User: "Users occasionally see stale data after updating their profile.
The issue is intermittent and happens more under load."
Claude (no thinking): "I'll look at the profile update endpoint."
→ Reads the endpoint, sees a cache invalidation call,
suggests adding a cache TTL. Misses the actual race condition
between the write and the cache invalidation.有思考(Claude 推理并发性问题):
Claude (thinking): "Intermittent + under load suggests a race
condition. The sequence is likely:
1. Request A writes to DB
2. Request B reads from cache (stale)
3. Request A invalidates cache (too late for B)
I should check if the cache invalidation is synchronous
and whether there's a read-after-write consistency issue..."
→ Identifies the actual race condition and proposes
a cache-aside pattern with write-through semantics.示例 2:API 设计决策
无思考(直接采用第一个方案,可能不是最优的):
User: "Design an API for our notification system that supports
email, SMS, push, and in-app notifications."
Claude (no thinking): Immediately starts writing a REST API
with /notifications endpoint. Produces a workable but
tightly-coupled design.有思考(评估多种方案):
Claude (thinking): "Let me consider the options:
- Single endpoint with type field: simple but mixes concerns
- Separate endpoints per channel: explicit but lots of duplication
- Event-driven with channel adapters: cleanest separation
The event-driven approach is best because:
- Channels have different delivery semantics
- Users want per-channel preferences
- We'll likely add channels in the future..."
→ Produces an event-driven design with channel adapters,
preference management, and retry logic per channel.在 Claude Code 和 Agent SDK 中的配置
Claude Code
在 Claude Code 中,Extended Thinking 由模型的能力控制。使用支持思考的模型时,Claude Code 会自动启用思考。你可以通过提示措辞影响思考行为:
# Encourage deeper thinking:
"Think carefully about the trade-offs before implementing."
"Consider at least 3 approaches before picking one."
"Analyze the performance implications before making changes."
# Discourage unnecessary thinking:
"This is a straightforward change — just do it."
"Quick fix: add the missing null check in src/utils.ts line 42."Agent SDK(Python)
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=16000,
thinking={
"type": "enabled",
"budget_tokens": 10000 # max thinking tokens
},
messages=[{
"role": "user",
"content": "Design a caching strategy for our microservices..."
}]
)
# Access thinking and response separately
for block in response.content:
if block.type == "thinking":
print(f"Thinking ({len(block.thinking)} chars):")
print(block.thinking[:200] + "...")
elif block.type == "text":
print(f"\nResponse:")
print(block.text)Agent SDK(TypeScript)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.messages.create({
model: "claude-sonnet-4-20250514",
max_tokens: 16000,
thinking: {
type: "enabled",
budget_tokens: 10000,
},
messages: [{
role: "user",
content: "Design a caching strategy for our microservices...",
}],
});
// Access thinking and response separately
for (const block of response.content) {
if (block.type === "thinking") {
console.log(`Thinking (${block.thinking.length} chars):`);
console.log(block.thinking.slice(0, 200) + "...");
} else if (block.type === "text") {
console.log(`\nResponse:`);
console.log(block.text);
}
}流式传输与 Extended Thinking
流式传输时,思考 Token 在回复 Token 之前到达。你可以在生成思考 Token 时显示"thinking..."指示器:
with client.messages.stream(
model="claude-sonnet-4-20250514",
max_tokens=16000,
thinking={"type": "enabled", "budget_tokens": 10000},
messages=[{"role": "user", "content": "..."}]
) as stream:
current_type = None
for event in stream:
if hasattr(event, "type"):
if event.type == "content_block_start":
block = event.content_block
if block.type == "thinking":
print("[Thinking...]", end="", flush=True)
current_type = "thinking"
elif block.type == "text":
print("\n[Response]")
current_type = "text"
elif event.type == "content_block_delta":
if current_type == "text":
print(event.delta.text, end="", flush=True)决策框架
使用以下框架来决定是否启用 Extended Thinking:
Is the task mechanically simple?
(rename, format, simple edit)
→ YES: Skip thinking. No benefit, just cost.
→ NO: Continue...
Does the task require evaluating trade-offs?
(architecture, design, multiple valid approaches)
→ YES: Enable thinking (medium-high budget).
→ NO: Continue...
Does the task require multi-step reasoning?
(debugging intermittent issues, complex refactoring)
→ YES: Enable thinking (medium budget).
→ NO: Continue...
Is this a high-throughput scenario?
(CI/CD pipeline, batch processing)
→ YES: Disable thinking. Latency and cost matter more.
→ NO: Enable thinking (low-medium budget).核心要点
- Extended Thinking 让 Claude 在回复前进行推理,在复杂任务上产生更好的结果,但代价是更高的延迟和 Token 用量。
- 将思考用于架构决策、复杂调试和多步规划。对于简单编辑、格式化和高吞吐量工作流则跳过。
- 思考 Token 作为输出 Token 计费(较贵的那种)。高思考预算可能使请求成本增加两倍。
- 在 Claude Code 的 Agent 循环中,思考按回合发生,而不是一次性的。每个工具调用周期都可以包含自己的思考阶段。
- 预算是上限,不是目标。Claude 只使用需要的量。设置合理的最大值,让模型自行调节。
- 提示措辞影响思考深度。"Think carefully about trade-offs" 鼓励更深入的推理;"Quick fix" 抑制不必要的深思。
参见:第 8 章 子 Agent 了解思考在多 Agent 场景中的工作方式,以及 A07 Token 经济学 了解成本管理策略。