A04: Agent 架构模式
相关章节: 第 8 章 Extended Thinking 与复杂工作流、第 9 章 多 Agent 编排、第 15 章 生产工作流
Claude Code 不是一个附加了额外功能的聊天机器人。它是一个自主 Agent -- 一个在循环中观察环境、推理下一步行动、执行操作并观察结果的系统。理解 AI Agent 背后的架构模式有助于你预测 Claude Code 的行为、设计与其优势对齐的任务,以及在出错时调试失败。
本附录介绍了现代 AI 系统中使用的主要 Agent 架构模式,重点是 Claude Code 如何实现和组合这些模式。这些模式来源于学术研究和 Anthropic 发布的关于构建高效 Agent 的指导。
Agent 循环:基础
每个 Agent 架构都建立在一个基本原语之上:观察-思考-行动循环。
The fundamental agentic loop:
┌──────────────────────────────────────┐
│ │
│ ┌──────────┐ │
│ │ Observe │ Read environment │
│ │ │ (tool results, │
│ │ │ user input) │
│ └────┬─────┘ │
│ │ │
│ ▼ │
│ ┌──────────┐ │
│ │ Think │ Reason about what │
│ │ │ to do next │
│ │ │ (next token │
│ │ │ prediction) │
│ └────┬─────┘ │
│ │ │
│ ▼ │
│ ┌──────────┐ │
│ │ Act │ Call a tool or │
│ │ │ respond to user │
│ └────┬─────┘ │
│ │ │
│ └────────── loop ────────────┘
│ │
└──────────────────────────────────────┘在 Claude Code 中,这个循环由 Harness(运行框架)管理。Harness 的工作是:
- 将完整上下文(系统提示、CLAUDE.md、历史记录、你的提示)发送到 Claude API
- 接收 Claude 的回复,其中可能包含工具调用
- 执行工具调用(读取文件、运行命令等)
- 将工具结果追加到上下文中
- 将更新后的上下文发回 Claude
- 重复直到 Claude 产生没有工具调用的回复("最终答案")
循环的每次迭代称为一个回合(Turn)。一条用户消息可能触发多个回合,因为 Claude 会读取文件、运行命令并迭代其方法。
模式 1:ReAct(推理 + 行动)
模式描述
ReAct(Yao 等人, 2022)是使用工具的 LLM Agent 的基础模式。它将推理轨迹与行动执行交替进行:
ReAct loop:
User: "Fix the failing test in src/auth/"
Thought: The user wants me to fix a failing test. I should first
find which test is failing by running the test suite.
Action: Bash("cd src/auth && npm test")
Observation: FAIL src/auth/__tests__/middleware.test.ts
● validateToken › should reject expired tokens
Expected: 401
Received: 200
Thought: The test expects a 401 for expired tokens but gets 200.
The issue is likely in the validateToken function. Let me
read the middleware to understand the current logic.
Action: Read("src/auth/middleware.ts")
Observation: [file contents showing validateToken function]
Thought: I see the issue. Line 24 checks `token.exp > Date.now()`
but Date.now() returns milliseconds while token.exp is in
seconds. The comparison is wrong.
Action: Edit("src/auth/middleware.ts", fix the comparison)
Observation: File edited successfully.
Thought: The fix is in place. Let me verify by running the tests again.
Action: Bash("cd src/auth && npm test")
Observation: PASS all tests
Response: "Fixed the bug. The issue was a unit mismatch..."为什么 ReAct 有效
ReAct 成功是因为它结合了两种互补的能力:
- LLM 推理处理歧义、规划和解读 -- 这些在确定性程序中难以编码。
- 工具执行提供事实真相 -- 读取实际文件而非猜测,运行实际测试而非预测结果。
推理轨迹还起到隐式思维链(Chain-of-Thought)的作用,提升后续决策的质量。每个"Thought"步骤迫使模型在行动前阐明其理解。
ReAct 在 Claude Code 中的体现
Claude Code 的核心循环就是 ReAct 实现。当你给 Claude 一个任务时,它在以下之间交替:
- 推理(工具调用之间生成的文本,包括启用时的 Extended Thinking)
- 行动(它进行的工具调用:Read、Edit、Bash 等)
- 观察(处理工具结果)
你可以在 Claude Code 的输出中直接看到这个模式。工具调用之间的文本是推理轨迹,工具调用是行动。
优势与局限
优势:
- 天然适合探索性任务(bug 调查、代码库理解)
- 自我纠正:行动的观察结果可以重新引导推理
- 透明:你可以跟随 Claude 的推理并在偏离轨道时介入
局限:
- 贪心策略:ReAct 基于局部推理决定下一步行动,而非全局计划。可能探索死胡同。
- Token 消耗大:每次推理轨迹消耗输出 Token;行动前的长思维链可能快速消耗上下文。
- 停机问题:没有外部限制的情况下,ReAct Agent 可能无限循环(Claude Code 强制执行回合限制来防止这一点)。
模式 2:计划执行(Plan-and-Execute)
模式描述
计划执行将规划与执行分离。Agent 不是在每步交替推理和行动,而是先创建一个全面的计划,然后逐步执行。
Plan-and-execute:
User: "Refactor the auth module to use the repository pattern"
Phase 1 — Plan:
1. Read current auth module structure
2. Identify data access patterns in auth code
3. Design repository interface
4. Create AuthRepository class
5. Refactor auth service to use repository
6. Update tests
7. Run tests to verify
Phase 2 — Execute:
Step 1: Read src/auth/ directory... done
Step 2: Found 3 direct database calls... done
Step 3: Designed IAuthRepository interface... done
Step 4: Created src/auth/repository.ts... done
Step 5: Refactored src/auth/service.ts... done
Step 6: Updated tests... done
Step 7: All tests pass... done计划执行在 Claude Code 中的体现
Claude Code 在以下情况下使用计划执行:
- 你要求复杂的多步任务
- 你显式请求计划("在实现前先规划")
- 启用了 Extended Thinking(思考步骤通常会产生一个隐式计划)
第 15 章的 R-P-E-R-S 工作流(Read-Plan-Execute-Review-Summarize)是一种正式化的计划执行模式:
R-P-E-R-S as plan-and-execute:
R (Read): Gather information by reading files and understanding context
P (Plan): Create a plan based on what was read
E (Execute): Implement the plan, one step at a time
R (Review): Verify the implementation by running tests
S (Summarize): Document what was done and why何时使用计划执行
在以下情况下使用计划执行(通过显式请求计划):
- 任务有许多步骤且步骤间有依赖关系
- 错误代价高昂(重构生产代码)
- 你想在 Claude 开始修改文件前审查方法
- 任务需要跨多个文件的协调变更
在以下情况下避免固定计划:
- 任务是探索性的(bug 调查)-- 你事先不知道步骤
- 任务很简单(单文件编辑)-- 规划开销浪费
- 执行过程中会出现改变计划的新信息
优势与局限
优势:
- 通过预先思考减少浪费的探索
- 给你在执行前审查和修改计划的机会
- 为长任务创建自然的检查点结构
局限:
- 计划会过时:随着执行揭示新信息,原始计划可能是错的
- 过度规划:对简单任务来说,规划开销消耗的 Token 比节省的更多
- 僵化:按计划执行的 Agent 可能忽略执行中发现的更好机会
模式 3:反思与自我纠正
模式描述
反思在 Agent 循环中添加了一个自我评估步骤。在采取行动后,Agent 显式评估结果是否满意以及下一步该做什么。
Standard loop: Observe → Think → Act → Observe → Think → Act
Reflection loop: Observe → Think → Act → Reflect → (retry or continue)反思在 Claude Code 中的体现
Claude Code 在以下情况下表现出反思行为:
- 工具调用返回错误,Claude 在重试前分析错误
- 测试失败,Claude 读取失败输出来诊断问题
- Claude 注意到其编辑没有产生预期结果并尝试不同方法
Example of reflection in action:
Action: Edit("src/auth/middleware.ts", incorrect edit)
Observation: Error: old_string not found in file
Reflection: "The file content has changed since I last read it.
Let me re-read the file to get the current content."
Action: Read("src/auth/middleware.ts")
Observation: [updated file contents]
Reflection: "I see -- the function was already refactored. My edit
target was based on the old version. Let me adjust."
Action: Edit("src/auth/middleware.ts", corrected edit)
Observation: File edited successfully.你可以通过组织提示来鼓励更显式的反思:
After each major change:
1. Re-read the modified file to verify the edit is correct
2. Run the relevant tests
3. If tests fail, analyze the failure before making more changes自我纠正的限制
Claude Code 的自我纠正有实际限制:
- 收益递减:在同一任务上失败 2-3 次后,Claude 的方法倾向于退化而非改善。它可能开始做越来越随机的更改。
- 沉没成本偏见:Claude 有时会坚持一个失败的方法,而不是退后一步尝试根本不同的方案。
- 上下文成本:每个重试循环消耗上下文。对失败的反思添加了加速上下文腐化的 Token。
实用建议:如果 Claude 在同一件事上失败了 3 次,请介入。提供提示、缩小问题范围,或带着已学到的教训开始新会话。
模式 4:多 Agent 模式
为什么需要多个 Agent?
单 Agent 系统(一个 Claude 实例处理所有事情)对聚焦任务效果很好。但它在以下方面有困难:
- 任务广度:单个上下文窗口无法容纳大型项目的所有信息
- 并行性:单个 Agent 顺序执行;多个 Agent 可以并行工作
- 专业化:不同子任务可能受益于不同的指令或模型
多 Agent 模式通过将工作分配到多个 Claude 实例来解决这些问题。
扇出模式
扇出(Fan-Out)将独立的子任务分配给并行工作的独立 Agent:
Fan-out pattern:
Orchestrator agent:
"Implement these 5 features"
├── Agent 1: "Implement user registration"
├── Agent 2: "Implement login endpoint"
├── Agent 3: "Implement password reset"
├── Agent 4: "Implement session management"
└── Agent 5: "Implement audit logging"
Each agent works independently with its own context window.
Results are collected by the orchestrator.在 Claude Code 中,扇出通过 claude --task 命令(无头模式)从父 Claude 会话中启动实现:
# Parent orchestrator launches sub-agents
claude --task "Implement user registration in src/auth/register.ts" &
claude --task "Implement login in src/auth/login.ts" &
claude --task "Implement password reset in src/auth/reset.ts" &
wait适用场景:真正独立的任务,Agent 不会编辑相同的文件。不同模块的功能实现、独立的测试套件、多服务部署。
流水线模式
流水线(Pipeline)将工作通过一系列专门的 Agent 传递:
Pipeline pattern:
Agent 1 (Planner):
Input: "Add authentication to the API"
Output: Detailed implementation plan
Agent 2 (Implementer):
Input: Plan from Agent 1
Output: Code changes
Agent 3 (Reviewer):
Input: Code changes from Agent 2
Output: Review findings
Agent 4 (Fixer):
Input: Review findings from Agent 3
Output: Fixed code流水线在不同阶段需要不同"思维模式"或你想要显式交接点让人类审查时很有用。
辩论模式
辩论(Debate)让多个 Agent 提出方案,然后为各自的方法辩论:
Debate pattern:
Agent A: "I'd solve this with a SQL migration that adds a new column"
Agent B: "I'd solve this with a separate lookup table"
Agent C (Judge): "Agent B's approach is better because it avoids
locking the main table during migration. However,
Agent A correctly identified that the column approach
is simpler for queries. Recommended: use the lookup
table with a denormalized cache column."辩论最适合存在真正权衡且没有单一"正确"答案的架构决策。实践中,这通过运行多个带有不同提示的 Claude 会话并让最终会话综合结果来实现。
共识模式
共识(Consensus)独立地将同一任务交给多个 Agent,取多数答案:
Consensus pattern (e.g., code review):
Agent 1: "I found 3 issues: X, Y, Z"
Agent 2: "I found 2 issues: X, Z"
Agent 3: "I found 4 issues: X, Y, Z, W"
Consensus: Issues X and Z are confirmed (3/3 agents).
Issue Y is likely real (2/3 agents).
Issue W needs human review (1/3 agents).这是 /review 技能的"ultra"模式的基础,它运行多次审查并综合发现。
Agent 系统中的状态管理
状态问题
Agent 需要跟踪状态:完成了什么、剩余什么、适用什么约束。在基于 LLM 的 Agent 中,状态管理具有挑战性,因为:
- 唯一的状态就是上下文窗口:Claude 在会话期间没有外部记忆。它"知道"的一切都必须在当前上下文中。
- 状态是隐式的:对话历史就是状态。没有单独的状态对象。
- 状态会退化:如 A02 上下文工程 所述,上下文腐化会降低 Agent 对自身状态的感知。
状态管理策略
策略 1:将状态外部化到文件
You: "Before starting each major step, write your current progress
to a file called PROGRESS.md. Include: completed steps, current step,
remaining steps, and any issues found."这创建了一个持久的状态产物,能在压缩中存活,并可在 Claude 失去追踪时重新读取。
策略 2:使用结构化检查点
You: "We're implementing 5 API endpoints. After each one, output a
status table:
| Endpoint | Status | Tests |
|----------|--------|-------|
| /users | Done | Pass |
| /auth | Done | Pass |
| /posts | WIP | - |
| /comments| TODO | - |
| /likes | TODO | - |"结构化格式比散文摘要更容易让 Claude 在后续轮次中解析。
策略 3:利用 git 作为状态
You: "Commit after each completed feature. Use descriptive commit
messages. If you need to understand what's been done, run git log."Git 提交创建了一个不受上下文腐化影响的外部状态记录。
失败恢复与重试策略
Agent 失败的类型
| 失败类型 | 示例 | 恢复方法 |
|---|---|---|
| 工具错误 | 文件未找到、命令失败 | 重新读取、调整路径、重试 |
| 逻辑错误 | 错误的方法、不正确的分析 | 反思、回溯、尝试替代方案 |
| 上下文溢出 | 失去对需求的追踪 | 总结、压缩或新会话 |
| 无限循环 | 重复同样失败的操作 | 需要人工介入 |
| 级联错误 | 早期错误在后续步骤中累积 | 回退到上一个已知的良好状态 |
恢复策略
调整后重试(用于临时工具错误):
Claude: [Bash] npm test
Error: ECONNREFUSED - database not running
Claude: "The database isn't running. Let me start it first."
Claude: [Bash] docker compose up -d postgres
Claude: [Bash] npm test
Result: All tests pass回溯并重新规划(用于逻辑错误):
You: "That approach isn't working. Stop, explain what went wrong,
and propose a completely different approach. Don't modify the
previous approach -- think of an alternative."检查点与回滚(用于级联错误):
You: "The last 3 changes made things worse. Run `git stash` to
undo them, then re-read the original code and try a simpler
approach."重新开始(用于上下文溢出):
You: "This session has gotten too complex. Summarize the current
state: what works, what doesn't, what's left. I'll start a new
session with your summary."何时使用单 Agent vs 多 Agent
决策框架
Should I use multiple agents?
Is the task decomposable into independent subtasks?
├── No → Single agent
└── Yes
├── Do subtasks share many files?
│ ├── Yes → Single agent (or pipeline)
│ └── No → Fan-out
├── Is the task complex enough to benefit from specialization?
│ ├── No → Single agent
│ └── Yes → Pipeline
└── Do you need high confidence in the result?
├── No → Single agent
└── Yes → Consensus实用指南
使用单 Agent 的场景:
- 任务涉及少于 10 个文件
- 变更是相互依赖的(编辑一个函数及其调用者)
- 任务可以舒适地放入一个上下文窗口
- 速度比彻底性更重要
使用多 Agent 的场景:
- 任务跨越多个模块或服务
- 子任务真正独立(没有共享的文件编辑)
- 你需要并行执行以提高速度
- 你想要多个视角(审查、架构决策)
- 任务超出单个上下文窗口的容量
开销权衡
多 Agent 系统增加了开销:
- 协调成本:编排 Agent 需要提示、上下文和逻辑
- 集成风险:合并多个 Agent 的工作可能产生冲突
- 调试复杂性:在多个 Agent 会话中追踪 bug 更困难
- Token 成本:每个 Agent 有自己的上下文窗口,总 Token 使用量增加
来自 Anthropic 的"Building effective agents"博客文章 (2024) 的经验法则是:从单 Agent 开始,只在遇到具体局限性时才添加 Agent。不要因为听起来高级就用多 Agent。在单 Agent 明显无法处理任务时才使用。
参考:Anthropic 的 Agent 设计原则
Anthropic 发布了题为"Building effective agents"的博客文章,概述了关键原则。以下是与 Claude Code 用户最相关的:
保持 Agent 架构简单:最有效的 Agent 使用最简单的可行架构。一个带有良好工具的 ReAct 循环在大多数任务上优于复杂的多 Agent 系统。
工具比提示更适合提供事实依据:当 Agent 需要事实信息时,应该使用工具(读取文件、运行命令)而不是依赖训练数据。这就是为什么 Claude Code 的工具调用方法比基于聊天的 AI 在软件工程方面更可靠。
显式边界提高可靠性:清晰的约束("不要修改 src/auth/ 之外的文件")比开放式指令("注意不要影响其他文件")产生更好的结果。
在关键节点引入人类:最可靠的 Agent 系统在不可逆操作(部署、数据库迁移、公共 API 变更)前包含人类检查点。Claude Code 的权限系统就是这一原则的实现。
快速失败、清楚恢复:Agent 应该快速检测失败并清楚地报告,而不是尝试越来越绝望的修复。这就是为什么你应该在 Claude 在同一件事上失败 3 次后介入。
核心要点
- Claude Code 的核心是一个 ReAct Agent。 理解观察-思考-行动循环有助于你预测其行为并设计更好的提示。
- 计划执行在复杂多步任务中表现突出。 当任务有许多相互依赖的步骤时,显式请求计划。
- 反思实现自我纠正,但收益递减。在 2-3 次失败尝试后介入,而不是期望最终成功。
- 多 Agent 是工具,不是默认选择。 从单 Agent 开始。只在遇到具体局限性时才添加 Agent。
- 状态管理是 Agent 系统中最难的问题。 将状态外部化到文件或 git 提交中,而不是仅依赖对话历史。
- 失败恢复需要将策略与失败类型匹配。 临时错误用重试,逻辑错误用回溯,上下文溢出用重新开始。
另见:A05 工具调用内部机制 了解 Agent 循环中行动步骤的实际工作方式,以及 A13 多 Agent 协调 了解高级多 Agent 模式。