Chapter 14: Managed Agents 与 Harness 架构
学习目标
- 三层虚拟化架构:Session、Harness、Sandbox
- 通过 Managed Agents API 创建和管理云端 agent
- 断连/重连的长时间运行任务(session 持久化)
- 自评估 agent 与 pass@k 指标
- 企业采用模式:Notion、Rakuten、Sentry、Atlassian
- 何时使用 CLI vs Agent SDK vs Managed Agents
为什么需要 Managed Agents
在 Managed Agents 之前,构建生产级 AI agent 需要你自己解决基础设施问题:
你想构建的: 你实际要构建的:
"重构代码的 Agent" 容器编排
状态管理
崩溃恢复
权限沙箱
上下文窗口管理
Prompt 缓存
工具执行路由
基础设施工作 >> Agent 业务逻辑Managed Agents 将所有这些移到 Anthropic 的云端。你定义 agent 该做什么,Anthropic 负责怎么运行。
研究参考:Anthropic 的"Building effective agents"博客文章(2024)确立了 Managed Agents 实现的设计原则:将 agent 逻辑与基础设施分离,优先使用简单的工具接口,让模型驱动控制流而不是硬编码到 harness 中。下文描述的三层架构是这些原则在基础设施层面的直接实现。
架构
三层虚拟化
这个设计借鉴了操作系统。每一层都是独立的——如果一个失败或需要替换,其他两个不受影响。
┌──────────────────────────────────────────────────┐
│ │
│ SESSION(事件日志) │
│ ┌─────────────────────────────────────────┐ │
│ │ 所有事件的追加写入日志 │ │
│ │ 独立于 Harness 和 Sandbox 持久存储 │ │
│ │ 随时可通过 getEvents() 查询 │ │
│ │ 即使其他一切崩溃也永远不丢 │ │
│ └─────────────────────────────────────────┘ │
│ │
│ HARNESS(编排循环) │
│ ┌─────────────────────────────────────────┐ │
│ │ 调用 Claude -> 路由工具调用 -> 循环 │ │
│ │ 无状态,可替换("牲畜"设计) │ │
│ │ 内置 prompt 缓存 + 上下文压缩 │ │
│ │ 崩溃后通过 wake(sessionId) 恢复 │ │
│ └─────────────────────────────────────────┘ │
│ │
│ SANDBOX(执行环境) │
│ ┌─────────────────────────────────────────┐ │
│ │ 暴露 execute(name, input) -> string │ │
│ │ 可以是容器、虚拟机、手机,任何执行目标 │ │
│ │ Harness 不知道也不关心具体实现 │ │
│ └─────────────────────────────────────────┘ │
│ │
└──────────────────────────────────────────────────┘"大脑与双手"分离——为什么重要
来自 Anthropic 工程博客:
Harness 编码的是模型做不到什么的假设,这些假设会过时。
具体例子:Sonnet 4.5 在接近上下文限制时会提前收工("上下文焦虑")。Anthropic 在 Harness 中加了上下文重置来应对。换到 Opus 4.5 后,这个行为消失了,重置变成了不必要的开销。
教训:将推理引擎(大脑)与执行环境(双手)解耦,这样可以独立升级任一部分。这与操作系统虚拟化是同一原理——为尚未想到的程序设计系统。
从"宠物"到"牲畜"
耦合设计(宠物): 解耦设计(牲畜):
┌─────────────────────┐ Harness -> 无状态,挂了换一个
│ 容器 │ Sandbox -> 容器,起新的
│ ├── Harness │ Session -> 持久化日志,永远不丢
│ ├── Claude │
│ ├── 工具执行 │ 宠物坏了你得救它。
│ └── Session 数据 │ 牲畜坏了换一头就行。
└─────────────────────┘
容器挂了 = 全部丢失解耦的性能提升:推理可以在容器准备好之前就开始。
- p50 TTFT(首 token 时间):下降约 60%
- p95 TTFT:下降超过 90%
四大 API 资源
| 资源 | 用途 | 生命周期 |
|---|---|---|
| Agent | 模型、系统提示、工具 | 可跨多个 session 复用 |
| Environment | 容器模板、网络规则 | 可复用 |
| Session | Agent + Environment 的运行实例 | 每个任务创建 |
| Events | 消息、状态更新、工具结果 | 通过 SSE 流式传输 |
CLI vs Agent SDK vs Managed Agents
| CLI | Agent SDK | Managed Agents | |
|---|---|---|---|
| 运行位置 | 你的终端 | 你的服务器 | Anthropic 云端 |
| 适合 | 交互式开发 | CI/CD、嵌入应用 | 异步云端任务 |
| 持续时间 | 会话级 | 取决于你 | 数小时到数天 |
| 基础设施 | 不需要 | 你管理 | Anthropic 管理 |
| 容错 | 手动 | 你实现 | 自动 |
| 定价 | 订阅制 | 按 token 计 | 按 token + $0.08/session-hour |
Demo 37: 真实代码库重构的 Managed Agent
场景
你有一个 Python 后端,其中有一个 1,500 行的 services.py 文件需要拆分为独立的 service 模块。这是一个多小时的任务,需要创建很多文件、更新导入和修复测试。这正是 Managed Agents 的用武之地——你启动它然后回来看结果。
前提
export ANTHROPIC_API_KEY="sk-ant-..."
pip install anthropic代码
# refactor_agent.py
from anthropic import Anthropic
client = Anthropic()
# Step 1: Create an agent specialized for refactoring
agent = client.beta.agents.create(
name="Python Refactoring Agent",
description="Splits monolithic Python modules into well-organized packages",
model="claude-sonnet-4-6",
system="""You are an expert Python refactoring agent. Your approach:
1. Read the target module completely before making any changes
2. Identify natural boundaries (classes, function groups, domain concepts)
3. Create the new package structure with __init__.py files
4. Move code in dependency order (leaf modules first)
5. Update all imports across the entire project
6. Run the test suite after each move to catch breakage early
7. Never change business logic -- this is a pure structural refactor
When you encounter circular imports, resolve them by:
- Extracting shared types into a types.py module
- Using TYPE_CHECKING blocks for type-only imports
- Restructuring the dependency graph
Commit after each successful module extraction with a clear message.""",
tools=[{"type": "agent_toolset_20260401"}],
)
# Step 2: Create an environment with git access
environment = client.beta.environments.create(
name="refactor-sandbox",
config={
"type": "cloud",
"networking": {"type": "unrestricted"},
},
setup_commands=[
"pip install pytest",
"git clone https://github.com/yourorg/yourproject.git /workspace",
"cd /workspace && pip install -e .",
],
)
# Step 3: Start the session
session = client.beta.sessions.create(
agent=agent.id,
environment_id=environment.id,
title="Split services.py into service package",
)
print(f"Session ID: {session.id}")
print("Save this -- you can disconnect and reconnect at any time.\n")
# Step 4: Send the task and stream initial results
with client.beta.sessions.events.stream(session.id) as stream:
client.beta.sessions.events.send(
session.id,
events=[{
"type": "user.message",
"content": [{
"type": "text",
"text": """Refactor /workspace/src/services.py into a services/ package.
The file currently contains:
- UserService (user CRUD, authentication)
- OrderService (order management, cart operations)
- PaymentService (payment processing, refunds)
- NotificationService (email, SMS, push notifications)
- AnalyticsService (event tracking, reporting)
Split into:
src/services/
__init__.py (re-exports for backward compatibility)
user_service.py
order_service.py
payment_service.py
notification_service.py
analytics_service.py
types.py (shared types/interfaces)
Requirements:
1. All existing imports like `from services import UserService` must still work
2. All tests must pass after the refactor
3. No business logic changes
4. Each module should have proper __all__ exports
5. Resolve any circular dependencies cleanly
Start by reading services.py and the test files to understand the current structure.""",
}],
}],
)
# Stream events as they come
for event in stream:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")
case "session.status_idle":
print("\n\n--- Agent finished ---")
break运行
python refactor_agent.py预期终端输出:
$ python refactor_agent.py
Session ID: sess_01JXK9M2ABC3DEF456
Save this -- you can disconnect and reconnect at any time.
Reading /workspace/src/services.py to understand the current structure...
[Tool: Read]
The file is 1,487 lines. I can see 5 service classes with clear boundaries.
Reading test files to understand the test structure...
[Tool: Glob]
[Tool: Read]
Found 3 test files: test_user.py (18 tests), test_orders.py (14 tests),
test_payments.py (9 tests). Total: 41 tests.
Starting extraction. I will work in dependency order:
1. types.py (shared types, no dependencies)
2. analytics_service.py (leaf node, no internal deps)
3. notification_service.py (leaf node)
4. payment_service.py (depends on notification)
5. order_service.py (depends on payment)
6. user_service.py (depends on notification)
7. __init__.py (re-exports everything)
Extracting shared types...
[Tool: Write]
[Tool: Edit]
Running tests after types extraction...
[Tool: Bash]
41/41 tests pass.
[Tool: Bash]
Committed: "Extract shared types into services/types.py"
Extracting AnalyticsService...
[Tool: Write]
[Tool: Edit]
[Tool: Bash]
41/41 tests pass.
[Tool: Bash]
Committed: "Extract AnalyticsService into services/analytics_service.py"
...
Creating __init__.py with backward-compatible re-exports...
[Tool: Write]
Final test run...
[Tool: Bash]
41/41 tests pass. All imports verified.
[Tool: Bash]
Committed: "Add __init__.py with re-exports for backward compatibility"
Refactoring complete. Summary:
- 1 monolithic file (1,487 lines) -> 7 focused modules
- All 41 tests pass
- All existing imports remain valid
- 7 commits, one per extraction step
--- Agent finished ---刚才发生了什么?
Demo 38: 长时间运行任务与 Session 持久化
要点
你启动一个任务,断开连接(合上笔记本、去吃午饭、随便做什么),agent 继续在云端运行。重新连接时,你从断开的地方继续——包括你不在时完成的所有工作。
启动任务
# start_task.py
from anthropic import Anthropic
client = Anthropic()
agent = client.beta.agents.create(
name="Migration Agent",
model="claude-sonnet-4-6",
system="You are a thorough migration agent. Work methodically through each file.",
tools=[{"type": "agent_toolset_20260401"}],
)
env = client.beta.environments.create(
name="migration-env",
config={"type": "cloud", "networking": {"type": "unrestricted"}},
setup_commands=[
"git clone https://github.com/yourorg/yourproject.git /workspace",
"cd /workspace && npm install",
],
)
session = client.beta.sessions.create(
agent=agent.id,
environment_id=env.id,
title="Migrate React class components to hooks",
)
print(f"\nSession ID: {session.id}")
print("The agent will keep running even after you close this script.\n")
# Send the task
with client.beta.sessions.events.stream(session.id) as stream:
client.beta.sessions.events.send(
session.id,
events=[{
"type": "user.message",
"content": [{
"type": "text",
"text": """Convert all React class components in /workspace/src/components/
to functional components with hooks. There are 47 files.
For each file:
1. Convert componentDidMount -> useEffect
2. Convert componentDidUpdate -> useEffect with deps
3. Convert componentWillUnmount -> useEffect cleanup
4. Convert this.state/this.setState -> useState
5. Convert static contextType -> useContext
6. Preserve all prop types (convert to TypeScript interfaces if .tsx)
7. Run npm test after every 5 files
Track progress by writing to /workspace/MIGRATION_PROGRESS.md after each file.""",
}],
}],
)
import time
start = time.time()
for event in stream:
# Watch for 60 seconds, then disconnect
if time.time() - start > 60:
print("\n\nDisconnecting. Agent continues in the cloud.")
break
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")启动时的预期终端输出:
$ python start_task.py
Session ID: sess_01JXK9N4XYZ789
The agent will keep running even after you close this script.
Starting migration of React class components to hooks.
Scanning /workspace/src/components/ for class components...
[Tool: Glob]
Found 47 files with class components.
[1/47] Converting Header.jsx
[Tool: Read]
[Tool: Edit]
Converted: 2 useState hooks, 1 useEffect hook
[Tool: Write]
Updated MIGRATION_PROGRESS.md
[2/47] Converting Sidebar.jsx
[Tool: Read]
[Tool: Edit]
Converted: 3 useState hooks, 2 useEffect hooks, 1 useContext hook
Disconnecting. Agent continues in the cloud.稍后重新连接
# reconnect.py
from anthropic import Anthropic
client = Anthropic()
session_id = "sess_..." # The ID you saved earlier
# Check if the agent is still working or finished
session = client.beta.sessions.retrieve(session_id)
print(f"Status: {session.status}") # "running" or "idle"
# Get all events, including ones generated while you were away
events = client.beta.sessions.events.list(session_id)
for event in events:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")
case "session.status_idle":
print("\n\n--- Agent finished while you were away ---")
# If it is still running, you can stream the remaining events
if session.status == "running":
print("\nAgent still working. Streaming remaining events...\n")
with client.beta.sessions.events.stream(session_id) as stream:
for event in stream:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "session.status_idle":
print("\n\n--- Agent finished ---")
break重新连接时的预期终端输出:
$ python reconnect.py
Status: idle
[1/47] Converting Header.jsx
Converted: 2 useState hooks, 1 useEffect hook
[2/47] Converting Sidebar.jsx
Converted: 3 useState hooks, 2 useEffect hooks, 1 useContext hook
...
[5/47] Converting Dashboard.jsx
Converted: 5 useState hooks, 3 useEffect hooks
[Tool: Bash]
npm test: 127/127 pass (checkpoint after 5 files)
[6/47] Converting UserProfile.jsx
...
[47/47] Converting LegacyModal.jsx
Converted: 1 useState hook, 1 useEffect hook
[Tool: Bash]
npm test: 127/127 pass (final run)
Migration complete.
- 47 files converted
- 127/127 tests pass
- 89 useState hooks, 63 useEffect hooks, 12 useContext hooks created
- See MIGRATION_PROGRESS.md for full details
--- Agent finished while you were away ---刚才发生了什么?
底层发生了什么
时间线:
0:00 你启动 session,发送任务
0:00-1:00 你流式接收事件,观看 Claude 工作
1:00 你断开连接(关闭脚本、合上笔记本)
1:00-?? Agent 继续在 Anthropic 云端运行
转换文件、运行测试、修复失败
所有事件追加到 session 日志
?? Agent 完成,session 变为 idle
后来 你运行 reconnect.py
getEvents() 返回完整历史
包括你不在时产生的所有内容关键区分:Session 不是上下文窗口。上下文窗口是 Claude 的工作记忆——它填满后会被压缩。Session 是持久化的事件日志——包含所有事件,永不丢失。
Demo 39: 自评估 Agent 与 pass@k 指标
pass@k 是什么
pass@k 衡量 agent 在 k 次尝试内正确完成任务的频率。pass@1 意味着第一次就成功。pass@5 意味着在 5 次内成功。ECC 项目的 eval-harness 使用这个指标追踪 Claude Code 在不同任务类型上的可靠性。
对于自评估 agent,你定义成功标准并让 agent 迭代直到所有标准通过——有效地将 pass@1 变成 pass@k,其中 k 自动确定。
代码
# self_eval_agent.py
from anthropic import Anthropic
client = Anthropic()
agent = client.beta.agents.create(
name="Self-Evaluating API Builder",
model="claude-sonnet-4-6",
system="""You are a quality-obsessed coding agent.
After completing any task, you MUST self-evaluate against the provided
success criteria. For each criterion:
- Run a concrete test or check (not just read the code)
- Record PASS or FAIL with evidence
- If any criterion fails, fix the issue and re-evaluate ALL criteria
Do not declare success until every criterion passes with evidence.
Track your iterations:
Iteration 1: implement, test, evaluate
Iteration 2: fix failures, re-test, re-evaluate
...
Iteration N: all criteria pass""",
tools=[{"type": "agent_toolset_20260401"}],
)
env = client.beta.environments.create(
name="eval-env",
config={"type": "cloud", "networking": {"type": "unrestricted"}},
setup_commands=["pip install flask pytest requests"],
)
session = client.beta.sessions.create(
agent=agent.id,
environment_id=env.id,
)
with client.beta.sessions.events.stream(session.id) as stream:
client.beta.sessions.events.send(
session.id,
events=[{
"type": "user.message",
"content": [{
"type": "text",
"text": """Build a REST API for a task management system. Here are
the SUCCESS CRITERIA -- every single one must pass:
FUNCTIONAL CRITERIA:
1. GET /tasks returns a JSON array of tasks (200)
2. POST /tasks creates a task with title (required), description (optional),
status (defaults to "pending") and returns 201
3. PUT /tasks/:id updates a task and returns 200
4. DELETE /tasks/:id removes a task and returns 204
5. GET /tasks/:id returns a single task or 404
6. POST /tasks with missing title returns 400 with error message
VALIDATION CRITERIA:
7. Title must be 1-200 characters; reject with 400 if invalid
8. Status must be one of: pending, in_progress, done; reject with 400 if invalid
9. All error responses use format: {"error": "message", "status": 400}
QUALITY CRITERIA:
10. All endpoints have tests (at least one test per endpoint)
11. ALL tests pass when run with pytest
12. All functions have type hints
13. API handles concurrent requests without data corruption
EVALUATION PROCESS:
After implementation:
1. Run ALL tests and record pass/fail for each
2. Manually test each endpoint with curl
3. Test error cases explicitly
4. Check type hints with a grep for 'def ' without '->'
5. Report: Iteration N -- Criteria 1-13: PASS/FAIL with evidence
If ANY criterion fails, fix it and re-run the FULL evaluation.""",
}],
}],
)
iteration = 0
for event in stream:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
text = block.text
if "Iteration" in text and "PASS" in text:
iteration += 1
print(text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")
case "session.status_idle":
print(f"\n\nAgent completed after {max(iteration, 1)} iteration(s)")
break预期行为
Iteration 1:
Building API implementation...
[Tool: Write]
[Tool: Write]
Writing tests...
[Tool: Write]
Running tests...
[Tool: Bash]
Criteria check:
1. GET /tasks returns JSON array: PASS (curl returns [])
2. POST /tasks creates task: PASS (returns 201 with task object)
3. PUT /tasks/:id updates task: PASS (returns 200 with updated task)
4. DELETE /tasks/:id removes task: PASS (returns 204)
5. GET /tasks/:id single task or 404: PASS (returns task or 404)
6. POST /tasks missing title returns 400: PASS (returns 400 with error)
7. Title 1-200 chars validation: PASS (empty and 201-char titles rejected)
8. Status enum validation: FAIL -- status "invalid" accepted
9. Error format consistent: FAIL -- DELETE 404 returns plain text
10. All endpoints have tests: PASS (13 test functions)
11. All tests pass: PASS (13/13)
12. All functions have type hints: FAIL -- 3 functions missing return types
13. Concurrent request safety: PASS (using thread lock on data store)
3 criteria failed. Fixing...
Iteration 2:
[Tool: Edit]
Added status enum validation to POST and PUT handlers
[Tool: Edit]
Fixed DELETE 404 to return JSON error format
[Tool: Edit]
Added missing type hints to 3 functions
[Tool: Bash]
Re-running ALL tests...
Criteria check:
1-7: PASS
8: PASS -- status "invalid" now returns 400
9: PASS -- all errors use consistent JSON format
10-11: PASS -- 15/15 tests (added 2 new tests for fixes)
12: PASS -- all functions have type hints
13: PASS
ALL 13 CRITERIA PASS. Implementation complete.
Agent completed after 2 iteration(s)刚才发生了什么?
在 Anthropic 的内部基准测试中,自评估 agent 在 SWE-bench 任务上将 pass@1 提高了最多 10 个百分点。
常见问题排查
Session 恢复失败
症状:你尝试重新连接到一个 session 但遇到错误或陈旧数据。
$ python reconnect.py
anthropic.NotFoundError: Session sess_01JXK9N4XYZ789 not found原因:Session 有最长生命周期(活跃 session 通常 24 小时,空闲 session 可能更早被清理)。如果你在 session 过期后重新连接,就找不到了。
修复:
# Always handle session expiry gracefully
try:
session = client.beta.sessions.retrieve(session_id)
except anthropic.NotFoundError:
print("Session expired or not found.")
print("Check your session ID and ensure it has not been idle too long.")
print("You may need to start a new session.")
# Optionally: create a new session and re-send the task预防:对于预期运行数小时的任务,将 session ID 保存到文件并设置定期健康检查:
import time
while True:
session = client.beta.sessions.retrieve(session_id)
print(f"Status: {session.status}, Last event: {session.last_event_at}")
if session.status == "idle":
break
time.sleep(300) # Check every 5 minutes沙箱限制
症状:Agent 尝试安装包或访问网络资源时失败。
[Tool: Bash]
ERROR: pip install torch failed: network access denied原因:环境的网络或文件系统权限对任务来说太严格了。
修复:创建环境时,确保网络和配置符合你的需求:
# If your agent needs network access (to install packages, clone repos, etc.)
environment = client.beta.environments.create(
name="my-env",
config={
"type": "cloud",
"networking": {"type": "unrestricted"}, # Allow network access
},
setup_commands=[
# Install everything the agent might need BEFORE it starts
"pip install torch numpy pandas",
"npm install -g typescript",
],
)如果你需要特定的包,在 setup_commands 中安装而不是依赖 agent 在运行时安装。Setup 命令在 agent 启动前以完全权限运行。
Managed Agent 超时
症状:Agent 在大型任务中途停止响应或 session 过早变为 idle。
[Tool: Bash]
[35/47] Converting SearchResults.jsx...
--- Agent finished --- (but only 35 of 47 files were converted!)原因:Agent 达到了上下文限制并激进地压缩,丢失了进度追踪。或者模型基于压缩后的上下文认为任务"够好了"。
修复:
- 添加显式进度追踪:让 agent 将进度写入一个能在压缩中幸存的文件:
# In your task prompt, add:
"""
IMPORTANT: After each file conversion, append a line to /workspace/PROGRESS.log:
DONE: filename.jsx -> filename.tsx (test status: pass/fail)
Before starting work, read PROGRESS.log to see what has already been done.
This ensures you resume correctly even after context compaction."""- 将大任务拆分为更小的 session:不要用一个 session 处理 47 个文件,创建每组 10 个的 session:
for batch_start in range(0, 47, 10):
batch_end = min(batch_start + 10, 47)
# Create a new session for each batch
session = client.beta.sessions.create(...)
# Task prompt: "Convert files {batch_start+1} through {batch_end}"- 使用 session 链接:启动新 session,通过读取进度文件接续上一个 session 的工作。
企业采用模式
这些不是假设。这些公司公开分享了他们的 Managed Agents 使用情况。
| 公司 | 场景 | 架构 | 结果 |
|---|---|---|---|
| Notion | 团队-agent 协作 | Agent 处理文档整理、交叉引用和模板生成 | 手动文档管理开销减少 60% |
| Rakuten | 部门专属 agent | 产品、销售、营销和 HR 各有专属 agent | 每个 agent 一周内部署 |
| Sentry | 调试 agent + 补丁 PR | Agent 读取错误报告、定位根因、提交修复 PR | 几周完成(原估几个月) |
| Atlassian | Jira 工作流 agent | 在 Jira 看板中直接将任务分配给 agent | Agent 与人类工程师并肩处理常规工单 |
| Asana | AI Teammates | Agent 参与项目管理工作流 | 显著加速高级功能开发 |
采用的共同模式
- 从只读 agent 开始——监控、分析、分流。先建立信任再给写入权限
- 每个领域一个 agent——不要构建"什么都做"的 agent。工具访问范围窄的专用 agent 更可靠
- 生产变更有人参与——agent 提交 PR,人类合并
- 长任务用 session 持久化——重构、迁移和大型测试套件受益于数小时的 session
定价
| 费用项 | 价格 | 说明 |
|---|---|---|
| Token 费用 | 标准 Claude API 费率 | 与直接调用 API 相同 |
| Session 运行时 | $0.08/小时 | 按毫秒计量 |
| 空闲时间 | 不计费 | 等待用户输入或工具返回 |
| Web 搜索 | $10/千次 | 如果 agent 使用 web search |
典型任务的成本估算:
- 快速分析(5 分钟):~$0.01(运行时)+ token 费
- 模块重构(1 小时):~$0.08(运行时)+ token 费
- 大型迁移(8 小时):~$0.64(运行时)+ token 费
对于大多数工作负载,session-hour 费用与 token 费相比可以忽略不计。
练习:端到端 Managed Agent 工作流
构建一个 Managed Agent:
- 克隆一个 GitHub 仓库(或创建一个有意的测试失败的模拟项目)
- 运行现有测试套件
- 识别失败的测试并分析根因
- 修复代码(不是测试)
- 重新运行测试套件确认全部通过
- 生成修复报告解释每个变更
使用自评估,成功标准:
- 所有原本通过的测试仍然通过(无回归)
- 之前失败的测试现在通过
- 修复报告解释了每个失败的根因和修复方法
- 除非绝对必要否则不创建新文件
本章小结
- Managed Agents 将 agent 基础设施移到 Anthropic 的云端——你定义任务,他们处理执行、状态管理和崩溃恢复
- 三层虚拟化(Session/Harness/Sandbox)意味着每个组件可以独立失败或替换
- Session 持久化让你断开并重新连接而不丢失工作——事件日志是追加写入的,永不丢失
- 自评估 agent 迭代直到成功标准通过,将 pass@1 提高最多 10 个百分点
- 企业采用遵循一个模式:从只读开始,按领域专业化,生产写入有人参与
附录链接:A15 Claude Code 内部机制 解释了 Managed Agents 构建于其上的 harness 循环、session 管理和压缩机制。A13 多 Agent 协调 涵盖共识、冲突解决和多 agent 部署的扩展模式。
研究参考:Anthropic 的"Building effective agents"博客文章(2024)是 Managed Agents 背后设计原则的基础文档。核心要点:优先使用简单、可组合的工具而不是复杂的工作流;让模型驱动控制流;保持 harness 精简,不要编码随着模型改进而过时的假设。
知识检测
下一章:Chapter 15: 生产级工作流设计——综合一切的收官章节。