Chapter 14: Managed Agents & Harness Architecture
What You Will Learn
- The three-layer virtualization architecture: Session, Harness, Sandbox
- Creating and managing cloud-based agents via the Managed Agents API
- Long-running tasks with disconnect/reconnect (session persistence)
- Self-evaluating agents with pass@k metrics
- Enterprise adoption patterns from Notion, Rakuten, Sentry, and Atlassian
- When to use CLI vs Agent SDK vs Managed Agents
Why Managed Agents Exist
Before Managed Agents, building a production AI agent required you to solve the infrastructure problem yourself:
You wanted to build: You actually had to build:
"Agent that refactors code" Container orchestration
State management
Crash recovery
Permission sandboxing
Context window management
Prompt caching
Tool execution routing
Infrastructure work >> Agent logicManaged Agents moves all of that to Anthropic's cloud. You define what the agent should do. Anthropic handles how it runs.
Research reference: Anthropic's "Building effective agents" blog post (2024) established the design principles that Managed Agents implements: keep agent logic separate from infrastructure, prefer simple tool interfaces, and let the model drive the control flow rather than hard-coding it into the harness. The three-layer architecture described below is a direct realization of those principles at infrastructure scale.
The Architecture
Three-Layer Virtualization
This design borrows from operating systems. Each layer is independent -- if one fails or needs replacing, the other two are unaffected.
┌──────────────────────────────────────────────────┐
│ │
│ SESSION (the event log) │
│ ┌─────────────────────────────────────────┐ │
│ │ Append-only log of all events │ │
│ │ Persisted independently of Harness │ │
│ │ and Sandbox │ │
│ │ Queryable via getEvents() at any time │ │
│ │ Never lost, even if everything else │ │
│ │ crashes │ │
│ └─────────────────────────────────────────┘ │
│ │
│ HARNESS (the orchestration loop) │
│ ┌─────────────────────────────────────────┐ │
│ │ Call Claude -> route tool calls -> loop │ │
│ │ Stateless and replaceable ("cattle") │ │
│ │ Built-in prompt caching + compaction │ │
│ │ Recovers via wake(sessionId) after │ │
│ │ a crash │ │
│ └─────────────────────────────────────────┘ │
│ │
│ SANDBOX (the execution environment) │
│ ┌─────────────────────────────────────────┐ │
│ │ Exposes execute(name, input) -> string │ │
│ │ Could be a container, a VM, a phone, │ │
│ │ any execution target │ │
│ │ Harness does not know or care about │ │
│ │ the implementation │ │
│ └─────────────────────────────────────────┘ │
│ │
└──────────────────────────────────────────────────┘"Brain vs. Hands" -- Why This Matters
From Anthropic's engineering blog:
The harness encodes assumptions about what the model cannot do, and those assumptions go stale.
Concrete example: Sonnet 4.5 would wrap up early when approaching the context limit ("context anxiety"). Anthropic added context resets to the harness to work around it. When they switched to Opus 4.5, that behavior was gone, and the resets became unnecessary overhead.
The lesson: decouple the reasoning engine (brain) from the execution environment (hands), so you can upgrade either one independently. This is the same principle as operating system virtualization -- design systems for programs that have not been imagined yet.
From "Pets" to "Cattle"
Coupled design (pets): Decoupled design (cattle):
┌─────────────────────┐ Harness -> stateless, swap if it dies
│ Container │ Sandbox -> container, spin up a new one
│ ├── Harness │ Session -> persisted log, never lost
│ ├── Claude │
│ ├── Tool execution │ If a pet gets sick, you nurse it.
│ └── Session data │ If cattle goes down, you replace it.
└─────────────────────┘
Container dies = everything lostPerformance gains from decoupling: inference starts before the container is ready.
- p50 TTFT (time to first token): dropped ~60%
- p95 TTFT: dropped over 90%
Four API Resources
| Resource | Purpose | Lifecycle |
|---|---|---|
| Agent | Model, system prompt, tools | Reusable across many sessions |
| Environment | Container template, network rules | Reusable |
| Session | A running instance of Agent + Environment | Created per task |
| Events | Messages, status updates, tool results | Streamed via SSE |
CLI vs Agent SDK vs Managed Agents
| CLI | Agent SDK | Managed Agents | |
|---|---|---|---|
| Runs on | Your terminal | Your server | Anthropic's cloud |
| Best for | Interactive dev | CI/CD, embedded apps | Async cloud tasks |
| Duration | Session-scoped | Up to you | Hours to days |
| Infrastructure | None | You manage | Anthropic manages |
| Fault tolerance | Manual | You implement | Automatic |
| Pricing | Subscription | Per-token | Per-token + $0.08/session-hour |
Demo 37: Managed Agent for Real Codebase Refactoring
Scenario
You have a Python backend with a 1,500-line services.py file that needs to be split into separate service modules. This is a multi-hour task with many files to create, imports to update, and tests to fix. It is exactly the kind of work where Managed Agents shines -- you kick it off and come back to the results.
Prerequisites
export ANTHROPIC_API_KEY="sk-ant-..."
pip install anthropicThe Code
# refactor_agent.py
from anthropic import Anthropic
client = Anthropic()
# Step 1: Create an agent specialized for refactoring
agent = client.beta.agents.create(
name="Python Refactoring Agent",
description="Splits monolithic Python modules into well-organized packages",
model="claude-sonnet-4-6",
system="""You are an expert Python refactoring agent. Your approach:
1. Read the target module completely before making any changes
2. Identify natural boundaries (classes, function groups, domain concepts)
3. Create the new package structure with __init__.py files
4. Move code in dependency order (leaf modules first)
5. Update all imports across the entire project
6. Run the test suite after each move to catch breakage early
7. Never change business logic -- this is a pure structural refactor
When you encounter circular imports, resolve them by:
- Extracting shared types into a types.py module
- Using TYPE_CHECKING blocks for type-only imports
- Restructuring the dependency graph
Commit after each successful module extraction with a clear message.""",
tools=[{"type": "agent_toolset_20260401"}],
)
# Step 2: Create an environment with git access
environment = client.beta.environments.create(
name="refactor-sandbox",
config={
"type": "cloud",
"networking": {"type": "unrestricted"},
},
setup_commands=[
"pip install pytest",
"git clone https://github.com/yourorg/yourproject.git /workspace",
"cd /workspace && pip install -e .",
],
)
# Step 3: Start the session
session = client.beta.sessions.create(
agent=agent.id,
environment_id=environment.id,
title="Split services.py into service package",
)
print(f"Session ID: {session.id}")
print("Save this -- you can disconnect and reconnect at any time.\n")
# Step 4: Send the task and stream initial results
with client.beta.sessions.events.stream(session.id) as stream:
client.beta.sessions.events.send(
session.id,
events=[{
"type": "user.message",
"content": [{
"type": "text",
"text": """Refactor /workspace/src/services.py into a services/ package.
The file currently contains:
- UserService (user CRUD, authentication)
- OrderService (order management, cart operations)
- PaymentService (payment processing, refunds)
- NotificationService (email, SMS, push notifications)
- AnalyticsService (event tracking, reporting)
Split into:
src/services/
__init__.py (re-exports for backward compatibility)
user_service.py
order_service.py
payment_service.py
notification_service.py
analytics_service.py
types.py (shared types/interfaces)
Requirements:
1. All existing imports like `from services import UserService` must still work
2. All tests must pass after the refactor
3. No business logic changes
4. Each module should have proper __all__ exports
5. Resolve any circular dependencies cleanly
Start by reading services.py and the test files to understand the current structure.""",
}],
}],
)
# Stream events as they come
for event in stream:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")
case "session.status_idle":
print("\n\n--- Agent finished ---")
breakRunning It
python refactor_agent.pyExpected terminal output:
$ python refactor_agent.py
Session ID: sess_01JXK9M2ABC3DEF456
Save this -- you can disconnect and reconnect at any time.
Reading /workspace/src/services.py to understand the current structure...
[Tool: Read]
The file is 1,487 lines. I can see 5 service classes with clear boundaries.
Reading test files to understand the test structure...
[Tool: Glob]
[Tool: Read]
Found 3 test files: test_user.py (18 tests), test_orders.py (14 tests),
test_payments.py (9 tests). Total: 41 tests.
Starting extraction. I will work in dependency order:
1. types.py (shared types, no dependencies)
2. analytics_service.py (leaf node, no internal deps)
3. notification_service.py (leaf node)
4. payment_service.py (depends on notification)
5. order_service.py (depends on payment)
6. user_service.py (depends on notification)
7. __init__.py (re-exports everything)
Extracting shared types...
[Tool: Write]
[Tool: Edit]
Running tests after types extraction...
[Tool: Bash]
41/41 tests pass.
[Tool: Bash]
Committed: "Extract shared types into services/types.py"
Extracting AnalyticsService...
[Tool: Write]
[Tool: Edit]
[Tool: Bash]
41/41 tests pass.
[Tool: Bash]
Committed: "Extract AnalyticsService into services/analytics_service.py"
...
Creating __init__.py with backward-compatible re-exports...
[Tool: Write]
Final test run...
[Tool: Bash]
41/41 tests pass. All imports verified.
[Tool: Bash]
Committed: "Add __init__.py with re-exports for backward compatibility"
Refactoring complete. Summary:
- 1 monolithic file (1,487 lines) -> 7 focused modules
- All 41 tests pass
- All existing imports remain valid
- 7 commits, one per extraction step
--- Agent finished ---What Just Happened?
Demo 38: Long-Running Task with Session Persistence
The Point
You start a task, disconnect (close your laptop, go to lunch, whatever), and the agent keeps running in the cloud. When you reconnect, you pick up right where it left off -- including all the work done while you were away.
Start the Task
# start_task.py
from anthropic import Anthropic
client = Anthropic()
agent = client.beta.agents.create(
name="Migration Agent",
model="claude-sonnet-4-6",
system="You are a thorough migration agent. Work methodically through each file.",
tools=[{"type": "agent_toolset_20260401"}],
)
env = client.beta.environments.create(
name="migration-env",
config={"type": "cloud", "networking": {"type": "unrestricted"}},
setup_commands=[
"git clone https://github.com/yourorg/yourproject.git /workspace",
"cd /workspace && npm install",
],
)
session = client.beta.sessions.create(
agent=agent.id,
environment_id=env.id,
title="Migrate React class components to hooks",
)
print(f"\nSession ID: {session.id}")
print("The agent will keep running even after you close this script.\n")
# Send the task
with client.beta.sessions.events.stream(session.id) as stream:
client.beta.sessions.events.send(
session.id,
events=[{
"type": "user.message",
"content": [{
"type": "text",
"text": """Convert all React class components in /workspace/src/components/
to functional components with hooks. There are 47 files.
For each file:
1. Convert componentDidMount -> useEffect
2. Convert componentDidUpdate -> useEffect with deps
3. Convert componentWillUnmount -> useEffect cleanup
4. Convert this.state/this.setState -> useState
5. Convert static contextType -> useContext
6. Preserve all prop types (convert to TypeScript interfaces if .tsx)
7. Run npm test after every 5 files
Track progress by writing to /workspace/MIGRATION_PROGRESS.md after each file.""",
}],
}],
)
import time
start = time.time()
for event in stream:
# Watch for 60 seconds, then disconnect
if time.time() - start > 60:
print("\n\nDisconnecting. Agent continues in the cloud.")
break
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")Expected terminal output when starting:
$ python start_task.py
Session ID: sess_01JXK9N4XYZ789
The agent will keep running even after you close this script.
Starting migration of React class components to hooks.
Scanning /workspace/src/components/ for class components...
[Tool: Glob]
Found 47 files with class components.
[1/47] Converting Header.jsx
[Tool: Read]
[Tool: Edit]
Converted: 2 useState hooks, 1 useEffect hook
[Tool: Write]
Updated MIGRATION_PROGRESS.md
[2/47] Converting Sidebar.jsx
[Tool: Read]
[Tool: Edit]
Converted: 3 useState hooks, 2 useEffect hooks, 1 useContext hook
Disconnecting. Agent continues in the cloud.Reconnect Later
# reconnect.py
from anthropic import Anthropic
client = Anthropic()
session_id = "sess_..." # The ID you saved earlier
# Check if the agent is still working or finished
session = client.beta.sessions.retrieve(session_id)
print(f"Status: {session.status}") # "running" or "idle"
# Get all events, including ones generated while you were away
events = client.beta.sessions.events.list(session_id)
for event in events:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")
case "session.status_idle":
print("\n\n--- Agent finished while you were away ---")
# If it is still running, you can stream the remaining events
if session.status == "running":
print("\nAgent still working. Streaming remaining events...\n")
with client.beta.sessions.events.stream(session_id) as stream:
for event in stream:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
print(block.text, end="")
case "session.status_idle":
print("\n\n--- Agent finished ---")
breakExpected terminal output when reconnecting:
$ python reconnect.py
Status: idle
[1/47] Converting Header.jsx
Converted: 2 useState hooks, 1 useEffect hook
[2/47] Converting Sidebar.jsx
Converted: 3 useState hooks, 2 useEffect hooks, 1 useContext hook
...
[5/47] Converting Dashboard.jsx
Converted: 5 useState hooks, 3 useEffect hooks
[Tool: Bash]
npm test: 127/127 pass (checkpoint after 5 files)
[6/47] Converting UserProfile.jsx
...
[47/47] Converting LegacyModal.jsx
Converted: 1 useState hook, 1 useEffect hook
[Tool: Bash]
npm test: 127/127 pass (final run)
Migration complete.
- 47 files converted
- 127/127 tests pass
- 89 useState hooks, 63 useEffect hooks, 12 useContext hooks created
- See MIGRATION_PROGRESS.md for full details
--- Agent finished while you were away ---What Just Happened?
What Is Happening Under the Hood
Timeline:
0:00 You start the session, send the task
0:00-1:00 You stream events, watch Claude work
1:00 You disconnect (close script, close laptop)
1:00-?? Agent keeps running on Anthropic's cloud
Converting files, running tests, fixing failures
All events appended to the session log
?? Agent finishes, session goes idle
Later You run reconnect.py
getEvents() returns the FULL history
Including everything done while you were awayKey distinction: a Session is not the context window. The context window is Claude's working memory -- it gets compacted as it fills up. The Session is a persisted event log -- it contains every event, forever.
Demo 39: Self-Evaluating Agent with pass@k Metrics
What pass@k Means
pass@k measures how often an agent gets a task right within k attempts. pass@1 means it succeeds on the first try. pass@5 means it succeeds within 5 attempts. The ECC project's eval-harness uses this metric to track Claude Code's reliability across different task types.
For self-evaluating agents, you define success criteria and let the agent iterate until all criteria pass -- effectively turning a pass@1 into a pass@k where k is determined automatically.
The Code
# self_eval_agent.py
from anthropic import Anthropic
client = Anthropic()
agent = client.beta.agents.create(
name="Self-Evaluating API Builder",
model="claude-sonnet-4-6",
system="""You are a quality-obsessed coding agent.
After completing any task, you MUST self-evaluate against the provided
success criteria. For each criterion:
- Run a concrete test or check (not just read the code)
- Record PASS or FAIL with evidence
- If any criterion fails, fix the issue and re-evaluate ALL criteria
Do not declare success until every criterion passes with evidence.
Track your iterations:
Iteration 1: implement, test, evaluate
Iteration 2: fix failures, re-test, re-evaluate
...
Iteration N: all criteria pass""",
tools=[{"type": "agent_toolset_20260401"}],
)
env = client.beta.environments.create(
name="eval-env",
config={"type": "cloud", "networking": {"type": "unrestricted"}},
setup_commands=["pip install flask pytest requests"],
)
session = client.beta.sessions.create(
agent=agent.id,
environment_id=env.id,
)
with client.beta.sessions.events.stream(session.id) as stream:
client.beta.sessions.events.send(
session.id,
events=[{
"type": "user.message",
"content": [{
"type": "text",
"text": """Build a REST API for a task management system. Here are
the SUCCESS CRITERIA -- every single one must pass:
FUNCTIONAL CRITERIA:
1. GET /tasks returns a JSON array of tasks (200)
2. POST /tasks creates a task with title (required), description (optional),
status (defaults to "pending") and returns 201
3. PUT /tasks/:id updates a task and returns 200
4. DELETE /tasks/:id removes a task and returns 204
5. GET /tasks/:id returns a single task or 404
6. POST /tasks with missing title returns 400 with error message
VALIDATION CRITERIA:
7. Title must be 1-200 characters; reject with 400 if invalid
8. Status must be one of: pending, in_progress, done; reject with 400 if invalid
9. All error responses use format: {"error": "message", "status": 400}
QUALITY CRITERIA:
10. All endpoints have tests (at least one test per endpoint)
11. ALL tests pass when run with pytest
12. All functions have type hints
13. API handles concurrent requests without data corruption
EVALUATION PROCESS:
After implementation:
1. Run ALL tests and record pass/fail for each
2. Manually test each endpoint with curl
3. Test error cases explicitly
4. Check type hints with a grep for 'def ' without '->'
5. Report: Iteration N -- Criteria 1-13: PASS/FAIL with evidence
If ANY criterion fails, fix it and re-run the FULL evaluation.""",
}],
}],
)
iteration = 0
for event in stream:
match event.type:
case "agent.message":
for block in event.content:
if hasattr(block, "text"):
text = block.text
if "Iteration" in text and "PASS" in text:
iteration += 1
print(text, end="")
case "agent.tool_use":
print(f"\n [Tool: {event.name}]")
case "session.status_idle":
print(f"\n\nAgent completed after {max(iteration, 1)} iteration(s)")
breakExpected Behavior
Iteration 1:
Building API implementation...
[Tool: Write]
[Tool: Write]
Writing tests...
[Tool: Write]
Running tests...
[Tool: Bash]
Criteria check:
1. GET /tasks returns JSON array: PASS (curl returns [])
2. POST /tasks creates task: PASS (returns 201 with task object)
3. PUT /tasks/:id updates task: PASS (returns 200 with updated task)
4. DELETE /tasks/:id removes task: PASS (returns 204)
5. GET /tasks/:id single task or 404: PASS (returns task or 404)
6. POST /tasks missing title returns 400: PASS (returns 400 with error)
7. Title 1-200 chars validation: PASS (empty and 201-char titles rejected)
8. Status enum validation: FAIL -- status "invalid" accepted
9. Error format consistent: FAIL -- DELETE 404 returns plain text
10. All endpoints have tests: PASS (13 test functions)
11. All tests pass: PASS (13/13)
12. All functions have type hints: FAIL -- 3 functions missing return types
13. Concurrent request safety: PASS (using thread lock on data store)
3 criteria failed. Fixing...
Iteration 2:
[Tool: Edit]
Added status enum validation to POST and PUT handlers
[Tool: Edit]
Fixed DELETE 404 to return JSON error format
[Tool: Edit]
Added missing type hints to 3 functions
[Tool: Bash]
Re-running ALL tests...
Criteria check:
1-7: PASS
8: PASS -- status "invalid" now returns 400
9: PASS -- all errors use consistent JSON format
10-11: PASS -- 15/15 tests (added 2 new tests for fixes)
12: PASS -- all functions have type hints
13: PASS
ALL 13 CRITERIA PASS. Implementation complete.
Agent completed after 2 iteration(s)What Just Happened?
In Anthropic's internal benchmarks, self-evaluating agents improved pass@1 by up to 10 percentage points on SWE-bench tasks.
When Things Go Wrong
Session Resume Failure
Symptom: You try to reconnect to a session but get an error or stale data.
$ python reconnect.py
anthropic.NotFoundError: Session sess_01JXK9N4XYZ789 not foundCause: Sessions have a maximum lifetime (typically 24 hours for active sessions, though idle sessions may be cleaned up sooner). If you reconnect after the session has expired, it will not be found.
Fix:
# Always handle session expiry gracefully
try:
session = client.beta.sessions.retrieve(session_id)
except anthropic.NotFoundError:
print("Session expired or not found.")
print("Check your session ID and ensure it has not been idle too long.")
print("You may need to start a new session.")
# Optionally: create a new session and re-send the taskPrevention: For tasks expected to run for many hours, save the session ID to a file and set up a periodic health check:
import time
while True:
session = client.beta.sessions.retrieve(session_id)
print(f"Status: {session.status}, Last event: {session.last_event_at}")
if session.status == "idle":
break
time.sleep(300) # Check every 5 minutesSandbox Limitations
Symptom: The agent tries to install a package or access a network resource and fails.
[Tool: Bash]
ERROR: pip install torch failed: network access deniedCause: The environment's networking or filesystem permissions are too restrictive for the task.
Fix: When creating the environment, ensure the networking and setup match your needs:
# If your agent needs network access (to install packages, clone repos, etc.)
environment = client.beta.environments.create(
name="my-env",
config={
"type": "cloud",
"networking": {"type": "unrestricted"}, # Allow network access
},
setup_commands=[
# Install everything the agent might need BEFORE it starts
"pip install torch numpy pandas",
"npm install -g typescript",
],
)If you need specific packages, install them in setup_commands rather than relying on the agent to install them at runtime. Setup commands run with full permissions before the agent starts.
Managed Agent Timeout
Symptom: The agent stops responding or the session goes idle prematurely in the middle of a large task.
[Tool: Bash]
[35/47] Converting SearchResults.jsx...
--- Agent finished --- (but only 35 of 47 files were converted!)Cause: The agent hit a context limit and compacted aggressively, losing track of its progress. Or the model decided the task was "done enough" based on its compacted context.
Fix:
- Add explicit progress tracking: Have the agent write progress to a file that survives compaction:
# In your task prompt, add:
"""
IMPORTANT: After each file conversion, append a line to /workspace/PROGRESS.log:
DONE: filename.jsx -> filename.tsx (test status: pass/fail)
Before starting work, read PROGRESS.log to see what has already been done.
This ensures you resume correctly even after context compaction."""- Break large tasks into smaller sessions: Instead of one session for 47 files, create sessions for groups of 10:
for batch_start in range(0, 47, 10):
batch_end = min(batch_start + 10, 47)
# Create a new session for each batch
session = client.beta.sessions.create(...)
# Task prompt: "Convert files {batch_start+1} through {batch_end}"- Use session chaining: Start a new session that picks up where the previous one left off by reading the progress file.
Enterprise Adoption Patterns
These are not hypothetical. These companies publicly shared their Managed Agents usage.
| Company | Use Case | Architecture | Results |
|---|---|---|---|
| Notion | Team-agent collaboration | Agents handle document organization, cross-referencing, and template generation | Reduced manual document management overhead by 60% |
| Rakuten | Department-specific agents | Product, sales, marketing, and HR each got a specialized agent | Each agent deployed within one week |
| Sentry | Debug agent + patch PRs | Agent reads error reports, identifies root cause, submits fix PRs | Completed in weeks (originally estimated months) |
| Atlassian | Jira workflow agents | Tasks assigned directly to agents in Jira boards | Agents handle routine tickets alongside human engineers |
| Asana | AI Teammates | Agents participate in project management workflows | Significantly accelerated advanced feature development |
Common Patterns Across Adoptions
- Start with read-only agents -- monitoring, analysis, triage. Build trust before giving write access
- One agent per domain -- do not build a "do everything" agent. Specialized agents with narrow tool access are more reliable
- Human-in-the-loop for production changes -- agents propose PRs, humans merge them
- Session persistence for long tasks -- refactoring, migrations, and large test suites benefit from hours-long sessions
Pricing
| Cost Item | Price | Notes |
|---|---|---|
| Token usage | Standard Claude API rates | Same as calling the API directly |
| Session runtime | $0.08/hour | Billed by the millisecond |
| Idle time | Not billed | Waiting for user input or tool responses |
| Web search | $10/1,000 queries | If the agent uses web search |
Cost estimates for typical tasks:
- Quick analysis (5 min): ~$0.01 (runtime) + token costs
- Module refactoring (1 hour): ~$0.08 (runtime) + token costs
- Large migration (8 hours): ~$0.64 (runtime) + token costs
The session-hour cost is negligible compared to token costs for most workloads.
Exercise: End-to-End Managed Agent Workflow
Build a Managed Agent that:
- Clones a GitHub repo (or creates a mock project with intentional test failures)
- Runs the existing test suite
- Identifies failing tests and analyzes root causes
- Fixes the code (not the tests)
- Re-runs the test suite to confirm all tests pass
- Generates a fix report explaining each change
Use self-evaluation with these criteria:
- All previously passing tests still pass (no regressions)
- Previously failing tests now pass
- The fix report explains the root cause and fix for each failure
- No new files created unless strictly necessary
Chapter Summary
- Managed Agents moves agent infrastructure to Anthropic's cloud -- you define the task, they handle execution, state management, and crash recovery
- Three-layer virtualization (Session/Harness/Sandbox) means each component can fail or be replaced independently
- Session persistence lets you disconnect and reconnect without losing work -- the event log is append-only and never lost
- Self-evaluating agents iterate until success criteria pass, improving pass@1 by up to 10 percentage points
- Enterprise adoption follows a pattern: start read-only, specialize per domain, human-in-the-loop for production writes
Appendix links: A15 Claude Code Internals explains the harness loop, session management, and compaction mechanics that Managed Agents builds on. A13 Multi-Agent Coordination covers consensus, conflict resolution, and scaling patterns for multi-agent deployments.
Research reference: Anthropic's "Building effective agents" blog post (2024) is the foundational document for the design principles behind Managed Agents. Key takeaways: prefer simple, composable tools over complex workflows; let the model drive control flow; and keep the harness thin so it does not encode assumptions that go stale as models improve.
Knowledge Check
Next: Chapter 15: Production Workflow Design -- the capstone chapter that brings everything together.