Skip to content

Chapter 14: Managed Agents & Harness Architecture ​

What You Will Learn ​

  • The three-layer virtualization architecture: Session, Harness, Sandbox
  • Creating and managing cloud-based agents via the Managed Agents API
  • Long-running tasks with disconnect/reconnect (session persistence)
  • Self-evaluating agents with pass@k metrics
  • Enterprise adoption patterns from Notion, Rakuten, Sentry, and Atlassian
  • When to use CLI vs Agent SDK vs Managed Agents

Why Managed Agents Exist ​

Before Managed Agents, building a production AI agent required you to solve the infrastructure problem yourself:

You wanted to build:                You actually had to build:
  "Agent that refactors code"         Container orchestration
                                      State management
                                      Crash recovery
                                      Permission sandboxing
                                      Context window management
                                      Prompt caching
                                      Tool execution routing
                                      
  Infrastructure work >> Agent logic

Managed Agents moves all of that to Anthropic's cloud. You define what the agent should do. Anthropic handles how it runs.

Research reference: Anthropic's "Building effective agents" blog post (2024) established the design principles that Managed Agents implements: keep agent logic separate from infrastructure, prefer simple tool interfaces, and let the model drive the control flow rather than hard-coding it into the harness. The three-layer architecture described below is a direct realization of those principles at infrastructure scale.

The Architecture ​

Three-Layer Virtualization ​

This design borrows from operating systems. Each layer is independent -- if one fails or needs replacing, the other two are unaffected.

┌──────────────────────────────────────────────────┐
│                                                  │
│  SESSION (the event log)                         │
│  ┌─────────────────────────────────────────┐     │
│  │ Append-only log of all events           │     │
│  │ Persisted independently of Harness      │     │
│  │ and Sandbox                             │     │
│  │ Queryable via getEvents() at any time   │     │
│  │ Never lost, even if everything else     │     │
│  │ crashes                                 │     │
│  └─────────────────────────────────────────┘     │
│                                                  │
│  HARNESS (the orchestration loop)                │
│  ┌─────────────────────────────────────────┐     │
│  │ Call Claude -> route tool calls -> loop  │     │
│  │ Stateless and replaceable ("cattle")    │     │
│  │ Built-in prompt caching + compaction    │     │
│  │ Recovers via wake(sessionId) after      │     │
│  │ a crash                                 │     │
│  └─────────────────────────────────────────┘     │
│                                                  │
│  SANDBOX (the execution environment)             │
│  ┌─────────────────────────────────────────┐     │
│  │ Exposes execute(name, input) -> string  │     │
│  │ Could be a container, a VM, a phone,    │     │
│  │ any execution target                    │     │
│  │ Harness does not know or care about     │     │
│  │ the implementation                      │     │
│  └─────────────────────────────────────────┘     │
│                                                  │
└──────────────────────────────────────────────────┘

"Brain vs. Hands" -- Why This Matters ​

From Anthropic's engineering blog:

The harness encodes assumptions about what the model cannot do, and those assumptions go stale.

Concrete example: Sonnet 4.5 would wrap up early when approaching the context limit ("context anxiety"). Anthropic added context resets to the harness to work around it. When they switched to Opus 4.5, that behavior was gone, and the resets became unnecessary overhead.

The lesson: decouple the reasoning engine (brain) from the execution environment (hands), so you can upgrade either one independently. This is the same principle as operating system virtualization -- design systems for programs that have not been imagined yet.

From "Pets" to "Cattle" ​

Coupled design (pets):               Decoupled design (cattle):
┌─────────────────────┐              Harness  -> stateless, swap if it dies
│ Container           │              Sandbox  -> container, spin up a new one
│ ├── Harness         │              Session  -> persisted log, never lost
│ ├── Claude          │
│ ├── Tool execution  │              If a pet gets sick, you nurse it.
│ └── Session data    │              If cattle goes down, you replace it.
└─────────────────────┘
Container dies = everything lost

Performance gains from decoupling: inference starts before the container is ready.

  • p50 TTFT (time to first token): dropped ~60%
  • p95 TTFT: dropped over 90%

Four API Resources ​

ResourcePurposeLifecycle
AgentModel, system prompt, toolsReusable across many sessions
EnvironmentContainer template, network rulesReusable
SessionA running instance of Agent + EnvironmentCreated per task
EventsMessages, status updates, tool resultsStreamed via SSE

CLI vs Agent SDK vs Managed Agents ​

CLIAgent SDKManaged Agents
Runs onYour terminalYour serverAnthropic's cloud
Best forInteractive devCI/CD, embedded appsAsync cloud tasks
DurationSession-scopedUp to youHours to days
InfrastructureNoneYou manageAnthropic manages
Fault toleranceManualYou implementAutomatic
PricingSubscriptionPer-tokenPer-token + $0.08/session-hour

Demo 37: Managed Agent for Real Codebase Refactoring ​

37
Cloud-Based Refactoring Agent
Advanced~20 min

Scenario ​

You have a Python backend with a 1,500-line services.py file that needs to be split into separate service modules. This is a multi-hour task with many files to create, imports to update, and tests to fix. It is exactly the kind of work where Managed Agents shines -- you kick it off and come back to the results.

Prerequisites ​

bash
export ANTHROPIC_API_KEY="sk-ant-..."
pip install anthropic

The Code ​

python
# refactor_agent.py

from anthropic import Anthropic

client = Anthropic()

# Step 1: Create an agent specialized for refactoring
agent = client.beta.agents.create(
    name="Python Refactoring Agent",
    description="Splits monolithic Python modules into well-organized packages",
    model="claude-sonnet-4-6",
    system="""You are an expert Python refactoring agent. Your approach:

1. Read the target module completely before making any changes
2. Identify natural boundaries (classes, function groups, domain concepts)
3. Create the new package structure with __init__.py files
4. Move code in dependency order (leaf modules first)
5. Update all imports across the entire project
6. Run the test suite after each move to catch breakage early
7. Never change business logic -- this is a pure structural refactor

When you encounter circular imports, resolve them by:
- Extracting shared types into a types.py module
- Using TYPE_CHECKING blocks for type-only imports
- Restructuring the dependency graph

Commit after each successful module extraction with a clear message.""",
    tools=[{"type": "agent_toolset_20260401"}],
)

# Step 2: Create an environment with git access
environment = client.beta.environments.create(
    name="refactor-sandbox",
    config={
        "type": "cloud",
        "networking": {"type": "unrestricted"},
    },
    setup_commands=[
        "pip install pytest",
        "git clone https://github.com/yourorg/yourproject.git /workspace",
        "cd /workspace && pip install -e .",
    ],
)

# Step 3: Start the session
session = client.beta.sessions.create(
    agent=agent.id,
    environment_id=environment.id,
    title="Split services.py into service package",
)

print(f"Session ID: {session.id}")
print("Save this -- you can disconnect and reconnect at any time.\n")

# Step 4: Send the task and stream initial results
with client.beta.sessions.events.stream(session.id) as stream:
    client.beta.sessions.events.send(
        session.id,
        events=[{
            "type": "user.message",
            "content": [{
                "type": "text",
                "text": """Refactor /workspace/src/services.py into a services/ package.

The file currently contains:
- UserService (user CRUD, authentication)
- OrderService (order management, cart operations)
- PaymentService (payment processing, refunds)
- NotificationService (email, SMS, push notifications)
- AnalyticsService (event tracking, reporting)

Split into:
  src/services/
    __init__.py          (re-exports for backward compatibility)
    user_service.py
    order_service.py
    payment_service.py
    notification_service.py
    analytics_service.py
    types.py             (shared types/interfaces)

Requirements:
1. All existing imports like `from services import UserService` must still work
2. All tests must pass after the refactor
3. No business logic changes
4. Each module should have proper __all__ exports
5. Resolve any circular dependencies cleanly

Start by reading services.py and the test files to understand the current structure.""",
            }],
        }],
    )

    # Stream events as they come
    for event in stream:
        match event.type:
            case "agent.message":
                for block in event.content:
                    if hasattr(block, "text"):
                        print(block.text, end="")
            case "agent.tool_use":
                print(f"\n  [Tool: {event.name}]")
            case "session.status_idle":
                print("\n\n--- Agent finished ---")
                break

Running It ​

bash
python refactor_agent.py

Expected terminal output:

terminal
$ python refactor_agent.py
Session ID: sess_01JXK9M2ABC3DEF456
Save this -- you can disconnect and reconnect at any time.

Reading /workspace/src/services.py to understand the current structure...
  [Tool: Read]
The file is 1,487 lines. I can see 5 service classes with clear boundaries.

Reading test files to understand the test structure...
  [Tool: Glob]
  [Tool: Read]
Found 3 test files: test_user.py (18 tests), test_orders.py (14 tests),
test_payments.py (9 tests). Total: 41 tests.

Starting extraction. I will work in dependency order:
1. types.py (shared types, no dependencies)
2. analytics_service.py (leaf node, no internal deps)
3. notification_service.py (leaf node)
4. payment_service.py (depends on notification)
5. order_service.py (depends on payment)
6. user_service.py (depends on notification)
7. __init__.py (re-exports everything)

Extracting shared types...
  [Tool: Write]
  [Tool: Edit]
Running tests after types extraction...
  [Tool: Bash]
41/41 tests pass.
  [Tool: Bash]
Committed: "Extract shared types into services/types.py"

Extracting AnalyticsService...
  [Tool: Write]
  [Tool: Edit]
  [Tool: Bash]
41/41 tests pass.
  [Tool: Bash]
Committed: "Extract AnalyticsService into services/analytics_service.py"

...

Creating __init__.py with backward-compatible re-exports...
  [Tool: Write]
Final test run...
  [Tool: Bash]
41/41 tests pass. All imports verified.
  [Tool: Bash]
Committed: "Add __init__.py with re-exports for backward compatibility"

Refactoring complete. Summary:
- 1 monolithic file (1,487 lines) -> 7 focused modules
- All 41 tests pass
- All existing imports remain valid
- 7 commits, one per extraction step

--- Agent finished ---

What Just Happened? ​

1
Anthropic API
agents.create()
↓
2
Anthropic API
environments.create()
↓
3
Anthropic API
sessions.create()
↓
4
SSE Stream
events.stream()

Demo 38: Long-Running Task with Session Persistence ​

38
Disconnect and Reconnect to a Running Agent
Advanced~15 min

The Point ​

You start a task, disconnect (close your laptop, go to lunch, whatever), and the agent keeps running in the cloud. When you reconnect, you pick up right where it left off -- including all the work done while you were away.

Start the Task ​

python
# start_task.py

from anthropic import Anthropic

client = Anthropic()

agent = client.beta.agents.create(
    name="Migration Agent",
    model="claude-sonnet-4-6",
    system="You are a thorough migration agent. Work methodically through each file.",
    tools=[{"type": "agent_toolset_20260401"}],
)

env = client.beta.environments.create(
    name="migration-env",
    config={"type": "cloud", "networking": {"type": "unrestricted"}},
    setup_commands=[
        "git clone https://github.com/yourorg/yourproject.git /workspace",
        "cd /workspace && npm install",
    ],
)

session = client.beta.sessions.create(
    agent=agent.id,
    environment_id=env.id,
    title="Migrate React class components to hooks",
)

print(f"\nSession ID: {session.id}")
print("The agent will keep running even after you close this script.\n")

# Send the task
with client.beta.sessions.events.stream(session.id) as stream:
    client.beta.sessions.events.send(
        session.id,
        events=[{
            "type": "user.message",
            "content": [{
                "type": "text",
                "text": """Convert all React class components in /workspace/src/components/
to functional components with hooks. There are 47 files.

For each file:
1. Convert componentDidMount -> useEffect
2. Convert componentDidUpdate -> useEffect with deps
3. Convert componentWillUnmount -> useEffect cleanup
4. Convert this.state/this.setState -> useState
5. Convert static contextType -> useContext
6. Preserve all prop types (convert to TypeScript interfaces if .tsx)
7. Run npm test after every 5 files

Track progress by writing to /workspace/MIGRATION_PROGRESS.md after each file.""",
            }],
        }],
    )

    import time
    start = time.time()
    for event in stream:
        # Watch for 60 seconds, then disconnect
        if time.time() - start > 60:
            print("\n\nDisconnecting. Agent continues in the cloud.")
            break

        match event.type:
            case "agent.message":
                for block in event.content:
                    if hasattr(block, "text"):
                        print(block.text, end="")
            case "agent.tool_use":
                print(f"\n  [Tool: {event.name}]")

Expected terminal output when starting:

terminal
$ python start_task.py

Session ID: sess_01JXK9N4XYZ789
The agent will keep running even after you close this script.

Starting migration of React class components to hooks.

Scanning /workspace/src/components/ for class components...
  [Tool: Glob]
Found 47 files with class components.

[1/47] Converting Header.jsx
  [Tool: Read]
  [Tool: Edit]
  Converted: 2 useState hooks, 1 useEffect hook
  [Tool: Write]
Updated MIGRATION_PROGRESS.md

[2/47] Converting Sidebar.jsx
  [Tool: Read]
  [Tool: Edit]
  Converted: 3 useState hooks, 2 useEffect hooks, 1 useContext hook

Disconnecting. Agent continues in the cloud.

Reconnect Later ​

python
# reconnect.py

from anthropic import Anthropic

client = Anthropic()
session_id = "sess_..."  # The ID you saved earlier

# Check if the agent is still working or finished
session = client.beta.sessions.retrieve(session_id)
print(f"Status: {session.status}")  # "running" or "idle"

# Get all events, including ones generated while you were away
events = client.beta.sessions.events.list(session_id)

for event in events:
    match event.type:
        case "agent.message":
            for block in event.content:
                if hasattr(block, "text"):
                    print(block.text, end="")
        case "agent.tool_use":
            print(f"\n  [Tool: {event.name}]")
        case "session.status_idle":
            print("\n\n--- Agent finished while you were away ---")

# If it is still running, you can stream the remaining events
if session.status == "running":
    print("\nAgent still working. Streaming remaining events...\n")
    with client.beta.sessions.events.stream(session_id) as stream:
        for event in stream:
            match event.type:
                case "agent.message":
                    for block in event.content:
                        if hasattr(block, "text"):
                            print(block.text, end="")
                case "session.status_idle":
                    print("\n\n--- Agent finished ---")
                    break

Expected terminal output when reconnecting:

terminal
$ python reconnect.py
Status: idle

[1/47] Converting Header.jsx
  Converted: 2 useState hooks, 1 useEffect hook
[2/47] Converting Sidebar.jsx
  Converted: 3 useState hooks, 2 useEffect hooks, 1 useContext hook
...
[5/47] Converting Dashboard.jsx
  Converted: 5 useState hooks, 3 useEffect hooks
  [Tool: Bash]
  npm test: 127/127 pass (checkpoint after 5 files)

[6/47] Converting UserProfile.jsx
...
[47/47] Converting LegacyModal.jsx
  Converted: 1 useState hook, 1 useEffect hook
  [Tool: Bash]
  npm test: 127/127 pass (final run)

Migration complete.
- 47 files converted
- 127/127 tests pass
- 89 useState hooks, 63 useEffect hooks, 12 useContext hooks created
- See MIGRATION_PROGRESS.md for full details

--- Agent finished while you were away ---

What Just Happened? ​

1
sessions.create()
Agent plus Environment
↓
2
events.stream()
Session event stream
↓
3
sessions.retrieve()
Session status check
↓
4
events.list()
Full event history

What Is Happening Under the Hood ​

Timeline:
  0:00    You start the session, send the task
  0:00-1:00  You stream events, watch Claude work
  1:00    You disconnect (close script, close laptop)

  1:00-??    Agent keeps running on Anthropic's cloud
             Converting files, running tests, fixing failures
             All events appended to the session log

  ??         Agent finishes, session goes idle

  Later      You run reconnect.py
             getEvents() returns the FULL history
             Including everything done while you were away

Key distinction: a Session is not the context window. The context window is Claude's working memory -- it gets compacted as it fills up. The Session is a persisted event log -- it contains every event, forever.


Demo 39: Self-Evaluating Agent with pass@k Metrics ​

39
Self-Evaluating Agent with Iterative Improvement
Advanced~20 min

What pass@k Means ​

pass@k measures how often an agent gets a task right within k attempts. pass@1 means it succeeds on the first try. pass@5 means it succeeds within 5 attempts. The ECC project's eval-harness uses this metric to track Claude Code's reliability across different task types.

For self-evaluating agents, you define success criteria and let the agent iterate until all criteria pass -- effectively turning a pass@1 into a pass@k where k is determined automatically.

The Code ​

python
# self_eval_agent.py

from anthropic import Anthropic

client = Anthropic()

agent = client.beta.agents.create(
    name="Self-Evaluating API Builder",
    model="claude-sonnet-4-6",
    system="""You are a quality-obsessed coding agent.

After completing any task, you MUST self-evaluate against the provided
success criteria. For each criterion:
- Run a concrete test or check (not just read the code)
- Record PASS or FAIL with evidence
- If any criterion fails, fix the issue and re-evaluate ALL criteria

Do not declare success until every criterion passes with evidence.

Track your iterations:
  Iteration 1: implement, test, evaluate
  Iteration 2: fix failures, re-test, re-evaluate
  ...
  Iteration N: all criteria pass""",
    tools=[{"type": "agent_toolset_20260401"}],
)

env = client.beta.environments.create(
    name="eval-env",
    config={"type": "cloud", "networking": {"type": "unrestricted"}},
    setup_commands=["pip install flask pytest requests"],
)

session = client.beta.sessions.create(
    agent=agent.id,
    environment_id=env.id,
)

with client.beta.sessions.events.stream(session.id) as stream:
    client.beta.sessions.events.send(
        session.id,
        events=[{
            "type": "user.message",
            "content": [{
                "type": "text",
                "text": """Build a REST API for a task management system. Here are
the SUCCESS CRITERIA -- every single one must pass:

FUNCTIONAL CRITERIA:
1. GET /tasks returns a JSON array of tasks (200)
2. POST /tasks creates a task with title (required), description (optional),
   status (defaults to "pending") and returns 201
3. PUT /tasks/:id updates a task and returns 200
4. DELETE /tasks/:id removes a task and returns 204
5. GET /tasks/:id returns a single task or 404
6. POST /tasks with missing title returns 400 with error message

VALIDATION CRITERIA:
7. Title must be 1-200 characters; reject with 400 if invalid
8. Status must be one of: pending, in_progress, done; reject with 400 if invalid
9. All error responses use format: {"error": "message", "status": 400}

QUALITY CRITERIA:
10. All endpoints have tests (at least one test per endpoint)
11. ALL tests pass when run with pytest
12. All functions have type hints
13. API handles concurrent requests without data corruption

EVALUATION PROCESS:
After implementation:
1. Run ALL tests and record pass/fail for each
2. Manually test each endpoint with curl
3. Test error cases explicitly
4. Check type hints with a grep for 'def ' without '->'
5. Report: Iteration N -- Criteria 1-13: PASS/FAIL with evidence

If ANY criterion fails, fix it and re-run the FULL evaluation.""",
            }],
        }],
    )

    iteration = 0
    for event in stream:
        match event.type:
            case "agent.message":
                for block in event.content:
                    if hasattr(block, "text"):
                        text = block.text
                        if "Iteration" in text and "PASS" in text:
                            iteration += 1
                        print(text, end="")
            case "agent.tool_use":
                print(f"\n  [Tool: {event.name}]")
            case "session.status_idle":
                print(f"\n\nAgent completed after {max(iteration, 1)} iteration(s)")
                break

Expected Behavior ​

terminal
Iteration 1:
  Building API implementation...
  [Tool: Write]
  [Tool: Write]
  Writing tests...
  [Tool: Write]
  Running tests...
  [Tool: Bash]

  Criteria check:
    1. GET /tasks returns JSON array:           PASS (curl returns [])
    2. POST /tasks creates task:                PASS (returns 201 with task object)
    3. PUT /tasks/:id updates task:             PASS (returns 200 with updated task)
    4. DELETE /tasks/:id removes task:          PASS (returns 204)
    5. GET /tasks/:id single task or 404:       PASS (returns task or 404)
    6. POST /tasks missing title returns 400:   PASS (returns 400 with error)
    7. Title 1-200 chars validation:            PASS (empty and 201-char titles rejected)
    8. Status enum validation:                  FAIL -- status "invalid" accepted
    9. Error format consistent:                 FAIL -- DELETE 404 returns plain text
    10. All endpoints have tests:               PASS (13 test functions)
    11. All tests pass:                         PASS (13/13)
    12. All functions have type hints:          FAIL -- 3 functions missing return types
    13. Concurrent request safety:              PASS (using thread lock on data store)

  3 criteria failed. Fixing...

Iteration 2:
  [Tool: Edit]
  Added status enum validation to POST and PUT handlers
  [Tool: Edit]
  Fixed DELETE 404 to return JSON error format
  [Tool: Edit]
  Added missing type hints to 3 functions
  [Tool: Bash]
  Re-running ALL tests...

  Criteria check:
    1-7:   PASS
    8:     PASS -- status "invalid" now returns 400
    9:     PASS -- all errors use consistent JSON format
    10-11: PASS -- 15/15 tests (added 2 new tests for fixes)
    12:    PASS -- all functions have type hints
    13:    PASS

  ALL 13 CRITERIA PASS. Implementation complete.

Agent completed after 2 iteration(s)

What Just Happened? ​

1
Write
app.py and test_app.py
↓
2
Bash
pytest and curl commands
↓
3
Edit
app.py (3 targeted fixes)
↓
4
Bash
Full re-evaluation

In Anthropic's internal benchmarks, self-evaluating agents improved pass@1 by up to 10 percentage points on SWE-bench tasks.


When Things Go Wrong ​

Session Resume Failure ​

Symptom: You try to reconnect to a session but get an error or stale data.

terminal
$ python reconnect.py
anthropic.NotFoundError: Session sess_01JXK9N4XYZ789 not found

Cause: Sessions have a maximum lifetime (typically 24 hours for active sessions, though idle sessions may be cleaned up sooner). If you reconnect after the session has expired, it will not be found.

Fix:

python
# Always handle session expiry gracefully
try:
    session = client.beta.sessions.retrieve(session_id)
except anthropic.NotFoundError:
    print("Session expired or not found.")
    print("Check your session ID and ensure it has not been idle too long.")
    print("You may need to start a new session.")
    # Optionally: create a new session and re-send the task

Prevention: For tasks expected to run for many hours, save the session ID to a file and set up a periodic health check:

python
import time

while True:
    session = client.beta.sessions.retrieve(session_id)
    print(f"Status: {session.status}, Last event: {session.last_event_at}")
    if session.status == "idle":
        break
    time.sleep(300)  # Check every 5 minutes

Sandbox Limitations ​

Symptom: The agent tries to install a package or access a network resource and fails.

terminal
  [Tool: Bash]
  ERROR: pip install torch failed: network access denied

Cause: The environment's networking or filesystem permissions are too restrictive for the task.

Fix: When creating the environment, ensure the networking and setup match your needs:

python
# If your agent needs network access (to install packages, clone repos, etc.)
environment = client.beta.environments.create(
    name="my-env",
    config={
        "type": "cloud",
        "networking": {"type": "unrestricted"},  # Allow network access
    },
    setup_commands=[
        # Install everything the agent might need BEFORE it starts
        "pip install torch numpy pandas",
        "npm install -g typescript",
    ],
)

If you need specific packages, install them in setup_commands rather than relying on the agent to install them at runtime. Setup commands run with full permissions before the agent starts.

Managed Agent Timeout ​

Symptom: The agent stops responding or the session goes idle prematurely in the middle of a large task.

terminal
  [Tool: Bash]
  [35/47] Converting SearchResults.jsx...
  
  --- Agent finished ---   (but only 35 of 47 files were converted!)

Cause: The agent hit a context limit and compacted aggressively, losing track of its progress. Or the model decided the task was "done enough" based on its compacted context.

Fix:

  1. Add explicit progress tracking: Have the agent write progress to a file that survives compaction:
python
# In your task prompt, add:
"""
IMPORTANT: After each file conversion, append a line to /workspace/PROGRESS.log:
  DONE: filename.jsx -> filename.tsx (test status: pass/fail)
Before starting work, read PROGRESS.log to see what has already been done.
This ensures you resume correctly even after context compaction."""
  1. Break large tasks into smaller sessions: Instead of one session for 47 files, create sessions for groups of 10:
python
for batch_start in range(0, 47, 10):
    batch_end = min(batch_start + 10, 47)
    # Create a new session for each batch
    session = client.beta.sessions.create(...)
    # Task prompt: "Convert files {batch_start+1} through {batch_end}"
  1. Use session chaining: Start a new session that picks up where the previous one left off by reading the progress file.

Enterprise Adoption Patterns ​

These are not hypothetical. These companies publicly shared their Managed Agents usage.

CompanyUse CaseArchitectureResults
NotionTeam-agent collaborationAgents handle document organization, cross-referencing, and template generationReduced manual document management overhead by 60%
RakutenDepartment-specific agentsProduct, sales, marketing, and HR each got a specialized agentEach agent deployed within one week
SentryDebug agent + patch PRsAgent reads error reports, identifies root cause, submits fix PRsCompleted in weeks (originally estimated months)
AtlassianJira workflow agentsTasks assigned directly to agents in Jira boardsAgents handle routine tickets alongside human engineers
AsanaAI TeammatesAgents participate in project management workflowsSignificantly accelerated advanced feature development

Common Patterns Across Adoptions ​

  1. Start with read-only agents -- monitoring, analysis, triage. Build trust before giving write access
  2. One agent per domain -- do not build a "do everything" agent. Specialized agents with narrow tool access are more reliable
  3. Human-in-the-loop for production changes -- agents propose PRs, humans merge them
  4. Session persistence for long tasks -- refactoring, migrations, and large test suites benefit from hours-long sessions

Pricing ​

Cost ItemPriceNotes
Token usageStandard Claude API ratesSame as calling the API directly
Session runtime$0.08/hourBilled by the millisecond
Idle timeNot billedWaiting for user input or tool responses
Web search$10/1,000 queriesIf the agent uses web search

Cost estimates for typical tasks:

  • Quick analysis (5 min): ~$0.01 (runtime) + token costs
  • Module refactoring (1 hour): ~$0.08 (runtime) + token costs
  • Large migration (8 hours): ~$0.64 (runtime) + token costs

The session-hour cost is negligible compared to token costs for most workloads.


Exercise: End-to-End Managed Agent Workflow ​

Build a Managed Agent that:

  1. Clones a GitHub repo (or creates a mock project with intentional test failures)
  2. Runs the existing test suite
  3. Identifies failing tests and analyzes root causes
  4. Fixes the code (not the tests)
  5. Re-runs the test suite to confirm all tests pass
  6. Generates a fix report explaining each change

Use self-evaluation with these criteria:

  • All previously passing tests still pass (no regressions)
  • Previously failing tests now pass
  • The fix report explains the root cause and fix for each failure
  • No new files created unless strictly necessary

Chapter Summary ​

  • Managed Agents moves agent infrastructure to Anthropic's cloud -- you define the task, they handle execution, state management, and crash recovery
  • Three-layer virtualization (Session/Harness/Sandbox) means each component can fail or be replaced independently
  • Session persistence lets you disconnect and reconnect without losing work -- the event log is append-only and never lost
  • Self-evaluating agents iterate until success criteria pass, improving pass@1 by up to 10 percentage points
  • Enterprise adoption follows a pattern: start read-only, specialize per domain, human-in-the-loop for production writes

Appendix links: A15 Claude Code Internals explains the harness loop, session management, and compaction mechanics that Managed Agents builds on. A13 Multi-Agent Coordination covers consensus, conflict resolution, and scaling patterns for multi-agent deployments.

Research reference: Anthropic's "Building effective agents" blog post (2024) is the foundational document for the design principles behind Managed Agents. Key takeaways: prefer simple, composable tools over complex workflows; let the model drive control flow; and keep the harness thin so it does not encode assumptions that go stale as models improve.


Knowledge Check ​

What is the key advantage of the three-layer architecture (Session, Harness, Sandbox) over a monolithic container design?
It is cheaper because each layer uses fewer resources
Each layer can fail or be replaced independently -- a harness crash does not lose the session log, and a new sandbox can be spun up without restarting
It allows running multiple models simultaneously
The three layers are required by the Anthropic API and cannot be changed
A self-evaluating agent fails 3 of 13 criteria on its first iteration. What should it do next?
Report the failures and stop, letting the user fix them
Fix only the 3 failing criteria and declare success
Fix the 3 failing criteria, then re-evaluate ALL 13 criteria from scratch
Start over from scratch with a completely new implementation
You disconnect from a Managed Agent session after 60 seconds. What happens to the agent?
The agent pauses and waits for you to reconnect
The agent keeps running in the cloud and logs all events to the session
The agent terminates after a 5-minute grace period
The agent restarts from the beginning when you reconnect

Next: Chapter 15: Production Workflow Design -- the capstone chapter that brings everything together.

Released under MIT License