Skip to content

A04: Agent Architecture Patterns ​

Related chapters: Ch8 Extended Thinking & Complex Workflows, Ch9 Multi-Agent Orchestration, Ch15 Production Workflows

Claude Code is not a chatbot with extra features. It is an autonomous agent -- a system that observes its environment, reasons about what to do, takes actions, and observes the results in a loop. Understanding the architectural patterns behind AI agents helps you predict Claude Code's behavior, design tasks that align with its strengths, and debug failures when things go wrong.

This appendix covers the major agent architecture patterns used in modern AI systems, with a focus on how Claude Code implements and combines them. The patterns are drawn from both academic research and Anthropic's own published guidance on building effective agents.

The Agentic Loop: The Foundation ​

Every agent architecture is built on a single primitive: the observe-think-act loop.

The fundamental agentic loop:
┌──────────────────────────────────────┐
│                                      │
│    ┌──────────┐                      │
│    │ Observe  │  Read environment    │
│    │          │  (tool results,      │
│    │          │   user input)        │
│    └────┬─────┘                      │
│         │                            │
│         ▼                            │
│    ┌──────────┐                      │
│    │  Think   │  Reason about what   │
│    │          │  to do next          │
│    │          │  (next token         │
│    │          │   prediction)        │
│    └────┬─────┘                      │
│         │                            │
│         ▼                            │
│    ┌──────────┐                      │
│    │   Act    │  Call a tool or      │
│    │          │  respond to user     │
│    └────┬─────┘                      │
│         │                            │
│         └────────── loop ────────────┘
│                                      │
└──────────────────────────────────────┘

In Claude Code, this loop is managed by the harness -- the runtime that wraps the Claude API. The harness:

  1. Sends the full context (system prompt, CLAUDE.md, history, your prompt) to the Claude API
  2. Receives Claude's response, which may include tool calls
  3. Executes the tool calls (reads files, runs commands, etc.)
  4. Appends the tool results to the context
  5. Sends the updated context back to Claude
  6. Repeats until Claude produces a response with no tool calls (the "final answer")

Each iteration of this loop is called a turn. A single user prompt may trigger many turns as Claude reads files, runs commands, and iterates on its approach.

Pattern 1: ReAct (Reason + Act) ​

The Pattern ​

ReAct (Yao et al., 2022) is the foundational pattern for tool-using LLM agents. It interleaves reasoning traces with action execution:

ReAct loop:

User: "Fix the failing test in src/auth/"

Thought: The user wants me to fix a failing test. I should first 
         find which test is failing by running the test suite.
Action:  Bash("cd src/auth && npm test")
Observation: FAIL src/auth/__tests__/middleware.test.ts
             ● validateToken › should reject expired tokens
             Expected: 401
             Received: 200

Thought: The test expects a 401 for expired tokens but gets 200.
         The issue is likely in the validateToken function. Let me 
         read the middleware to understand the current logic.
Action:  Read("src/auth/middleware.ts")
Observation: [file contents showing validateToken function]

Thought: I see the issue. Line 24 checks `token.exp > Date.now()` 
         but Date.now() returns milliseconds while token.exp is in 
         seconds. The comparison is wrong.
Action:  Edit("src/auth/middleware.ts", fix the comparison)
Observation: File edited successfully.

Thought: The fix is in place. Let me verify by running the tests again.
Action:  Bash("cd src/auth && npm test")
Observation: PASS all tests

Response: "Fixed the bug. The issue was a unit mismatch..."

Why ReAct Works ​

ReAct succeeds because it combines two complementary strengths:

  1. LLM reasoning handles ambiguity, planning, and interpretation -- things that are hard to encode in deterministic programs.
  2. Tool execution provides ground truth -- reading actual files instead of guessing, running actual tests instead of predicting outcomes.

The reasoning traces also serve as an implicit chain-of-thought, improving the quality of subsequent decisions. Each "Thought" step forces the model to articulate its understanding before acting.

ReAct in Claude Code ​

Claude Code's core loop is a ReAct implementation. When you give Claude a task, it alternates between:

  • Reasoning (the text it generates between tool calls, including extended thinking when enabled)
  • Acting (the tool calls it makes: Read, Edit, Bash, etc.)
  • Observing (processing the tool results)

You can see this pattern directly in Claude Code's output. The text between tool calls is the reasoning trace, and the tool calls are the actions.

Strengths and Limitations ​

Strengths:

  • Natural fit for exploratory tasks (bug investigation, codebase understanding)
  • Self-correcting: observations from actions can redirect reasoning
  • Transparent: you can follow Claude's reasoning and intervene if it goes off track

Limitations:

  • Greedy: ReAct decides the next action based on local reasoning, not a global plan. It may explore dead ends.
  • Token-expensive: each reasoning trace consumes output tokens; long chains of thought before action can burn through context.
  • Halting problem: without external limits, a ReAct agent can loop indefinitely (Claude Code enforces turn limits to prevent this).

Pattern 2: Plan-and-Execute ​

The Pattern ​

Plan-and-execute separates planning from execution. Instead of interleaving reasoning and action at each step, the agent first creates a comprehensive plan, then executes it step by step.

Plan-and-execute:

User: "Refactor the auth module to use the repository pattern"

Phase 1 — Plan:
  1. Read current auth module structure
  2. Identify data access patterns in auth code
  3. Design repository interface
  4. Create AuthRepository class
  5. Refactor auth service to use repository
  6. Update tests
  7. Run tests to verify

Phase 2 — Execute:
  Step 1: Read src/auth/ directory... done
  Step 2: Found 3 direct database calls... done
  Step 3: Designed IAuthRepository interface... done
  Step 4: Created src/auth/repository.ts... done
  Step 5: Refactored src/auth/service.ts... done
  Step 6: Updated tests... done
  Step 7: All tests pass... done

Plan-and-Execute in Claude Code ​

Claude Code uses plan-and-execute when:

  • You ask for a complex, multi-step task
  • You explicitly request a plan ("plan this before implementing")
  • Extended thinking is enabled (the thinking step often produces an implicit plan)

The R-P-E-R-S workflow from Ch15 (Read-Plan-Execute-Review-Summarize) is a formalized plan-and-execute pattern:

R-P-E-R-S as plan-and-execute:

R (Read):    Gather information by reading files and understanding context
P (Plan):    Create a plan based on what was read
E (Execute): Implement the plan, one step at a time
R (Review):  Verify the implementation by running tests
S (Summarize): Document what was done and why

When to Use Plan-and-Execute ​

Use plan-and-execute (by explicitly requesting a plan) when:

  • The task has many steps with dependencies between them
  • Mistakes are expensive (refactoring production code)
  • You want to review the approach before Claude starts changing files
  • The task requires coordinated changes across multiple files

Avoid rigid plans when:

  • The task is exploratory (bug investigation) -- you do not know the steps in advance
  • The task is simple (single file edit) -- planning overhead is wasted
  • Information will emerge during execution that changes the plan

Strengths and Limitations ​

Strengths:

  • Reduces wasted exploration by thinking ahead
  • Gives you a chance to review and modify the plan before execution
  • Creates a natural checkpoint structure for long tasks

Limitations:

  • Plans become stale: as execution reveals new information, the original plan may be wrong
  • Over-planning: for simple tasks, the planning overhead costs more tokens than it saves
  • Rigidity: an agent following a plan may ignore better opportunities discovered during execution

Pattern 3: Reflection and Self-Correction ​

The Pattern ​

Reflection adds a self-evaluation step to the agentic loop. After taking an action, the agent explicitly evaluates whether the result is satisfactory and what to do next.

Standard loop:    Observe → Think → Act → Observe → Think → Act
Reflection loop:  Observe → Think → Act → Reflect → (retry or continue)

Reflection in Claude Code ​

Claude Code exhibits reflection behavior when:

  • A tool call returns an error and Claude analyzes the error before retrying
  • Tests fail and Claude reads the failure output to diagnose the issue
  • Claude notices that its edit did not produce the expected result and tries a different approach
Example of reflection in action:

Action:  Edit("src/auth/middleware.ts", incorrect edit)
Observation: Error: old_string not found in file

Reflection: "The file content has changed since I last read it. 
            Let me re-read the file to get the current content."

Action:  Read("src/auth/middleware.ts")
Observation: [updated file contents]

Reflection: "I see -- the function was already refactored. My edit 
            target was based on the old version. Let me adjust."

Action:  Edit("src/auth/middleware.ts", corrected edit)
Observation: File edited successfully.

You can encourage more explicit reflection by structuring your prompts:

After each major change:
1. Re-read the modified file to verify the edit is correct
2. Run the relevant tests
3. If tests fail, analyze the failure before making more changes

Self-Correction Limits ​

Claude Code's self-correction has practical limits:

  • Diminishing returns: After 2-3 failed attempts at the same task, Claude's approach tends to degrade rather than improve. It may start making increasingly random changes.
  • Sunk cost bias: Claude sometimes commits to a failing approach rather than stepping back and trying something fundamentally different.
  • Context cost: Each retry cycle consumes context. Reflection on failures adds tokens that accelerate context rot.

Practical advice: If Claude has failed at the same thing 3 times, intervene. Provide a hint, narrow the problem, or start a fresh session with lessons learned.

Pattern 4: Multi-Agent Patterns ​

Why Multiple Agents? ​

Single-agent systems (one Claude instance handling everything) work well for focused tasks. But they struggle with:

  • Task breadth: A single context window cannot hold all information for a large project
  • Parallelism: A single agent executes sequentially; multiple agents can work in parallel
  • Specialization: Different subtasks may benefit from different instructions or models

Multi-agent patterns address these by distributing work across multiple Claude instances.

Fan-Out Pattern ​

Fan-out assigns independent subtasks to separate agents working in parallel:

Fan-out pattern:

Orchestrator agent:
  "Implement these 5 features"
  ├── Agent 1: "Implement user registration"
  ├── Agent 2: "Implement login endpoint"
  ├── Agent 3: "Implement password reset"
  ├── Agent 4: "Implement session management"
  └── Agent 5: "Implement audit logging"

Each agent works independently with its own context window.
Results are collected by the orchestrator.

In Claude Code, fan-out is achieved through the claude --task command (headless mode) launched from a parent Claude session:

bash
# Parent orchestrator launches sub-agents
claude --task "Implement user registration in src/auth/register.ts" &
claude --task "Implement login in src/auth/login.ts" &
claude --task "Implement password reset in src/auth/reset.ts" &
wait

When to use fan-out: Tasks that are truly independent, where the agents will not be editing the same files. Feature implementations in separate modules, independent test suites, multi-service deployments.

Pipeline Pattern ​

Pipeline passes work through a sequence of specialized agents:

Pipeline pattern:

Agent 1 (Planner):
  Input: "Add authentication to the API"
  Output: Detailed implementation plan

Agent 2 (Implementer):
  Input: Plan from Agent 1
  Output: Code changes

Agent 3 (Reviewer):
  Input: Code changes from Agent 2
  Output: Review findings

Agent 4 (Fixer):
  Input: Review findings from Agent 3
  Output: Fixed code

Pipeline is useful when different stages require different "mindsets" or when you want explicit handoff points where a human can review.

Debate Pattern ​

Debate has multiple agents propose solutions, then argue for their approaches:

Debate pattern:

Agent A: "I'd solve this with a SQL migration that adds a new column"
Agent B: "I'd solve this with a separate lookup table"

Agent C (Judge): "Agent B's approach is better because it avoids 
                  locking the main table during migration. However, 
                  Agent A correctly identified that the column approach 
                  is simpler for queries. Recommended: use the lookup 
                  table with a denormalized cache column."

Debate is most useful for architectural decisions where there are genuine trade-offs and no single "correct" answer. In practice, this is implemented by running multiple Claude sessions with different prompts and having a final session synthesize the results.

Consensus Pattern ​

Consensus runs the same task through multiple agents independently and takes the majority answer:

Consensus pattern (e.g., code review):

Agent 1: "I found 3 issues: X, Y, Z"
Agent 2: "I found 2 issues: X, Z"
Agent 3: "I found 4 issues: X, Y, Z, W"

Consensus: Issues X and Z are confirmed (3/3 agents).
           Issue Y is likely real (2/3 agents).
           Issue W needs human review (1/3 agents).

This is the basis of the /review skill's "ultra" mode, which runs multiple review passes and synthesizes findings.

State Management in Agent Systems ​

The State Problem ​

Agents need to track state: what have they done, what remains, what constraints apply. In LLM-based agents, state management is challenging because:

  1. The only state is the context window: Claude does not have external memory during a session. Everything it "knows" must be in the current context.
  2. State is implicit: The conversation history IS the state. There is no separate state object.
  3. State degrades: As described in A02 Context Engineering, context rot degrades the agent's awareness of its own state.

State Management Strategies ​

Strategy 1: Externalize state to files

You: "Before starting each major step, write your current progress 
to a file called PROGRESS.md. Include: completed steps, current step, 
remaining steps, and any issues found."

This creates a persistent state artifact that survives compaction and can be re-read if Claude loses track.

Strategy 2: Use structured checkpoints

You: "We're implementing 5 API endpoints. After each one, output a 
status table:

| Endpoint | Status | Tests |
|----------|--------|-------|
| /users   | Done   | Pass  |
| /auth    | Done   | Pass  |
| /posts   | WIP    | -     |
| /comments| TODO   | -     |
| /likes   | TODO   | -     |"

The structured format is easier for Claude to parse in later turns than prose summaries.

Strategy 3: Leverage git as state

You: "Commit after each completed feature. Use descriptive commit 
messages. If you need to understand what's been done, run git log."

Git commits create an external state record that is immune to context rot.

Failure Recovery and Retry Strategies ​

Types of Agent Failures ​

Failure TypeExampleRecovery
Tool errorFile not found, command failedRe-read, adjust path, retry
Logic errorWrong approach, incorrect analysisReflect, backtrack, try alternative
Context overflowLost track of requirementsSummarize, compact, or new session
Infinite loopRepeating the same failed actionIntervention required
Cascading errorEarly mistake compounds through later stepsRevert to last known good state

Recovery Strategies ​

Retry with adjustment (for transient tool errors):

Claude: [Bash] npm test
Error: ECONNREFUSED - database not running

Claude: "The database isn't running. Let me start it first."
Claude: [Bash] docker compose up -d postgres
Claude: [Bash] npm test
Result: All tests pass

Backtrack and replan (for logic errors):

You: "That approach isn't working. Stop, explain what went wrong, 
and propose a completely different approach. Don't modify the 
previous approach -- think of an alternative."

Checkpoint and rollback (for cascading errors):

You: "The last 3 changes made things worse. Run `git stash` to 
undo them, then re-read the original code and try a simpler 
approach."

Fresh start (for context overflow):

You: "This session has gotten too complex. Summarize the current 
state: what works, what doesn't, what's left. I'll start a new 
session with your summary."

When to Use Single-Agent vs Multi-Agent ​

Decision Framework ​

Should I use multiple agents?

Is the task decomposable into independent subtasks?
├── No → Single agent
└── Yes
    ├── Do subtasks share many files?
    │   ├── Yes → Single agent (or pipeline)
    │   └── No → Fan-out
    ├── Is the task complex enough to benefit from specialization?
    │   ├── No → Single agent
    │   └── Yes → Pipeline
    └── Do you need high confidence in the result?
        ├── No → Single agent
        └── Yes → Consensus

Practical Guidelines ​

Use single agent when:

  • Task involves fewer than 10 files
  • Changes are interdependent (editing a function and its callers)
  • Task fits comfortably in one context window
  • Speed matters more than thoroughness

Use multi-agent when:

  • Task spans multiple modules or services
  • Subtasks are truly independent (no shared file edits)
  • You need parallel execution for speed
  • You want multiple perspectives (review, architecture decisions)
  • The task exceeds a single context window's capacity

The Overhead Trade-Off ​

Multi-agent systems add overhead:

  • Coordination cost: orchestrating agents requires prompts, context, and logic
  • Integration risk: merging work from multiple agents can create conflicts
  • Debugging complexity: tracing a bug through multiple agent sessions is harder
  • Token cost: each agent gets its own context window, so total token usage increases

The rule of thumb from Anthropic's "Building effective agents" blog post (2024): start with single agent, add agents only when you hit a specific limitation. Do not use multi-agent because it sounds impressive. Use it because a single agent demonstrably cannot handle the task.

Reference: Anthropic's Agent Design Principles ​

Anthropic published a blog post titled "Building effective agents" that outlines key principles. Here are the most relevant for Claude Code users:

  1. Keep agent architecture simple: The most effective agents use the simplest architecture that works. A single ReAct loop with good tools outperforms a complex multi-agent system for most tasks.

  2. Tools are better than prompts for grounding: When an agent needs factual information, it should use a tool (read a file, run a command) rather than rely on its training data. This is why Claude Code's tool-calling approach is more reliable than a chat-based AI for software engineering.

  3. Explicit boundaries improve reliability: Clear constraints ("do not modify files outside src/auth/") produce better results than open-ended instructions ("be careful with other files").

  4. Humans in the loop at critical points: The most reliable agent systems include human checkpoints before irreversible actions (deployments, database migrations, public API changes). Claude Code's permission system is an implementation of this principle.

  5. Fail fast, recover clearly: Agents should detect failures quickly and report them clearly, rather than attempting increasingly desperate fixes. This is why you should intervene when Claude has failed at the same thing 3 times.

Key Takeaways ​

  1. Claude Code is a ReAct agent at its core. Understanding the observe-think-act loop helps you predict its behavior and design better prompts.
  2. Plan-and-execute shines for complex, multi-step tasks. Request a plan explicitly when the task has many interdependent steps.
  3. Reflection enables self-correction, but it has diminishing returns. Intervene after 2-3 failed attempts rather than hoping for eventual success.
  4. Multi-agent is a tool, not a default. Start with single agent. Add agents only when you hit a specific limitation.
  5. State management is the hardest problem in agentic systems. Externalize state to files or git commits rather than relying on conversation history alone.
  6. Failure recovery requires matching the strategy to the failure type. Retry for transient errors, backtrack for logic errors, fresh start for context overflow.

See also: A05 Tool Calling Internals for how the action step of the agentic loop actually works, and A13 Multi-Agent Coordination for advanced multi-agent patterns.

Released under MIT License