Skip to content

Chapter 9: Agent Teams — Parallel Workflows That Actually Ship ​

What You'll Build ​

Two real parallel workflows from production codebases:

  1. Autoplan review pipeline — Three agents review from product, design, and engineering perspectives, then synthesize (from gstack's CEO/Design/Eng pipeline)
  2. Parallel feature development — Frontend + backend + tests built simultaneously by separate agents, with dependency-aware orchestration

These patterns come from gstack (68k stars) and ECC's dmux-workflows. They're how real teams use Claude Code to parallelize development.

Appendix links: A13 Multi-Agent Coordination covers consensus protocols, conflict resolution strategies, and scaling patterns for teams of 5+ agents.

Research note: The orchestrator-worker and parallelization patterns in this chapter are formalized in Anthropic's Building effective agents guide. That guide identifies five key multi-agent patterns; this chapter focuses on two of them: parallel delegation (Demo 24) and dependency-aware orchestration (Demo 25).

How Agent Teams Work ​

In Chapter 8, you used subagents — single agents spawned for specific tasks, returning results to the main session. Agent Teams take this further: multiple agents running in parallel, each with their own session, able to communicate with each other.

Lead (your Claude session)
  |
  +---> Agent A (Frontend)    --|
  |       Own session           |  Running in parallel
  |       Own 200K context      |  Each has full independence
  |                             |
  +---> Agent B (Backend)     --|
  |       Own session           |
  |       Own 200K context      |
  |                             |
  +---> Agent C (Testing)     --|  Waits for A and B
          Own session
          Own 200K context
  |
  <--- Lead collects all results, synthesizes

Teams vs Subagents ​

Subagents (Ch8)Agent Teams (Ch9)
SessionsShare main session's processEach teammate gets its own session
CommunicationReturn results to main onlyTeammates can message each other
LifecycleOne task, then doneCan pick up multiple tasks
Best forFocused analysis tasksParallel development work
DisplayInline in main sessiontmux panes or interleaved output

GSD's Wave Execution Model ​

GSD (50k stars) formalizes this as "Wave Execution": parallel tasks run in fresh 200K contexts. Each agent gets a clean context window, avoiding the "Context Rot" problem where quality degrades past 50% context usage.

The key insight: a fresh 200K context produces better code than a polluted 150K context, even if the fresh one has less history. This is why parallel agents outperform sequential work in a single session for large tasks.

Communication: The Mailbox System ​

Teammates communicate through a Mailbox:

Frontend Agent --> Backend Agent: "What's the response format for POST /api/tasks?"
Backend Agent --> Frontend Agent: "{ id: string, title: string, status: 'pending' | 'done', createdAt: ISO8601 }"
Frontend Agent: "Got it, building the type interface now"

This happens automatically when an agent needs information from another. The lead (your session) can also send messages to any teammate.

Task Dependencies ​

Tasks can have dependencies:

Task 1: Build API endpoints        (no deps)     --> Agent B
Task 2: Build UI components         (no deps)     --> Agent A
Task 3: Write integration tests     (needs 1 + 2) --> Agent C (waits)
Task 4: Write deployment config     (needs 3)     --> Agent A (after)

The system automatically runs 1 and 2 in parallel, starts 3 when both finish, and queues 4 after 3.


Demo 24: Autoplan Review Pipeline ​

24
Multi-Perspective Review Pipeline
Intermediate~20 min

Adapted from gstack's CEO -> Design -> Eng review pipeline and "Boil the Lake" completeness methodology.

The Problem ​

A feature spec or PR lands. In a good team, it gets reviewed from multiple angles — product (is this worth building?), design (is the UX right?), and engineering (is the architecture sound?). Usually this takes 3 different people and a day of calendar time. With an agent team, you get all three perspectives in 2 minutes.

Build It ​

1. Set Up ​

bash
mkdir -p ~/claude-demos/demo-24/src && cd ~/claude-demos/demo-24
git init && git branch -M main

# A feature to review: task management API
cat > src/task-service.js << 'EOF'
import { db } from './db.js';
import { sendNotification } from './notify.js';

export class TaskService {
  async createTask(req, res) {
    const { title, description, assigneeId, priority, dueDate } = req.body;

    // No validation on priority range
    // No check if assigneeId exists
    const task = await db.query(
      `INSERT INTO tasks (title, description, assignee_id, priority, due_date, status)
       VALUES ($1, $2, $3, $4, $5, 'pending') RETURNING *`,
      [title, description, assigneeId, priority, dueDate]
    );

    // Fire and forget — if notification fails, user never knows
    sendNotification(assigneeId, `New task: ${title}`).catch(() => {});

    return res.status(201).json(task.rows[0]);
  }

  async listTasks(req, res) {
    const { status, assignee, sort } = req.query;

    // Building query dynamically — but safely with parameterized queries
    let query = 'SELECT * FROM tasks WHERE 1=1';
    const params = [];
    let paramCount = 0;

    if (status) { paramCount++; query += ` AND status = $${paramCount}`; params.push(status); }
    if (assignee) { paramCount++; query += ` AND assignee_id = $${paramCount}`; params.push(assignee); }

    // No pagination — will return ALL tasks
    query += ` ORDER BY ${sort || 'created_at'} DESC`;

    const result = await db.query(query, params);
    return res.json(result.rows);
  }

  async updateStatus(req, res) {
    const { id } = req.params;
    const { status } = req.body;

    // No validation on status values
    // No check if task exists
    // No authorization — anyone can update any task
    await db.query('UPDATE tasks SET status = $1, updated_at = NOW() WHERE id = $2', [status, id]);

    return res.json({ success: true });
  }

  async deleteTask(req, res) {
    const { id } = req.params;
    // Hard delete — no soft delete, no audit trail
    await db.query('DELETE FROM tasks WHERE id = $1', [id]);
    return res.json({ success: true });
  }
}
EOF

# Feature spec document
cat > FEATURE_SPEC.md << 'EOF'
# Feature: Task Management System

## User Story
As a team member, I want to create, assign, and track tasks so that our team can coordinate work effectively.

## Requirements
- Create tasks with title, description, assignee, priority (1-5), and due date
- List tasks with filtering by status and assignee
- Update task status (pending -> in_progress -> done)
- Delete tasks
- Notify assignees when tasks are created

## Technical Notes
- REST API using Express
- PostgreSQL database
- Notification via internal service
EOF

cat > CLAUDE.md << 'EOF'
# Task Management API
- Node.js + Express + PostgreSQL
- Parameterized queries for SQL
- RESTful conventions
EOF

git add -A && git commit -m "feat: task management API and feature spec"

2. Create the Three Review Agents ​

bash
mkdir -p .claude/agents

# Product (CEO) Review Agent
cat > .claude/agents/product-reviewer.md << 'EOF'
---
name: product-reviewer
description: Reviews from a product/business perspective — is this worth building?
tools:
  - Read
  - Glob
  - Grep
model: claude-sonnet-4-6
---

You are a product manager reviewing a feature implementation. Your job is to evaluate whether this feature delivers real value and is ready to ship to users.

## Review Criteria

### Value Assessment
- Does the implementation match what users actually need?
- Are there missing capabilities that would make this useless? (e.g., can create tasks but can't set due dates)
- What's the minimum viable version of this feature?

### User Experience Gaps
- What happens when things go wrong? (error messages, edge cases)
- Is the API intuitive? Would a developer integrating this be confused?
- Are there missing features that users would immediately ask for?

### Business Risk
- Could this cause data loss?
- Could this cause customer-facing outages?
- Is there anything that would block a launch?

## Output Format

Product Review ​

Verdict: SHIP / ITERATE / BLOCK ​

Completeness Rating: X/10 ​

(gstack's "Boil the Lake" score — how complete is this feature?)

What Works ​

  • [Things the implementation gets right]

Gaps ​

  • [P0] [Missing capability that blocks ship]
  • [P1] [Missing capability that hurts adoption]
  • [P2] [Nice-to-have for v2]

User Impact Assessment ​

  • Who benefits from this feature?
  • What workflow does it enable?
  • What's missing for the workflow to actually work end-to-end?
EOF

# Design Review Agent
cat > .claude/agents/design-reviewer.md << 'EOF'
---
name: design-reviewer
description: Reviews API design, UX patterns, and developer experience
tools:
  - Read
  - Glob
  - Grep
model: claude-sonnet-4-6
---

You are a senior API designer reviewing an implementation for usability, consistency, and developer experience.

## Review Criteria

### API Design
- Are endpoints RESTful and consistent?
- Are HTTP status codes correct? (201 for create, 404 for not found, etc.)
- Are error responses structured and helpful?
- Is the request/response format consistent across endpoints?

### Developer Experience
- Could someone use this API without reading the source code?
- Are required vs optional fields clear?
- Is pagination supported for list endpoints?
- Are there reasonable defaults?

### Consistency
- Do naming conventions match across endpoints? (camelCase vs snake_case)
- Are similar operations handled the same way?
- Does the API follow the principle of least surprise?

## Output Format

Design Review ​

Verdict: APPROVE / REQUEST CHANGES ​

API Consistency Score: X/10 ​

Findings ​

  • [P1] [Issue with specific recommendation]
  • [P2] [Issue with specific recommendation]

Suggested API Contract ​

[For any endpoint with issues, show the correct request/response format]

EOF

# Engineering Review Agent
cat > .claude/agents/eng-reviewer.md << 'EOF'
---
name: eng-reviewer
description: Reviews architecture, security, performance, and code quality
tools:
  - Read
  - Glob
  - Grep
model: claude-sonnet-4-6
---

You are a senior engineer reviewing code for production readiness. Focus on things that cause incidents.

## Review Criteria

### Security (P0)
- SQL injection (even with parameterized queries, check for string interpolation)
- Missing authentication or authorization
- Missing input validation
- Data exposure in responses

### Reliability (P1)
- Missing error handling
- Silent failures (catch-and-ignore patterns)
- Race conditions
- Missing transactions for multi-step operations

### Performance (P1)
- N+1 queries
- Missing pagination
- Missing caching for hot paths
- Unbounded query results

### Code Quality (P2)
- Missing input validation
- Hardcoded values that should be configurable
- Missing logging/observability
- Dead code or unreachable branches

## Output Format

Engineering Review ​

Verdict: APPROVE / REQUEST CHANGES / BLOCK ​

Findings ​

[P0] [Title] ​

  • File: path:line
  • Confidence: N%
  • Problem: [What's wrong]
  • Fix: [Exact code change]

[P1] [Title] ​

...

Architecture Notes ​

[Any structural concerns or recommendations]

Effort Estimate ​

  • Critical fixes: ~Xh
  • Full recommendations: ~Xh (gstack's human-vs-CC effort comparison)
EOF

3. Run the Pipeline ​

bash
claude
I need a comprehensive review of the task management feature. 
Run product-reviewer, design-reviewer, and eng-reviewer agents in parallel on 
src/task-service.js and FEATURE_SPEC.md, then synthesize their findings into 
a single actionable review.

4. Expected Terminal Output ​

terminal
$ claude
> I need a comprehensive review of the task management feature. Run
  product-reviewer, design-reviewer, and eng-reviewer agents in parallel on
  src/task-service.js and FEATURE_SPEC.md, then synthesize their findings
  into a single actionable review.

Claude: I'll launch all three review agents in parallel, then synthesize
their findings.

  Spawning 3 agents in parallel...
  
  [product-reviewer] Starting (claude-sonnet-4-6, read-only)
  [design-reviewer]  Starting (claude-sonnet-4-6, read-only)
  [eng-reviewer]     Starting (claude-sonnet-4-6, read-only)
  
  [product-reviewer] Read: FEATURE_SPEC.md
  [design-reviewer]  Read: src/task-service.js
  [eng-reviewer]     Read: src/task-service.js
  [product-reviewer] Read: src/task-service.js
  [design-reviewer]  Read: FEATURE_SPEC.md
  [eng-reviewer]     Grep: 'req.body' in src/ -> 3 matches
  [eng-reviewer]     Grep: 'req.params' in src/ -> 2 matches
  [eng-reviewer]     Grep: 'req.query' in src/ -> 1 match
  
  [product-reviewer] Complete. (2.1s)
  [design-reviewer]  Complete. (1.8s)
  [eng-reviewer]     Complete. (2.4s)
  
  All 3 agents finished. Synthesizing results...

---

# Synthesized Review

## Overall Verdict: ITERATE (do not ship as-is)

## Product Review (Verdict: ITERATE, Completeness: 6/10)
### What Works
- Core CRUD operations are implemented
- Parameterized SQL queries (secure foundation)
- Notification on task creation

### Gaps
- [P0] No way to list MY tasks specifically -- assignee filter exists
  but no auth context to default to current user
- [P1] No task reassignment -- once assigned, stuck with that person
- [P1] deleteTask has no confirmation or undo -- data loss risk
- [P2] No due date reminders or overdue indicators

## Design Review (Verdict: REQUEST CHANGES, Consistency: 4/10)
### Findings
- [P1] Inconsistent responses: createTask returns the full task object,
  updateStatus returns {success: true}, deleteTask returns {success: true}.
  All mutations should return the affected resource.
- [P1] No pagination on listTasks -- will break with >1000 tasks.
  Add ?page=1&limit=20 with Link headers.
- [P2] Mixed naming: assignee_id (snake_case in DB) vs assigneeId
  (camelCase in API). Pick one convention for the API layer.

## Engineering Review (Verdict: BLOCK, Critical fixes: ~3h)
### [P0] SQL Injection via sort parameter (95% confidence)
- File: src/task-service.js:28
- The sort query parameter is interpolated directly into SQL:
  ORDER BY ${sort || 'created_at'} DESC
- PoC: GET /api/tasks?sort=created_at;DROP TABLE tasks;--
- Fix: Whitelist allowed sort columns

### [P1] Missing authorization on updateStatus (90% confidence)
- File: src/task-service.js:36
- Any authenticated user can update any task status
- No ownership or role check

### [P1] Silent notification failure (88% confidence)
- File: src/task-service.js:16
- sendNotification failure is caught and discarded
- Users are never told their notification was not delivered

## Priority Action Plan
1. FIX NOW: Whitelist sort columns to prevent SQL injection (~30min)
2. FIX NOW: Add authorization check to updateStatus (~1h)
3. Before ship: Add pagination to listTasks (~2h)
4. Before ship: Make response format consistent across all endpoints (~1h)
5. Next iteration: Add task reassignment and due date reminders

Total critical effort: ~3.5h (estimated ~1 day without Claude Code)

What Just Happened? ​

1
TaskAgent (x3)
product-reviewer, design-reviewer, eng-reviewer
↓
2
Read (parallel)
src/task-service.js + FEATURE_SPEC.md
↓
3
Grep (eng only)
User input patterns
↓
4
Synthesize
Lead session

The three agents ran in parallel, finishing in ~2.4 seconds total (not 2.4 x 3 = 7.2 seconds). Each found different issues from its perspective: the product reviewer caught the missing "my tasks" flow, the design reviewer caught inconsistent response formats, and the engineering reviewer found the SQL injection. No single agent would have caught all three categories.

Why This Pattern Works ​

This is gstack's "Boil the Lake" approach: review from every angle until the completeness rating hits 9+/10. Each reviewer catches different things:

  • Product: "This feature is incomplete — no task assignment notification actually works"
  • Design: "The API is inconsistent — createTask returns the object, updateStatus returns {success: true}"
  • Engineering: "The sort parameter is injected directly into SQL — that's a P0"

No single reviewer would catch all three. The parallel execution means you get all perspectives in the time it takes to run one.


Demo 25: Parallel Feature Development ​

25
Parallel Feature Development with Dependency Orchestration
Intermediate~25 min

Adapted from ECC's dmux-workflows and GSD's Wave Execution pattern.

The Problem ​

You need to build a feature that spans frontend, backend, and tests. Doing it sequentially in one Claude session means the context fills up fast, and quality drops as you switch between concerns. Parallel agents each get a fresh 200K context and focus on one thing.

Build It ​

1. Set Up ​

bash
mkdir -p ~/claude-demos/demo-25/{src/api,src/pages,tests} && cd ~/claude-demos/demo-25
git init && git branch -M main

cat > CLAUDE.md << 'EOF'
# Bookmark Manager

## Tech Stack
- Backend: Node.js + Express + PostgreSQL
- Frontend: Vanilla HTML + JavaScript (no framework)
- Testing: Vitest

## Feature: Bookmark CRUD
Users can save, list, tag, and search bookmarks.

## API Contract
POST /api/bookmarks
  Body: { url: string, title: string, tags: string[] }
  Response: { id, url, title, tags, createdAt }

GET /api/bookmarks?tag=&search=
  Response: { bookmarks: [...], total: number }

DELETE /api/bookmarks/:id
  Response: { success: true }

## Coding Standards
- Use parameterized SQL queries
- Validate all input
- Return appropriate HTTP status codes
- Tests must mock the database layer
EOF

cat > src/api/index.js << 'EOF'
import express from 'express';
const app = express();
app.use(express.json());
app.use(express.static('src/pages'));
// Routes will be added by the backend agent
export default app;
EOF

cat > package.json << 'EOF'
{
  "name": "bookmark-manager",
  "version": "1.0.0",
  "type": "module",
  "scripts": {
    "start": "node src/api/index.js",
    "test": "vitest run"
  },
  "dependencies": {
    "express": "^4.18.0"
  },
  "devDependencies": {
    "vitest": "^1.0.0"
  }
}
EOF

git add -A && git commit -m "feat: project skeleton with API contract"

2. Launch Parallel Development ​

bash
claude
Build the bookmark manager feature using parallel agents:

1. Backend Agent: Implement the API endpoints in src/api/bookmarks.js
   following the contract in CLAUDE.md. Wire up routes in src/api/index.js.
   Focus on: input validation, parameterized queries, proper error handling,
   correct status codes.

2. Frontend Agent: Build src/pages/index.html — a single-page app with:
   - Form to add bookmarks (URL, title, tags)
   - List of bookmarks with tag filtering
   - Search box
   - Delete button per bookmark
   - Clean, functional design (no framework needed, vanilla JS is fine)

3. Test Agent: AFTER the backend agent finishes, write tests in 
   tests/bookmarks.test.js covering:
   - Happy path for each endpoint
   - Validation failures (missing URL, invalid input)
   - Edge cases (duplicate bookmarks, empty search results)
   Mock the database layer.

Run the backend and frontend agents in parallel. Start the test agent 
after the backend finishes.

3. Expected Terminal Output ​

terminal
$ claude
> Build the bookmark manager feature using parallel agents...

Claude: I'll orchestrate three agents with dependency-aware scheduling:
  - Backend and Frontend run in parallel (no dependencies)
  - Test agent starts after Backend completes (needs the API code to test)

  Task graph:
    [Backend]  ----+---> [Test Agent]
    [Frontend] ----+

  Spawning Backend agent...
  Spawning Frontend agent...

  [Backend]  Read: CLAUDE.md (API contract)
  [Frontend] Read: CLAUDE.md (API contract)
  [Backend]  Read: src/api/index.js (entry point)
  [Frontend] Write: src/pages/index.html (creating SPA)
  [Backend]  Write: src/api/bookmarks.js (3 endpoints)
  [Frontend] Adding fetch calls for POST, GET, DELETE...
  [Backend]  Edit: src/api/index.js (wiring routes)
  [Frontend] Adding tag filter dropdown and search input...
  [Backend]  Complete. 3 endpoints with validation. (3.2s)
  
  Backend finished. Spawning Test agent...

  [Frontend] Adding delete button with confirmation...
  [Test]     Read: src/api/bookmarks.js (understanding the API)
  [Test]     Read: package.json (detecting vitest)
  [Frontend] Complete. Single-page app with search and filtering. (4.1s)
  [Test]     Write: tests/bookmarks.test.js (14 test cases)
  [Test]     Bash: npx vitest run --reporter=verbose
  
  stdout | tests/bookmarks.test.js
    POST /api/bookmarks
      + creates bookmark with valid input (2ms)
      + returns 400 when URL is missing (1ms)
      + returns 400 when URL format is invalid (1ms)
      + stores tags as array (1ms)
      + returns 201 with created bookmark (1ms)
    GET /api/bookmarks
      + returns all bookmarks with total count (2ms)
      + filters by tag when tag parameter provided (1ms)
      + searches title and URL when search parameter provided (1ms)
      + returns empty array when no matches (1ms)
      + handles pagination parameters (1ms)
    DELETE /api/bookmarks/:id
      + deletes existing bookmark (1ms)
      + returns 404 for non-existent bookmark (1ms)
      + returns success true on deletion (1ms)
    Edge cases
      + handles special characters in search query (2ms)

  Test Files  1 passed (1)
       Tests  14 passed (14)
    Duration  0.28s

  [Test]     Complete. 14/14 tests passing. (2.8s)

  All agents finished. Total time: 6.0s (vs ~18s sequential)

  Files created:
    src/api/bookmarks.js  (API endpoints with validation)
    src/pages/index.html  (Single-page app with search/filter/delete)
    tests/bookmarks.test.js (14 test cases, all passing)
  
  Files modified:
    src/api/index.js (route registration added)

What Just Happened? ​

1
Plan
Task dependency graph
↓
2
TaskAgent (x2)
Backend + Frontend agents
↓
3
Write (Backend)
src/api/bookmarks.js + index.js
↓
4
Write (Frontend)
src/pages/index.html
↓
5
TaskAgent
Test agent (wave 2)
↓
6
Write + Bash (Test)
tests/bookmarks.test.js

The CLAUDE.md file acted as a shared contract between agents. The Backend and Frontend agents never communicated directly — they both read the API contract and built to spec. This is the "shared artifact" coordination pattern: instead of agents talking to each other, they agree on an interface document.

4. Verify ​

bash
# Check all files were created
ls src/api/bookmarks.js src/pages/index.html tests/bookmarks.test.js

# Run the tests
npm test

# Start the server and check manually
node src/api/index.js &
# Open http://localhost:3000 in a browser

How This Maps to GSD's 6-Step Process ​

GSD's process is: Initialize -> Discuss -> Plan -> Execute -> Verify -> Ship

In this demo:

  1. Initialize: CLAUDE.md defines the project and API contract
  2. Discuss: The prompt defines requirements for each agent
  3. Plan: Claude (the lead) creates the task dependency graph
  4. Execute: Three agents run in parallel (Wave Execution)
  5. Verify: Test agent runs tests, lead checks results
  6. Ship: All files created, tests passing, ready to deploy

The key GSD insight: each agent gets a fresh 200K context ("wave"), so quality stays high even as the total work grows. A single session doing all three tasks sequentially would be at 60%+ context by the time it reaches tests, producing worse output.


When Things Go Wrong ​

Agent Communication Timeout ​

terminal
$ claude
> Build a feature with frontend and backend agents. The frontend needs to
  know the exact response schema from the backend.

Claude: Launching Backend and Frontend agents in parallel...

  [Frontend] Building form component...
  [Frontend] --> Backend: What fields does POST /api/items return?
  [Frontend] Waiting for response from Backend agent...
  [Backend]  Still implementing input validation...
  [Backend]  Still writing database queries...
  [Frontend] Timeout waiting for Backend response (30s).
  [Frontend] Falling back to CLAUDE.md contract definition.
  [Frontend] Using { id, name, createdAt } from the spec.
  [Frontend] Complete.
  
  [Backend]  Complete. POST /api/items returns { id, name, description, 
             createdAt, updatedAt }
  
  Warning: Frontend built against spec contract, but Backend returns 
  additional fields (description, updatedAt) not in the original spec.

What happened: The Frontend agent asked the Backend agent for the response schema, but the Backend agent was busy and did not respond in time. The Frontend fell back to the CLAUDE.md contract, which was incomplete.

Fix: The root cause is that the API contract in CLAUDE.md was not detailed enough. The best prevention is a thorough shared contract that both agents can rely on independently, without needing to communicate at runtime. Update CLAUDE.md to include all response fields:

markdown
## API Contract
POST /api/items
  Body: { name: string, description: string }
  Response: { id: string, name: string, description: string, createdAt: ISO8601, updatedAt: ISO8601 }

When agents share a complete contract, they do not need real-time communication. This is more reliable than mailbox-based coordination.

Merge Conflict Between Parallel Agents ​

terminal
$ claude
> Have two agents both improve src/api/index.js — one adds error handling
  and one adds logging.

Claude: Launching two agents in parallel...

  [error-handler] Read: src/api/index.js
  [logger]        Read: src/api/index.js
  [error-handler] Edit: src/api/index.js (adding try-catch wrappers)
  [logger]        Edit: src/api/index.js (adding request logging middleware)
  
  Error: Merge conflict in src/api/index.js
  
  [error-handler] wrote:
    app.use((err, req, res, next) => {
      console.error(err.stack);
      res.status(500).json({ error: 'Internal server error' });
    });
  
  [logger] wrote:
    app.use((req, res, next) => {
      console.log(`${req.method} ${req.path}`);
      next();
    });
  
  Both agents modified the same region of the file.
  Lead session resolving conflict...

What happened: Two agents edited the same file simultaneously. When the second agent tried to write, the file had already been changed by the first agent, causing a conflict.

Fix: Do not assign two agents to modify the same file. Instead, structure your tasks so each agent owns distinct files:

Agent A: Create src/middleware/errorHandler.js (new file, no conflict)
Agent B: Create src/middleware/logger.js (new file, no conflict)
Lead: Wire both into src/api/index.js after both agents finish

If agents must touch the same file, make them sequential, not parallel. The dependency system handles this: set the second agent to depend on the first.

Context Overflow in Team Orchestration ​

terminal
$ claude
> Launch 8 agents to review every module in the monorepo, then synthesize
  all their findings into one report.

Claude: Launching 8 review agents in parallel...

  [agent-1] Complete. Report: 2,400 tokens.
  [agent-2] Complete. Report: 1,800 tokens.
  [agent-3] Complete. Report: 3,100 tokens.
  [agent-4] Complete. Report: 2,200 tokens.
  [agent-5] Complete. Report: 2,600 tokens.
  [agent-6] Complete. Report: 1,900 tokens.
  [agent-7] Complete. Report: 2,800 tokens.
  [agent-8] Complete. Report: 2,500 tokens.
  
  Collecting results into lead session...
  Total report content: ~19,300 tokens
  Lead session context before synthesis: 145,000 / 200,000 tokens
  
  Warning: Lead session at 82% context capacity after collecting 8 reports.
  Synthesis quality may be degraded (Context Rot threshold: 50%).
  
  The synthesized report missed findings from agents 6-8 and produced
  generic recommendations instead of specific file:line references.

What happened: Each agent returned a detailed report. Collecting all 8 reports into the lead session pushed context usage past the 50% quality threshold. The synthesis was shallow because the lead session was overloaded.

Fix: Use hierarchical synthesis. Instead of one lead collecting all reports, group agents and use intermediate summarizers:

Wave 1: 8 review agents (parallel)
Wave 2: 2 summarizer agents (each takes 4 reports, produces 500-token summary)
Wave 3: Lead collects 2 summaries (~1,000 tokens total)

Alternatively, tell each agent to produce a shorter report (under 500 tokens) by requiring only P0 and P1 findings. A focused 500-token report per agent keeps the lead session under 50% context even with 8 agents:

markdown
## Output Rules
- Maximum 500 tokens per report
- Only P0 and P1 findings
- One line per finding: [severity] file:line — description

Advanced: tmux-Based Orchestration ​

ECC's dmux-workflows uses tmux to give each agent a visible pane. On macOS:

bash
# Install tmux
brew install tmux

# Start a tmux session
tmux new -s agents

# Split into panes and run agents
tmux split-window -h
tmux split-window -v
tmux select-pane -t 0 && tmux split-window -v

# Now you have 4 panes — one lead + three agents
# Each pane runs its own `claude` process

On Windows, use Windows Terminal tabs or WSL2 with tmux.


Exercise: Build a Bug Investigation Team ​

A bug report comes in: "Users can see other users' bookmarks after logging in with a different account."

Create three investigation agents:

  1. log-analyst: Search the code for authentication and session handling. Trace how user identity flows through requests.
  2. code-tracer: Start from the GET /api/bookmarks endpoint and trace every function call. Map the data flow.
  3. exploit-writer: Based on findings from the other two, write a test that reproduces the bug.

Run log-analyst and code-tracer in parallel. Start exploit-writer after both finish. Produce a root-cause analysis report.

Success Criteria ​

  • [ ] Three agent files in .claude/agents/
  • [ ] log-analyst and code-tracer run in parallel
  • [ ] exploit-writer depends on the first two completing
  • [ ] Final report identifies the root cause with specific file:line references
  • [ ] Exploit test actually demonstrates the vulnerability

Knowledge Check ​

Why does GSD recommend fresh 200K contexts (Wave Execution) instead of one long-running session?
Fresh contexts are cheaper because they use fewer API tokens
Fresh contexts produce better output because quality degrades past 50% context usage
Long-running sessions are not supported by the Claude API
Fresh contexts allow using different models for each wave
In Demo 25, the Frontend and Backend agents never communicated directly. How did they build compatible code?
The lead session relayed messages between them in real time
They both read the API contract from CLAUDE.md and built to the same spec
The Frontend waited for the Backend to finish, then read its output
They used the Mailbox system to agree on the interface
Two parallel agents both need to modify src/api/index.js. What is the safest approach?
Let both agents edit the file and resolve conflicts automatically
Have each agent create a separate new file, then wire them together in a sequential step
Lock the file so only one agent can access it at a time
Merge the two agent contexts into one so they share the file state
You launched 8 review agents and collected all reports into the lead session. The synthesis was shallow and missed findings from later agents. What went wrong?
The later agents produced lower-quality reports
The lead session hit the Context Rot threshold after collecting too many reports
The agents interfered with each other by reading the same files
Eight agents exceeded the maximum allowed parallel sessions

Summary ​

Agent Teams turn Claude Code from a single developer into a parallel development shop. The two patterns in this chapter — multi-perspective review and parallel feature development — are the highest-ROI applications of agent teams.

Key points:

  • GSD's Wave Execution: fresh 200K contexts produce better output than polluted long-running sessions
  • gstack's multi-role review: product + design + engineering catches issues that single-perspective review misses
  • Task dependencies: parallel where possible, sequential where necessary
  • Shared contracts (CLAUDE.md) are more reliable than real-time agent communication
  • Mailbox communication lets agents coordinate when real-time data exchange is needed
  • ECC's dmux-workflows provide visual orchestration via tmux
  • Watch for context overflow when collecting many agent reports — use hierarchical synthesis

Congratulations on completing the Intermediate stage.

Next, tackle the Capstone Project: PR Review Bot to put everything from Ch5-9 together.

Or jump ahead to Stage 3: Advanced — Chapter 10: Permissions and Security.

Released under MIT License