Chapter 15: Production Workflow Design
This is the capstone chapter of the tutorial. Everything from Chapters 1 through 14 converges here into production-grade patterns for deploying Claude Code at scale.
What You Will Learn
- The R-P-E-R-S workflow used by top-performing teams across the Claude Code ecosystem
- GSD's wave execution pattern for parallel task processing
- HumanLayer's RPI pattern for large codebases (300k+ LOC)
- gstack's effort table: real productivity multipliers by task type
- When to use each framework: ECC, gstack, GSD, BMAD, HumanLayer
- Full walkthrough: adding OAuth to an Express app using the production workflow
- Cost comparison: single session vs subagents vs teams vs waves
Research reference: The R-P-E-R-S workflow documented in this chapter draws from multiple sources: Anthropic's "Building effective agents" blog post (which advocates for structured agent workflows over ad-hoc prompting), the ECC project's empirical evaluation data across dozens of agents, and community frameworks like gstack and GSD that independently converged on similar phase structures. The workflow is not prescribed by any single authority -- it emerged from practitioners observing what works at scale.
The Production Workflow
Every major Claude Code framework has converged on the same fundamental pattern, though they call it different things:
Research -> Plan -> Execute -> Review -> Ship
| | | | |
| | | | |
Explore Create Code with Verify Commit
the space the map the plan quality and deployThis is not a suggestion. Teams that skip phases -- especially Research and Review -- produce measurably worse output. The ECC project tracked this across multiple agents and found that agents following R-P-E-R-S had significantly higher pass@1 scores than those that jumped straight to implementation (the improvement was consistent across task types, though the exact magnitude varied by task complexity).
Phase Breakdown
| Phase | Tool | What Happens | Context Impact |
|---|---|---|---|
| Research | Explore subagent | Read files, understand architecture, map dependencies | Subagent isolates context -- main agent stays clean |
| Plan | Plan subagent or /plan | Produce a reviewable specification: files to change, approach, risks | Human reviews plan before execution starts |
| Execute | Main agent + Skills + Hooks | Implement the plan, run tests at each step | Skills automate repetitive operations |
| Review | Code reviewer subagent | Independent review: security, correctness, performance | Separate context prevents bias from implementation phase |
| Ship | /commit + CI/CD | Commit with structured message, create PR, run pipeline | Automated quality gates |
Framework Patterns
gstack's "Boil the Lake" Principle
The gstack framework (68k stars) operates on a completeness principle: do the job thoroughly, with full transparency about effort. Its key contribution is the effort table -- empirical data from teams using gstack in production on how much Claude Code accelerates different task types:
Effort Compression Ratios -- Your Mileage May Vary
The following compression ratios were reported by specific teams using gstack in production environments. Your actual results will depend on codebase complexity, language, test coverage, and how well your CLAUDE.md describes your project. Treat these as directional indicators, not guarantees.
| Task Type | Human Time | With Claude Code | Compression Ratio |
|---|---|---|---|
| Boilerplate (CRUD, forms, configs) | 4 hours | 2-3 minutes | ~100x |
| Tests (unit, integration, e2e) | 3 hours | 3-4 minutes | ~50x |
| Features (new functionality) | 8 hours | 15-20 minutes | ~30x |
| Bug Fixes (diagnose + fix) | 2 hours | 5-10 minutes | ~20x |
| Architecture (design + refactor) | 2 days | 2-3 hours | ~8x |
The compression ratio drops as tasks require more judgment and less typing.
gstack also introduces auto-decision logic -- six principles for when Claude should make decisions autonomously vs. ask the human:
- Reversible decisions: auto-decide (can be undone easily)
- Established patterns: auto-decide (matches existing codebase conventions)
- Performance trade-offs: ask (human needs to weigh constraints)
- Architecture changes: ask (high blast radius)
- External API contracts: ask (affects other teams)
- Security boundaries: ask (risk of credential exposure)
GSD's Wave Execution
The GSD framework (50k stars, adopted at Amazon, Google, Shopify) solves a specific problem: how do you execute a large task with many independent subtasks without blowing up the context window?
Wave execution groups tasks into waves of parallel work. Each wave runs in fresh contexts:
Wave 1 (parallel, fresh contexts):
├── Agent A: Create user model + migration
├── Agent B: Create order model + migration
└── Agent C: Create product model + migration
Wait for all to complete. Verify with tests.
Wave 2 (parallel, fresh contexts):
├── Agent A: User API endpoints + tests
├── Agent B: Order API endpoints + tests
└── Agent C: Product API endpoints + tests
Wait for all to complete. Integration test.
Wave 3 (sequential, builds on wave 2):
└── Agent A: Wire up cross-model relationships,
add auth middleware, run full test suiteContext Utilization "Golden Zone" -- Community-Reported Experience
The 40-60% context utilization range is widely cited across Claude Code community frameworks (GSD, HumanLayer, and others) as the effective operating range. Below 40%, Claude may lack sufficient context to make good decisions. Above 60%, quality tends to degrade as the context window fills with noise. This is based on community-reported experience across many projects, not a formal benchmark. Your optimal range may differ depending on task type and context composition.
HumanLayer's RPI Pattern
HumanLayer got Claude to handle a 300k LOC Rust codebase in about one hour using the Research-Plan-Implement pattern with frequent intentional compaction:
Research (in a subagent):
Read the codebase structure
Identify all files related to the change
Map the dependency graph
Output: structured summary (not the raw files)
--> Subagent exits, its context is freed
Plan (main agent, with research summary):
Design the approach based on the summary
List specific files to modify and how
Identify risks and edge cases
Output: step-by-step plan
--> Compact the context (remove research details, keep plan)
Implement (main agent, plan in context):
Execute the plan step by step
After each step, run tests
If context exceeds 40%, compact again
--> Repeat: compact after each major stepThe critical technique is intentional compaction -- do not wait for the automatic compaction at 83.5% context. Proactively compact at 40-50% to keep Claude's working memory clean and focused. In practice, this means:
# In a Claude Code session, compact proactively:
/compact Keep only: (1) the current plan, (2) files modified so far,
(3) test results. Drop all exploration context.BMAD-METHOD's Agent Personas
BMAD (44k stars) takes a different approach: instead of phases, it defines 12 agent personas that handle different aspects of software delivery:
| Persona | Role | Claude Code Equivalent |
|---|---|---|
| Analyst | Requirements gathering | Explore subagent with domain prompts |
| PM | Feature prioritization | Plan subagent with business context |
| Architect | System design | Plan subagent with architecture focus |
| Developer | Implementation | Main agent with Skills |
| QA | Testing | Code reviewer subagent |
| Scrum Master | Progress tracking | Scheduled Desktop App task |
BMAD's contribution is Scale-Domain-Adaptive complexity: the same framework adjusts its overhead based on project scope. A weekend project uses 3 personas. An enterprise migration uses all 12.
Demo 40: Full R-P-E-R-S Workflow -- Adding OAuth to an Express App
The Scenario
You have an Express.js API with JWT-based authentication. You need to add OAuth2 support (Google and GitHub providers) without breaking existing auth. This is a real feature that touches middleware, routes, database schema, and frontend callbacks.
Phase 1: Research
claude
> Use a subagent to research the current auth implementation. I need to know:
>
> 1. How is JWT auth currently implemented? (middleware, token generation, validation)
> 2. What is the user model schema? (database, ORM, fields)
> 3. How are routes protected? (middleware chain, role checks)
> 4. What session/cookie handling exists?
> 5. Are there any existing OAuth-related files or dependencies?
> 6. What test coverage exists for auth?
>
> Research only. Do not change anything.Claude spawns an Explore subagent that reads ~15 files and returns a structured summary:
Auth Architecture Summary:
- JWT middleware in src/middleware/auth.ts using jsonwebtoken
- User model in src/models/user.ts (Prisma ORM, PostgreSQL)
Fields: id, email, passwordHash, name, role, createdAt
- Route protection via requireAuth() middleware
- No session/cookie handling -- pure stateless JWT
- No existing OAuth code or dependencies
- 23 auth-related tests in tests/auth/
- Token refresh handled client-sideWhat Just Happened? (Research Phase)
Phase 2: Plan
> Based on the research, create a plan for adding Google and GitHub OAuth.
>
> The plan must cover:
> - Database schema changes (new fields or tables)
> - New dependencies
> - New routes and middleware
> - How OAuth users coexist with password users
> - How to handle account linking (same email, different auth methods)
> - Test plan
> - Rollback strategy
>
> Write the plan to OAUTH_PLAN.md so I can review it before you implement anything.Claude generates a detailed plan. You review it:
## OAuth Implementation Plan
### Database Changes
- Add to User model: googleId, githubId, authProvider (enum: local|google|github)
- Add OAuthToken model: userId, provider, accessToken, refreshToken, expiresAt
- Migration: ALTER TABLE users ADD COLUMN auth_provider DEFAULT 'local'
### New Dependencies
- passport, passport-google-oauth20, passport-github2
### New Routes
- GET /auth/google (redirect to Google consent screen)
- GET /auth/google/callback (handle Google callback)
- GET /auth/github (redirect to GitHub consent screen)
- GET /auth/github/callback (handle GitHub callback)
### Account Linking
- If OAuth email matches existing user: link accounts (add googleId/githubId)
- If no match: create new user with authProvider set
- Users can have both password + OAuth (linked account)
### Files to Modify
1. prisma/schema.prisma (User model + OAuthToken model)
2. src/middleware/passport.ts (NEW - passport strategies)
3. src/routes/auth.ts (add OAuth routes)
4. src/services/auth.service.ts (add OAuth user creation/linking)
5. src/config/oauth.ts (NEW - provider configs)
### Files to Create Tests For
1. tests/auth/oauth-google.test.ts
2. tests/auth/oauth-github.test.ts
3. tests/auth/account-linking.test.ts
### Rollback
- Migration is additive only (new columns, new table)
- Rollback = remove new routes, drop new columns
- Existing JWT auth is unchangedYou approve the plan: "Looks good. Implement it."
What Just Happened? (Plan Phase)
Phase 3: Execute
> Implement the OAuth plan from OAUTH_PLAN.md. Rules:
>
> 1. One file at a time, in dependency order
> 2. Run `npm test` after each file change
> 3. If tests fail, fix before moving on
> 4. Commit after each successful step
> 5. Do not modify existing auth logic -- add alongside itClaude works through the plan:
Step 1: Database migration
- Updated prisma/schema.prisma
- Ran npx prisma migrate dev --name add-oauth
- npm test: 23/23 pass (no regressions)
- Committed: "Add OAuth fields to User model and OAuthToken table"
Step 2: OAuth config
- Created src/config/oauth.ts
- npm test: 23/23 pass
- Committed: "Add OAuth provider configuration"
Step 3: Passport strategies
- Created src/middleware/passport.ts
- npm test: 23/23 pass
- Committed: "Add Google and GitHub passport strategies"
Step 4: Auth service updates
- Updated src/services/auth.service.ts (added findOrCreateOAuthUser, linkAccount)
- npm test: 23/23 pass
- Committed: "Add OAuth user creation and account linking service"
Step 5: OAuth routes
- Updated src/routes/auth.ts
- npm test: 23/23 pass
- Committed: "Add OAuth routes for Google and GitHub"
Step 6: Tests
- Created tests/auth/oauth-google.test.ts (8 tests)
- Created tests/auth/oauth-github.test.ts (8 tests)
- Created tests/auth/account-linking.test.ts (6 tests)
- npm test: 45/45 pass (22 new tests + 23 existing)
- Committed: "Add OAuth and account linking tests"What Just Happened? (Execute Phase)
Phase 4: Review
> Use a code reviewer subagent to review all changes made in this session.
> Focus on:
> 1. Security: Are OAuth tokens stored securely? Is CSRF protection in place?
> 2. Edge cases: What happens if Google returns no email? Duplicate accounts?
> 3. Test coverage: Are error paths tested?The reviewer subagent analyzes the diff and reports:
Review Findings:
SECURITY:
- ISSUE: OAuth callback routes missing CSRF state parameter validation
Fix: Add state parameter generation and verification
- OK: OAuth tokens stored in separate table with proper encryption
- OK: Callback URLs validated against whitelist
EDGE CASES:
- ISSUE: No handling for Google accounts without email (rare but possible)
Fix: Return 400 with clear message if no email in OAuth profile
- OK: Duplicate email detection and account linking works correctly
TEST COVERAGE:
- ISSUE: Missing test for expired OAuth token refresh
Fix: Add test case for token expiry handling
- OK: All success and error paths testedWhat Just Happened? (Review Phase)
Phase 5: Ship
> Fix the three issues from the review, then create a PR.Claude fixes the issues, runs tests one more time (48/48 pass), and creates the PR:
$ git add -A
$ git commit -m "Fix OAuth CSRF, missing email handling, and token refresh test"
$ gh pr create --title "Add Google and GitHub OAuth support" \
--body "## Summary
- Adds OAuth2 login via Google and GitHub
- Account linking for users with matching email
- 25 new tests, all passing
- No changes to existing JWT auth flow
## Security
- CSRF state parameter on all OAuth flows
- OAuth tokens stored in separate encrypted table
- Callback URL whitelist validation
## Test Plan
- [x] Google login flow (8 tests)
- [x] GitHub login flow (8 tests)
- [x] Account linking (6 tests)
- [x] Edge cases: no email, expired tokens, CSRF (3 tests)"
Creating pull request for feature/oauth into main...
https://github.com/myorg/myapp/pull/43What Just Happened? (Ship Phase)
Demo 41: Cost Comparison
Same task -- "add OAuth to an Express app" -- executed four different ways:
Single Session (Baseline)
Approach: One long session, everything in sequence
Tokens: ~180,000 input + ~45,000 output
Cost: ~$2.70
Time: ~25 minutes
Pass rate: ~75% (sometimes misses edge cases without review phase)Subagent Isolation
Approach: Explore subagent for research, main agent for implementation
Tokens: ~120,000 input + ~40,000 output
Cost: ~$1.90 (30% savings)
Time: ~28 minutes (slightly slower due to subagent overhead)
Pass rate: ~85% (research phase catches more issues upfront)Agent Teams (Parallel)
Approach: 3 agents: backend, tests, review (parallel where possible)
Tokens: ~150,000 input + ~50,000 output (more total, but parallel)
Cost: ~$2.40
Time: ~15 minutes (parallel execution)
Pass rate: ~90% (independent review agent catches issues)GSD Waves
Approach: Wave 1: schema + config, Wave 2: routes + service, Wave 3: tests + review
Tokens: ~100,000 input + ~35,000 output
Cost: ~$1.60 (lowest)
Time: ~20 minutes
Pass rate: ~92% (fresh contexts per wave, best context hygiene)The Takeaway
| Approach | Cost | Time | Quality |
|---|---|---|---|
| Single session | $2.70 | 25 min | 75% pass |
| Subagent isolation | $1.90 | 28 min | 85% pass |
| Agent teams | $2.40 | 15 min | 90% pass |
| GSD waves | $1.60 | 20 min | 92% pass |
There is no universally "best" approach. The right choice depends on what you are optimizing for:
- Minimize cost: GSD waves (fresh contexts = fewer redundant tokens)
- Minimize time: Agent teams (parallel execution)
- Maximize quality: GSD waves or agent teams (both include independent review)
- Minimize complexity: Single session (simplest setup)
What Just Happened?
When Things Go Wrong
Workflow Phase Failure
Symptom: The Execute phase fails partway through -- tests break after step 3 of 6, and Claude gets stuck in a fix loop.
Step 3: Passport strategies
- Created src/middleware/passport.ts
- npm test: 18/23 pass -- 5 FAILURES
- Attempting fix...
- npm test: 16/23 pass -- 7 FAILURES (made it worse!)
- Attempting fix...
- npm test: 12/23 pass -- REGRESSIONCause: The plan had a dependency error (e.g., passport strategies imported from a file that does not exist yet), or the plan assumed an API that the library version does not support.
Fix:
- Stop the current phase immediately. Do not let Claude keep trying to fix cascading failures:
> Stop. Do not make any more changes.
> Run git diff to show me what has changed since the last passing commit.
> Then run git stash to save the changes and go back to the last good state.- Return to the Plan phase. Update the plan with what you learned from the failure:
> The plan failed at step 3. The issue is that passport-google-oauth20 v3
> changed its callback signature. Update the plan to account for the v3 API
> and re-order the steps if needed.- Resume Execute from the last good checkpoint. This is why committing after each step matters -- you can always roll back to a known-good state.
Cost Overrun
Symptom: A task you expected to cost $2-3 is burning through $10+ in tokens with no end in sight.
# Check your token usage in the Claude Code session
> /cost
Session tokens: 485,000 input / 120,000 output
Estimated cost: $8.40 and countingCause: Claude is re-reading large files repeatedly because the context is being compacted and it loses track of what it already read. Or the task scope expanded beyond what was planned.
Fix:
- Compact immediately with specific instructions about what to keep:
> /compact Keep only: the current plan, what has been completed (steps 1-3),
> and the specific error from step 4. Drop all file contents -- I will re-read
> only what is needed for step 4.- Check for re-reading loops: If Claude keeps reading the same files, it means compaction is dropping information it needs. Write critical information to a file:
> Write a STATUS.md file with: (1) what has been done, (2) what remains,
> (3) the current error. Then compact. After compaction, read STATUS.md
> to resume.- Switch to Sonnet for mechanical steps: If the remaining work is straightforward (adding type annotations, writing tests from a template), drop to a cheaper model:
claude --model claude-sonnet-4-6
> Read STATUS.md and continue the migration from where it left off.Context Overflow in Production
Symptom: In a CI/CD pipeline using the Agent SDK, the agent starts producing lower-quality output or making errors it would not normally make, late in a long-running task.
# In your CI logs:
[Agent] Converting file 38/50...
[Agent] ERROR: Created duplicate function name in user_service.ts
[Agent] ERROR: Import path references old file structure
[Agent] WARNING: Agent attempted to read a file it already read 3 turns agoCause: The context window is full. Automatic compaction has kicked in and is dropping information the agent needs. This is especially common in migration tasks that touch many files.
Fix for SDK users: Break the task into chunks that fit comfortably in context:
# Instead of one query() call for 50 files:
files_to_migrate = get_file_list()
chunk_size = 10
for i in range(0, len(files_to_migrate), chunk_size):
chunk = files_to_migrate[i:i + chunk_size]
# Each chunk gets a fresh context
async for msg in query(
prompt=f"""Migrate these files to TypeScript: {chunk}
Read MIGRATION_STATUS.md for context on what has been done.
After completing this batch, update MIGRATION_STATUS.md.""",
options=ClaudeCodeOptions(
allowed_tools=["Read", "Write", "Edit", "Bash", "Glob", "Grep"],
),
):
if msg.type == "text":
print(msg.text, end="")Fix for Managed Agents: Use the progress file pattern described in Ch14 and break large tasks into multiple sessions with explicit handoff via files on disk.
Framework Selection Guide
Eight major frameworks exist in the Claude Code ecosystem. Here is when to use each:
| Framework | Stars | Best For | Key Idea |
|---|---|---|---|
| Everything Claude Code (ECC) | 148k | Complete reference, security, evals | 47 agents, 181 skills, enterprise governance |
| gstack | 68k | Thoroughness, effort transparency | "Boil the lake" + auto-decision logic |
| GSD | 50k | Parallel execution, context hygiene | Wave execution, 40-60% context zone |
| BMAD-METHOD | 44k | Structured team workflows | 12 agent personas, scale-adaptive |
| HumanLayer | - | Large codebases (100k+ LOC) | RPI pattern, intentional compaction at 40% |
| Codex (OpenAI) | - | Second opinion on architecture | gstack uses it for multi-AI decisions |
| Cursor | - | IDE-native AI coding | Tighter IDE integration, different tool |
| Claude Code (vanilla) | - | Getting started, simple projects | No framework overhead |
Decision Tree
Is your project under 10k LOC?
Yes -> Vanilla Claude Code or gstack (simple mode)
No -> Continue
Do you need parallel execution across many files?
Yes -> GSD waves
No -> Continue
Do you need enterprise governance (audit, compliance, roles)?
Yes -> ECC framework
No -> Continue
Do you need structured team workflows with defined roles?
Yes -> BMAD-METHOD
No -> Continue
Is the codebase over 100k LOC?
Yes -> HumanLayer RPI pattern
No -> gstack (covers most cases well)Cost Optimization Strategies
| Strategy | Token Savings | How |
|---|---|---|
| Subagent context isolation | ~40% | Read large files in subagents; only summaries reach the main agent |
| Precise CLAUDE.md | ~15% | Fewer irrelevant instructions = fewer tokens consumed per turn |
| Sonnet for routine tasks | ~60% cost | Use --model claude-sonnet-4-6 for tests, boilerplate, simple edits |
| Pipe mode | ~30% | git diff | claude -p "review" avoids loading the full project |
| Proactive compaction | ~20% | /compact at 50% beats the automatic compaction at 83.5% |
| GSD wave execution | ~35% | Fresh contexts per wave eliminate accumulated noise |
| Batches API (SDK) | 50% cost | For bulk processing, submit as a batch for half-price tokens |
Monthly Cost Estimates (Per Developer)
| Usage Level | Tasks/Day | Estimated Monthly Cost |
|---|---|---|
| Light (simple edits, reviews) | 5-10 | $50-100 |
| Medium (features, bug fixes) | 10-20 | $150-300 |
| Heavy (architecture, migrations) | 20+ | $400-800 |
| Enterprise (with Managed Agents) | Varies | $500-2,000 |
Team Configuration Template
Here is a complete .claude/ directory setup for a production team:
.claude/
├── settings.json # Permissions, hooks, MCP servers
├── CLAUDE.md # Project conventions (checked into git)
├── agents/
│ ├── explore.md # Research agent prompt
│ ├── reviewer.md # Code review agent prompt
│ └── test-writer.md # Test generation agent prompt
├── skills/
│ ├── commit.md # Commit with conventional format
│ ├── pr-review.md # PR review checklist
│ └── migrate.md # Database migration helper
└── hooks/
├── agent-shield.py # Security runtime filter
└── audit-logger.sh # Append-only audit trailExercise
Design the complete Claude Code deployment for your team:
- Write
.claude/settings.jsonwith role-based permissions (junior/senior/lead) - Write
CLAUDE.mdwith your project's coding conventions - Create three agent definitions in
.claude/agents/ - Create three skill definitions in
.claude/skills/ - Set up a GitHub Actions workflow for automated PR review
- Run the full R-P-E-R-S workflow on a real feature
- Track the token usage and calculate the cost
Chapter Summary
- R-P-E-R-S (Research, Plan, Execute, Review, Ship) is the universal workflow across all major frameworks -- teams that skip phases produce worse output
- gstack's effort table shows compression ratios ranging from ~100x for boilerplate to ~8x for architecture work (reported by specific teams, your results will vary)
- GSD wave execution maximizes quality and minimizes cost through parallel work in fresh contexts
- HumanLayer's RPI pattern handles 300k+ LOC codebases by compacting at 40% rather than waiting for automatic compaction
- Framework choice depends on project size, team structure, and optimization priority
- Cost optimization is about context hygiene: isolate reads in subagents, compact proactively, use the right model for the right task
Appendix links: A01 Prompt Engineering covers how to write effective prompts in a tool-calling context -- the foundation for every phase of R-P-E-R-S. A02 Context Engineering explains the 2025-2026 paradigm shift from prompt engineering to context engineering. A07 Token Economics provides the cost models and budget optimization strategies behind the cost comparison in Demo 41.
Research reference: The convergence of independent frameworks (ECC, gstack, GSD, HumanLayer, BMAD) on the same Research-Plan-Execute-Review-Ship pattern is notable. While each framework has different terminology and emphasis, the underlying workflow is remarkably consistent. Anthropic's "Building effective agents" blog post (2024) provides the theoretical foundation: agents work best when they explore before acting, plan before implementing, and verify before shipping. The R-P-E-R-S pattern is this principle operationalized for software engineering.
Knowledge Check
You have completed all 15 chapters. Head to the Advanced Capstone Project to build a production-grade AI code review pipeline that ties everything together.