Capstone Project: AI Code Review Pipeline
Build a production-grade code review system that combines Managed Agents, the Agent SDK, hooks, and framework patterns from the entire Advanced stage.
What You Will Build
A complete AI code review pipeline that:
- Triggers on every PR via GitHub Actions
- Runs multi-perspective review using three specialized subagents (security, correctness, architecture)
- Enforces quality gates via hooks that block merges on critical findings
- Uses role-based review perspectives inspired by gstack's CEO/Design/Eng review pattern
- Posts structured review comments with severity, file references, and fix suggestions
- Tracks review quality with pass@k metrics over time
This is not a toy project. When you finish, you will have a review pipeline that a team of 10-50 engineers can deploy on a real repository.
Architecture
GitHub PR Event
|
v
GitHub Actions Workflow
|
v
Review Orchestrator (Agent SDK)
|
├── Security Reviewer (subagent)
| Checks: injection, auth bypass, secrets, SSRF
|
├── Correctness Reviewer (subagent)
| Checks: logic errors, error handling, race conditions
|
├── Architecture Reviewer (subagent)
| Checks: coupling, naming, duplication, test coverage
|
v
Result Aggregator
|
├── Format review comment (Markdown table)
├── Calculate severity score
├── Post to PR as review comment
└── If critical findings: block merge (exit code 1)Skills Exercised
| Skill | Chapter | How It Is Used |
|---|---|---|
| Permission rules | Ch10 | Reviewers get read-only access. Orchestrator has write access for posting comments |
| IDE workflow | Ch11 | VS Code for developing and testing the pipeline locally |
| GitHub Actions | Ch12 | Trigger the pipeline on PR events |
| Agent SDK | Ch13 | Python SDK for orchestrating subagents and posting results |
| Managed Agents | Ch14 | Optional: run the pipeline as a Managed Agent for long-running reviews |
| Production patterns | Ch15 | R-P-E-R-S workflow, cost optimization, framework patterns |
Implementation Plan
Phase 1: Research
Study the target repository's conventions before writing the pipeline.
claude
> I am building an AI code review pipeline for this repository. Before writing
> any code, research:
>
> 1. What languages and frameworks does this repo use?
> 2. What is the existing CI/CD setup? (.github/workflows/)
> 3. What test framework is used? What is the coverage?
> 4. Are there existing code review guidelines? (CONTRIBUTING.md, PR templates)
> 5. What are the most common PR review comments from the last 20 merged PRs?
> (Use `gh pr list --state merged --limit 20` and check comments)
>
> This research will inform what the review agents should focus on.Phase 2: Plan
Design the pipeline based on research findings.
Create a REVIEW_PIPELINE_PLAN.md with:
- Which review perspectives to include (security, correctness, architecture)
- What each reviewer checks (specific to this repo's tech stack)
- Severity levels and what blocks merge vs. what is advisory
- Cost budget per review (target: under $0.20 per average PR)
- How to handle large PRs (over 50 files changed)
Phase 3: Execute
Step 1: Create the Review Orchestrator
# review_pipeline/orchestrator.py
import asyncio
import json
import os
import subprocess
from claude_code_sdk import query, ClaudeCodeOptions
class ReviewOrchestrator:
"""
Orchestrates multi-perspective code review on a PR.
Spawns specialized subagents in parallel, aggregates findings,
and posts a structured review comment.
"""
def __init__(self, repo: str, pr_number: int):
self.repo = repo
self.pr_number = pr_number
self.diff = ""
self.pr_info = {}
self.findings = []
async def run(self) -> int:
"""
Run the full review pipeline.
Returns exit code: 0 = pass, 1 = critical findings (block merge).
"""
# Fetch PR data
self.diff = self._fetch_diff()
self.pr_info = self._fetch_pr_info()
if not self.diff:
print("No diff found. Skipping review.")
return 0
# Truncate large diffs to stay within budget
max_diff_chars = 60000 # ~15k tokens
if len(self.diff) > max_diff_chars:
self.diff = self.diff[:max_diff_chars]
self.diff += "\n\n[DIFF TRUNCATED -- only first 60k chars reviewed]"
# Run all reviewers in parallel
security, correctness, architecture = await asyncio.gather(
self._run_security_review(),
self._run_correctness_review(),
self._run_architecture_review(),
)
# Aggregate findings
all_findings = []
for category, findings_json in [
("Security", security),
("Correctness", correctness),
("Architecture", architecture),
]:
parsed = self._parse_findings(findings_json)
for f in parsed:
f["category"] = category
all_findings.extend(parsed)
# Sort by severity
severity_order = {"critical": 0, "high": 1, "medium": 2, "low": 3}
all_findings.sort(
key=lambda f: severity_order.get(f.get("severity", "low"), 4)
)
# Format and post the review
comment = self._format_review(all_findings)
self._post_comment(comment)
# Determine exit code
has_critical = any(
f.get("severity") == "critical" for f in all_findings
)
if has_critical:
print(f"BLOCKING: {sum(1 for f in all_findings if f.get('severity') == 'critical')} critical finding(s)")
return 1
print(f"Review complete: {len(all_findings)} finding(s), none critical")
return 0
def _fetch_diff(self) -> str:
result = subprocess.run(
["gh", "api", f"repos/{self.repo}/pulls/{self.pr_number}",
"--header", "Accept: application/vnd.github.v3.diff"],
capture_output=True, text=True,
)
return result.stdout
def _fetch_pr_info(self) -> dict:
result = subprocess.run(
["gh", "api", f"repos/{self.repo}/pulls/{self.pr_number}"],
capture_output=True, text=True,
)
try:
return json.loads(result.stdout)
except json.JSONDecodeError:
return {}
async def _run_security_review(self) -> str:
output = []
async for msg in query(
prompt=f"""You are a security reviewer. Analyze this PR diff.
PR: {self.pr_info.get('title', '')}
{self.diff}
Check for:
1. SQL injection (string concatenation in queries, unsanitized input)
2. XSS (unescaped user input in HTML/JSX rendering)
3. Authentication bypass (missing auth middleware, broken access control)
4. Path traversal (user input in file paths without sanitization)
5. Secrets in code (hardcoded API keys, passwords, connection strings)
6. SSRF (user-controlled URLs in server-side HTTP requests)
7. Command injection (user input in shell commands or exec calls)
8. Insecure deserialization (pickle, eval, yaml.load without SafeLoader)
Output a JSON array. Each finding:
{{"severity":"critical|high|medium","file":"path","line":N,"issue":"description","fix":"suggestion"}}
If nothing found, output: []
Only report real issues with specific file and line references.""",
options=ClaudeCodeOptions(
allowed_tools=[],
max_tokens=4096,
),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)
async def _run_correctness_review(self) -> str:
output = []
async for msg in query(
prompt=f"""You are a correctness reviewer. Analyze this PR diff.
PR: {self.pr_info.get('title', '')}
Description: {self.pr_info.get('body', '')[:2000]}
{self.diff}
Check for:
1. Logic errors (wrong conditions, off-by-one, missing edge cases)
2. Error handling gaps (unhandled promise rejections, missing try/catch)
3. Race conditions in async/concurrent code
4. Null/undefined access without guards
5. Incorrect API contract (wrong HTTP methods, missing required fields)
6. Resource leaks (unclosed connections, file handles, event listeners)
7. Incorrect error propagation (swallowed errors, wrong error types)
Output a JSON array. Each finding:
{{"severity":"critical|high|medium","file":"path","line":N,"issue":"description","suggestion":"what to do"}}
If nothing found, output: []""",
options=ClaudeCodeOptions(
allowed_tools=[],
max_tokens=4096,
),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)
async def _run_architecture_review(self) -> str:
output = []
async for msg in query(
prompt=f"""You are an architecture reviewer. Analyze this PR diff for
maintainability and design issues. Do NOT nitpick style -- focus on
structural problems that will cause pain in 6 months.
PR: {self.pr_info.get('title', '')}
{self.diff}
Check for:
1. Functions over 50 lines that should be decomposed
2. Duplicated logic (same pattern appearing in multiple places)
3. Missing abstractions (hardcoded values that should be config)
4. Tight coupling (modules directly depending on implementation details)
5. Missing tests for new functionality
6. API contract changes without migration path
7. Dead code or unreachable branches
Output a JSON array. Each finding:
{{"severity":"medium|low","file":"path","line":N,"issue":"description","suggestion":"what to do"}}
If nothing found, output: []""",
options=ClaudeCodeOptions(
allowed_tools=[],
max_tokens=4096,
),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)
def _parse_findings(self, raw: str) -> list:
try:
start = raw.find("[")
end = raw.rfind("]") + 1
if start >= 0 and end > start:
return json.loads(raw[start:end])
except json.JSONDecodeError:
pass
return []
def _format_review(self, findings: list) -> str:
title = self.pr_info.get("title", f"PR #{self.pr_number}")
parts = [f"## AI Code Review: {title}\n"]
criticals = [f for f in findings if f.get("severity") == "critical"]
highs = [f for f in findings if f.get("severity") == "high"]
mediums = [f for f in findings if f.get("severity") == "medium"]
lows = [f for f in findings if f.get("severity") == "low"]
if criticals:
parts.append(
f"**BLOCKING**: {len(criticals)} critical issue(s) must be "
f"resolved before merge.\n"
)
elif not findings:
parts.append("No significant issues found. Looks good.\n")
else:
parts.append(f"Found {len(findings)} issue(s) to review.\n")
# Summary table
if findings:
parts.append("| Severity | Category | File | Issue |")
parts.append("|----------|----------|------|-------|")
for f in findings:
sev = f.get("severity", "?")
cat = f.get("category", "?")
file_ref = f"`{f.get('file', '?')}:{f.get('line', '?')}`"
issue = f.get("issue", "")
parts.append(f"| {sev} | {cat} | {file_ref} | {issue} |")
# Detailed suggestions
actionable = [f for f in findings if f.get("fix") or f.get("suggestion")]
if actionable:
parts.append("\n### Suggested Fixes\n")
for f in actionable:
fix = f.get("fix") or f.get("suggestion", "")
parts.append(
f"**{f.get('file', '?')}:{f.get('line', '?')}** "
f"({f.get('category', '')})\n{fix}\n"
)
parts.append("\n---\n*Review generated by AI Code Review Pipeline*")
return "\n".join(parts)
def _post_comment(self, body: str):
subprocess.run(
["gh", "api", f"repos/{self.repo}/issues/{self.pr_number}/comments",
"--method", "POST",
"--field", f"body={body}"],
)
async def main():
repo = os.environ.get("GITHUB_REPOSITORY", "")
pr_number = int(os.environ.get("PR_NUMBER", "0"))
if not repo or not pr_number:
print("GITHUB_REPOSITORY and PR_NUMBER environment variables required")
raise SystemExit(1)
orchestrator = ReviewOrchestrator(repo, pr_number)
exit_code = await orchestrator.run()
raise SystemExit(exit_code)
if __name__ == "__main__":
asyncio.run(main())Step 2: Create the GitHub Actions Workflow
# .github/workflows/ai-review-pipeline.yml
name: AI Code Review Pipeline
on:
pull_request:
types: [opened, synchronize, reopened]
permissions:
contents: read
pull-requests: write
issues: write
jobs:
ai-review:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install dependencies
run: |
pip install claude-code-sdk
npm install -g @anthropic-ai/claude-code
- name: Run AI Review Pipeline
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GITHUB_REPOSITORY: ${{ github.repository }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: python review_pipeline/orchestrator.pyStep 3: Add Quality Gate Hooks
Create a PreToolUse hook that validates review output quality:
# review_pipeline/hooks/validate_review.py
"""
Validates that review findings are specific enough to be actionable.
Rejects vague findings like "code could be improved" without file/line refs.
"""
import sys
import json
def validate(tool_input: str):
try:
data = json.loads(tool_input)
findings = data if isinstance(data, list) else []
for finding in findings:
# Every finding must have a file reference
if not finding.get("file"):
print("REJECTED: Finding missing file reference", file=sys.stderr)
sys.exit(2)
# Every finding must have a concrete issue description
issue = finding.get("issue", "")
if len(issue) < 20:
print(f"REJECTED: Issue too vague: '{issue}'", file=sys.stderr)
sys.exit(2)
sys.exit(0)
except (json.JSONDecodeError, TypeError):
sys.exit(0) # Not a findings output, allow
if __name__ == "__main__":
validate(sys.argv[1] if len(sys.argv) > 1 else "")Phase 4: Review
Run the pipeline on your own PRs first:
# Test locally
GITHUB_REPOSITORY=yourorg/yourrepo PR_NUMBER=42 python review_pipeline/orchestrator.pyCheck the review output for:
- Are findings specific (file + line + description)?
- Are severity ratings appropriate (not everything marked critical)?
- Are suggestions actionable (can you implement the fix from the description)?
- Is the cost within budget (under $0.20 per PR)?
Phase 5: Ship
- Commit the pipeline code
- Create a PR with the pipeline itself -- the pipeline will review its own PR
- Make the security gate a required check in branch protection settings
- Monitor the first 10 PRs for false positives and tune the review prompts
Extension: Role-Based Reviews (gstack Pattern)
gstack's review pattern uses different "lenses" for different stakeholders:
# Additional reviewer for product-impacting changes
async def _run_product_review(self) -> str:
"""Reviews from a product perspective -- UX impact, feature completeness."""
output = []
async for msg in query(
prompt=f"""You are a product reviewer. This is NOT a code review.
Analyze this PR from a product perspective:
1. Does this change affect the user experience? How?
2. Are there accessibility implications?
3. Is the feature complete, or are there missing states (loading, error, empty)?
4. Are there backward compatibility concerns for existing users?
5. Should this be behind a feature flag?
Only report product-level concerns, not code issues.
Output JSON array or [] if no concerns.""",
options=ClaudeCodeOptions(allowed_tools=[], max_tokens=2048),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)Acceptance Criteria
- [ ] Pipeline triggers on PR events via GitHub Actions
- [ ] Three subagents run in parallel (security, correctness, architecture)
- [ ] Results aggregated into a single structured PR comment
- [ ] Critical findings block merge (exit code 1)
- [ ] Cost per average review is under $0.20
- [ ] Pipeline successfully reviews its own PR
- [ ] Quality gate hook rejects vague findings
- [ ] Full R-P-E-R-S workflow documented in commit history
What You Have Learned
By completing this capstone, you have demonstrated:
- Ch10: Permission rules scoping reviewer agents to read-only access
- Ch11: Using VS Code to develop and debug the pipeline locally
- Ch12: GitHub Actions integration with proper secrets management
- Ch13: Agent SDK for orchestrating parallel subagents
- Ch14: Optional Managed Agents for long-running reviews on large PRs
- Ch15: R-P-E-R-S workflow, cost optimization, and framework patterns
You now have the skills to deploy Claude Code at production scale.