Skip to content

Capstone Project: AI Code Review Pipeline ​

Build a production-grade code review system that combines Managed Agents, the Agent SDK, hooks, and framework patterns from the entire Advanced stage.

What You Will Build ​

A complete AI code review pipeline that:

  1. Triggers on every PR via GitHub Actions
  2. Runs multi-perspective review using three specialized subagents (security, correctness, architecture)
  3. Enforces quality gates via hooks that block merges on critical findings
  4. Uses role-based review perspectives inspired by gstack's CEO/Design/Eng review pattern
  5. Posts structured review comments with severity, file references, and fix suggestions
  6. Tracks review quality with pass@k metrics over time

This is not a toy project. When you finish, you will have a review pipeline that a team of 10-50 engineers can deploy on a real repository.


Architecture ​

GitHub PR Event
     |
     v
GitHub Actions Workflow
     |
     v
Review Orchestrator (Agent SDK)
     |
     ├── Security Reviewer (subagent)
     |     Checks: injection, auth bypass, secrets, SSRF
     |
     ├── Correctness Reviewer (subagent)
     |     Checks: logic errors, error handling, race conditions
     |
     ├── Architecture Reviewer (subagent)
     |     Checks: coupling, naming, duplication, test coverage
     |
     v
Result Aggregator
     |
     ├── Format review comment (Markdown table)
     ├── Calculate severity score
     ├── Post to PR as review comment
     └── If critical findings: block merge (exit code 1)

Skills Exercised ​

SkillChapterHow It Is Used
Permission rulesCh10Reviewers get read-only access. Orchestrator has write access for posting comments
IDE workflowCh11VS Code for developing and testing the pipeline locally
GitHub ActionsCh12Trigger the pipeline on PR events
Agent SDKCh13Python SDK for orchestrating subagents and posting results
Managed AgentsCh14Optional: run the pipeline as a Managed Agent for long-running reviews
Production patternsCh15R-P-E-R-S workflow, cost optimization, framework patterns

Implementation Plan ​

Phase 1: Research ​

Study the target repository's conventions before writing the pipeline.

bash
claude

> I am building an AI code review pipeline for this repository. Before writing
> any code, research:
>
> 1. What languages and frameworks does this repo use?
> 2. What is the existing CI/CD setup? (.github/workflows/)
> 3. What test framework is used? What is the coverage?
> 4. Are there existing code review guidelines? (CONTRIBUTING.md, PR templates)
> 5. What are the most common PR review comments from the last 20 merged PRs?
>    (Use `gh pr list --state merged --limit 20` and check comments)
>
> This research will inform what the review agents should focus on.

Phase 2: Plan ​

Design the pipeline based on research findings.

Create a REVIEW_PIPELINE_PLAN.md with:

  • Which review perspectives to include (security, correctness, architecture)
  • What each reviewer checks (specific to this repo's tech stack)
  • Severity levels and what blocks merge vs. what is advisory
  • Cost budget per review (target: under $0.20 per average PR)
  • How to handle large PRs (over 50 files changed)

Phase 3: Execute ​

Step 1: Create the Review Orchestrator ​

python
# review_pipeline/orchestrator.py

import asyncio
import json
import os
import subprocess
from claude_code_sdk import query, ClaudeCodeOptions

class ReviewOrchestrator:
    """
    Orchestrates multi-perspective code review on a PR.
    
    Spawns specialized subagents in parallel, aggregates findings,
    and posts a structured review comment.
    """
    
    def __init__(self, repo: str, pr_number: int):
        self.repo = repo
        self.pr_number = pr_number
        self.diff = ""
        self.pr_info = {}
        self.findings = []
    
    async def run(self) -> int:
        """
        Run the full review pipeline.
        Returns exit code: 0 = pass, 1 = critical findings (block merge).
        """
        # Fetch PR data
        self.diff = self._fetch_diff()
        self.pr_info = self._fetch_pr_info()
        
        if not self.diff:
            print("No diff found. Skipping review.")
            return 0
        
        # Truncate large diffs to stay within budget
        max_diff_chars = 60000  # ~15k tokens
        if len(self.diff) > max_diff_chars:
            self.diff = self.diff[:max_diff_chars]
            self.diff += "\n\n[DIFF TRUNCATED -- only first 60k chars reviewed]"
        
        # Run all reviewers in parallel
        security, correctness, architecture = await asyncio.gather(
            self._run_security_review(),
            self._run_correctness_review(),
            self._run_architecture_review(),
        )
        
        # Aggregate findings
        all_findings = []
        for category, findings_json in [
            ("Security", security),
            ("Correctness", correctness),
            ("Architecture", architecture),
        ]:
            parsed = self._parse_findings(findings_json)
            for f in parsed:
                f["category"] = category
            all_findings.extend(parsed)
        
        # Sort by severity
        severity_order = {"critical": 0, "high": 1, "medium": 2, "low": 3}
        all_findings.sort(
            key=lambda f: severity_order.get(f.get("severity", "low"), 4)
        )
        
        # Format and post the review
        comment = self._format_review(all_findings)
        self._post_comment(comment)
        
        # Determine exit code
        has_critical = any(
            f.get("severity") == "critical" for f in all_findings
        )
        
        if has_critical:
            print(f"BLOCKING: {sum(1 for f in all_findings if f.get('severity') == 'critical')} critical finding(s)")
            return 1
        
        print(f"Review complete: {len(all_findings)} finding(s), none critical")
        return 0
    
    def _fetch_diff(self) -> str:
        result = subprocess.run(
            ["gh", "api", f"repos/{self.repo}/pulls/{self.pr_number}",
             "--header", "Accept: application/vnd.github.v3.diff"],
            capture_output=True, text=True,
        )
        return result.stdout
    
    def _fetch_pr_info(self) -> dict:
        result = subprocess.run(
            ["gh", "api", f"repos/{self.repo}/pulls/{self.pr_number}"],
            capture_output=True, text=True,
        )
        try:
            return json.loads(result.stdout)
        except json.JSONDecodeError:
            return {}
    
    async def _run_security_review(self) -> str:
        output = []
        async for msg in query(
            prompt=f"""You are a security reviewer. Analyze this PR diff.

PR: {self.pr_info.get('title', '')}

{self.diff}

Check for:
1. SQL injection (string concatenation in queries, unsanitized input)
2. XSS (unescaped user input in HTML/JSX rendering)
3. Authentication bypass (missing auth middleware, broken access control)
4. Path traversal (user input in file paths without sanitization)
5. Secrets in code (hardcoded API keys, passwords, connection strings)
6. SSRF (user-controlled URLs in server-side HTTP requests)
7. Command injection (user input in shell commands or exec calls)
8. Insecure deserialization (pickle, eval, yaml.load without SafeLoader)

Output a JSON array. Each finding:
{{"severity":"critical|high|medium","file":"path","line":N,"issue":"description","fix":"suggestion"}}

If nothing found, output: []
Only report real issues with specific file and line references.""",
            options=ClaudeCodeOptions(
                allowed_tools=[],
                max_tokens=4096,
            ),
        ):
            if msg.type == "text":
                output.append(msg.text)
        return "".join(output)
    
    async def _run_correctness_review(self) -> str:
        output = []
        async for msg in query(
            prompt=f"""You are a correctness reviewer. Analyze this PR diff.

PR: {self.pr_info.get('title', '')}
Description: {self.pr_info.get('body', '')[:2000]}

{self.diff}

Check for:
1. Logic errors (wrong conditions, off-by-one, missing edge cases)
2. Error handling gaps (unhandled promise rejections, missing try/catch)
3. Race conditions in async/concurrent code
4. Null/undefined access without guards
5. Incorrect API contract (wrong HTTP methods, missing required fields)
6. Resource leaks (unclosed connections, file handles, event listeners)
7. Incorrect error propagation (swallowed errors, wrong error types)

Output a JSON array. Each finding:
{{"severity":"critical|high|medium","file":"path","line":N,"issue":"description","suggestion":"what to do"}}

If nothing found, output: []""",
            options=ClaudeCodeOptions(
                allowed_tools=[],
                max_tokens=4096,
            ),
        ):
            if msg.type == "text":
                output.append(msg.text)
        return "".join(output)
    
    async def _run_architecture_review(self) -> str:
        output = []
        async for msg in query(
            prompt=f"""You are an architecture reviewer. Analyze this PR diff for
maintainability and design issues. Do NOT nitpick style -- focus on
structural problems that will cause pain in 6 months.

PR: {self.pr_info.get('title', '')}

{self.diff}

Check for:
1. Functions over 50 lines that should be decomposed
2. Duplicated logic (same pattern appearing in multiple places)
3. Missing abstractions (hardcoded values that should be config)
4. Tight coupling (modules directly depending on implementation details)
5. Missing tests for new functionality
6. API contract changes without migration path
7. Dead code or unreachable branches

Output a JSON array. Each finding:
{{"severity":"medium|low","file":"path","line":N,"issue":"description","suggestion":"what to do"}}

If nothing found, output: []""",
            options=ClaudeCodeOptions(
                allowed_tools=[],
                max_tokens=4096,
            ),
        ):
            if msg.type == "text":
                output.append(msg.text)
        return "".join(output)
    
    def _parse_findings(self, raw: str) -> list:
        try:
            start = raw.find("[")
            end = raw.rfind("]") + 1
            if start >= 0 and end > start:
                return json.loads(raw[start:end])
        except json.JSONDecodeError:
            pass
        return []
    
    def _format_review(self, findings: list) -> str:
        title = self.pr_info.get("title", f"PR #{self.pr_number}")
        
        parts = [f"## AI Code Review: {title}\n"]
        
        criticals = [f for f in findings if f.get("severity") == "critical"]
        highs = [f for f in findings if f.get("severity") == "high"]
        mediums = [f for f in findings if f.get("severity") == "medium"]
        lows = [f for f in findings if f.get("severity") == "low"]
        
        if criticals:
            parts.append(
                f"**BLOCKING**: {len(criticals)} critical issue(s) must be "
                f"resolved before merge.\n"
            )
        elif not findings:
            parts.append("No significant issues found. Looks good.\n")
        else:
            parts.append(f"Found {len(findings)} issue(s) to review.\n")
        
        # Summary table
        if findings:
            parts.append("| Severity | Category | File | Issue |")
            parts.append("|----------|----------|------|-------|")
            for f in findings:
                sev = f.get("severity", "?")
                cat = f.get("category", "?")
                file_ref = f"`{f.get('file', '?')}:{f.get('line', '?')}`"
                issue = f.get("issue", "")
                parts.append(f"| {sev} | {cat} | {file_ref} | {issue} |")
        
        # Detailed suggestions
        actionable = [f for f in findings if f.get("fix") or f.get("suggestion")]
        if actionable:
            parts.append("\n### Suggested Fixes\n")
            for f in actionable:
                fix = f.get("fix") or f.get("suggestion", "")
                parts.append(
                    f"**{f.get('file', '?')}:{f.get('line', '?')}** "
                    f"({f.get('category', '')})\n{fix}\n"
                )
        
        parts.append("\n---\n*Review generated by AI Code Review Pipeline*")
        
        return "\n".join(parts)
    
    def _post_comment(self, body: str):
        subprocess.run(
            ["gh", "api", f"repos/{self.repo}/issues/{self.pr_number}/comments",
             "--method", "POST",
             "--field", f"body={body}"],
        )


async def main():
    repo = os.environ.get("GITHUB_REPOSITORY", "")
    pr_number = int(os.environ.get("PR_NUMBER", "0"))
    
    if not repo or not pr_number:
        print("GITHUB_REPOSITORY and PR_NUMBER environment variables required")
        raise SystemExit(1)
    
    orchestrator = ReviewOrchestrator(repo, pr_number)
    exit_code = await orchestrator.run()
    raise SystemExit(exit_code)


if __name__ == "__main__":
    asyncio.run(main())

Step 2: Create the GitHub Actions Workflow ​

yaml
# .github/workflows/ai-review-pipeline.yml
name: AI Code Review Pipeline
on:
  pull_request:
    types: [opened, synchronize, reopened]

permissions:
  contents: read
  pull-requests: write
  issues: write

jobs:
  ai-review:
    runs-on: ubuntu-latest
    steps:
      - name: Checkout
        uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"

      - name: Install dependencies
        run: |
          pip install claude-code-sdk
          npm install -g @anthropic-ai/claude-code

      - name: Run AI Review Pipeline
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
          GITHUB_REPOSITORY: ${{ github.repository }}
          PR_NUMBER: ${{ github.event.pull_request.number }}
        run: python review_pipeline/orchestrator.py

Step 3: Add Quality Gate Hooks ​

Create a PreToolUse hook that validates review output quality:

python
# review_pipeline/hooks/validate_review.py
"""
Validates that review findings are specific enough to be actionable.
Rejects vague findings like "code could be improved" without file/line refs.
"""

import sys
import json

def validate(tool_input: str):
    try:
        data = json.loads(tool_input)
        findings = data if isinstance(data, list) else []
        
        for finding in findings:
            # Every finding must have a file reference
            if not finding.get("file"):
                print("REJECTED: Finding missing file reference", file=sys.stderr)
                sys.exit(2)
            
            # Every finding must have a concrete issue description
            issue = finding.get("issue", "")
            if len(issue) < 20:
                print(f"REJECTED: Issue too vague: '{issue}'", file=sys.stderr)
                sys.exit(2)
        
        sys.exit(0)
    except (json.JSONDecodeError, TypeError):
        sys.exit(0)  # Not a findings output, allow

if __name__ == "__main__":
    validate(sys.argv[1] if len(sys.argv) > 1 else "")

Phase 4: Review ​

Run the pipeline on your own PRs first:

bash
# Test locally
GITHUB_REPOSITORY=yourorg/yourrepo PR_NUMBER=42 python review_pipeline/orchestrator.py

Check the review output for:

  • Are findings specific (file + line + description)?
  • Are severity ratings appropriate (not everything marked critical)?
  • Are suggestions actionable (can you implement the fix from the description)?
  • Is the cost within budget (under $0.20 per PR)?

Phase 5: Ship ​

  1. Commit the pipeline code
  2. Create a PR with the pipeline itself -- the pipeline will review its own PR
  3. Make the security gate a required check in branch protection settings
  4. Monitor the first 10 PRs for false positives and tune the review prompts

Extension: Role-Based Reviews (gstack Pattern) ​

gstack's review pattern uses different "lenses" for different stakeholders:

python
# Additional reviewer for product-impacting changes
async def _run_product_review(self) -> str:
    """Reviews from a product perspective -- UX impact, feature completeness."""
    output = []
    async for msg in query(
        prompt=f"""You are a product reviewer. This is NOT a code review.
Analyze this PR from a product perspective:

1. Does this change affect the user experience? How?
2. Are there accessibility implications?
3. Is the feature complete, or are there missing states (loading, error, empty)?
4. Are there backward compatibility concerns for existing users?
5. Should this be behind a feature flag?

Only report product-level concerns, not code issues.
Output JSON array or [] if no concerns.""",
        options=ClaudeCodeOptions(allowed_tools=[], max_tokens=2048),
    ):
        if msg.type == "text":
            output.append(msg.text)
    return "".join(output)

Acceptance Criteria ​

  • [ ] Pipeline triggers on PR events via GitHub Actions
  • [ ] Three subagents run in parallel (security, correctness, architecture)
  • [ ] Results aggregated into a single structured PR comment
  • [ ] Critical findings block merge (exit code 1)
  • [ ] Cost per average review is under $0.20
  • [ ] Pipeline successfully reviews its own PR
  • [ ] Quality gate hook rejects vague findings
  • [ ] Full R-P-E-R-S workflow documented in commit history

What You Have Learned ​

By completing this capstone, you have demonstrated:

  • Ch10: Permission rules scoping reviewer agents to read-only access
  • Ch11: Using VS Code to develop and debug the pipeline locally
  • Ch12: GitHub Actions integration with proper secrets management
  • Ch13: Agent SDK for orchestrating parallel subagents
  • Ch14: Optional Managed Agents for long-running reviews on large PRs
  • Ch15: R-P-E-R-S workflow, cost optimization, and framework patterns

You now have the skills to deploy Claude Code at production scale.

Released under MIT License