阶段性大项目:AI 代码审查流水线
构建一个生产级代码审查系统,结合 Managed Agents、Agent SDK、hooks 和整个高级阶段的框架模式。
你将构建什么
一个完整的 AI 代码审查流水线(Pipeline):
- 在每个 PR 上触发 -- 通过 GitHub Actions
- 运行多视角审查 -- 使用三个专用 subagent(安全、正确性、架构)
- 强制质量门禁 -- 通过 hooks 在 critical 发现时阻止合并
- 使用基于角色的审查视角 -- 受 gstack 的 CEO/Design/Eng 审查模式启发
- 发布结构化审查评论 -- 包含严重性、文件引用和修复建议
- 追踪审查质量 -- 使用 pass@k 指标随时间监控
这不是玩具项目。完成后,你将拥有一个 10-50 人工程团队可以部署到真实仓库的审查流水线。
架构
GitHub PR Event
|
v
GitHub Actions Workflow
|
v
Review Orchestrator (Agent SDK)
|
├── Security Reviewer (subagent)
| Checks: injection, auth bypass, secrets, SSRF
|
├── Correctness Reviewer (subagent)
| Checks: logic errors, error handling, race conditions
|
├── Architecture Reviewer (subagent)
| Checks: coupling, naming, duplication, test coverage
|
v
Result Aggregator
|
├── Format review comment (Markdown table)
├── Calculate severity score
├── Post to PR as review comment
└── If critical findings: block merge (exit code 1)涉及的技能
| 技能 | 章节 | 如何使用 |
|---|---|---|
| 权限规则 | Ch10 | 审查员只有只读权限。编排器有发布评论的写权限 |
| IDE 工作流 | Ch11 | VS Code 用于本地开发和测试流水线 |
| GitHub Actions | Ch12 | 在 PR 事件上触发流水线 |
| Agent SDK | Ch13 | Python SDK 编排 subagent 并发布结果 |
| Managed Agents | Ch14 | 可选:将流水线作为 Managed Agent 运行用于长时间审查 |
| 生产模式 | Ch15 | R-P-E-R-S 工作流、成本优化、框架模式 |
实施计划
阶段 1:Research
在编写流水线之前研究目标仓库的惯例。
bash
claude
> I am building an AI code review pipeline for this repository. Before writing
> any code, research:
>
> 1. What languages and frameworks does this repo use?
> 2. What is the existing CI/CD setup? (.github/workflows/)
> 3. What test framework is used? What is the coverage?
> 4. Are there existing code review guidelines? (CONTRIBUTING.md, PR templates)
> 5. What are the most common PR review comments from the last 20 merged PRs?
> (Use `gh pr list --state merged --limit 20` and check comments)
>
> This research will inform what the review agents should focus on.阶段 2:Plan
基于研究发现设计流水线。
创建 REVIEW_PIPELINE_PLAN.md,包含:
- 包含哪些审查视角(安全、正确性、架构)
- 每个审查员检查什么(针对这个仓库的技术栈)
- 严重性级别以及什么阻止合并 vs 什么是建议性的
- 每次审查的成本预算(目标:平均 PR 低于 $0.20)
- 如何处理大型 PR(超过 50 个文件变更)
阶段 3:Execute
步骤 1:创建审查编排器
python
# review_pipeline/orchestrator.py
import asyncio
import json
import os
import subprocess
from claude_code_sdk import query, ClaudeCodeOptions
class ReviewOrchestrator:
"""
Orchestrates multi-perspective code review on a PR.
Spawns specialized subagents in parallel, aggregates findings,
and posts a structured review comment.
"""
def __init__(self, repo: str, pr_number: int):
self.repo = repo
self.pr_number = pr_number
self.diff = ""
self.pr_info = {}
self.findings = []
async def run(self) -> int:
"""
Run the full review pipeline.
Returns exit code: 0 = pass, 1 = critical findings (block merge).
"""
# Fetch PR data
self.diff = self._fetch_diff()
self.pr_info = self._fetch_pr_info()
if not self.diff:
print("No diff found. Skipping review.")
return 0
# Truncate large diffs to stay within budget
max_diff_chars = 60000 # ~15k tokens
if len(self.diff) > max_diff_chars:
self.diff = self.diff[:max_diff_chars]
self.diff += "\n\n[DIFF TRUNCATED -- only first 60k chars reviewed]"
# Run all reviewers in parallel
security, correctness, architecture = await asyncio.gather(
self._run_security_review(),
self._run_correctness_review(),
self._run_architecture_review(),
)
# Aggregate findings
all_findings = []
for category, findings_json in [
("Security", security),
("Correctness", correctness),
("Architecture", architecture),
]:
parsed = self._parse_findings(findings_json)
for f in parsed:
f["category"] = category
all_findings.extend(parsed)
# Sort by severity
severity_order = {"critical": 0, "high": 1, "medium": 2, "low": 3}
all_findings.sort(
key=lambda f: severity_order.get(f.get("severity", "low"), 4)
)
# Format and post the review
comment = self._format_review(all_findings)
self._post_comment(comment)
# Determine exit code
has_critical = any(
f.get("severity") == "critical" for f in all_findings
)
if has_critical:
print(f"BLOCKING: {sum(1 for f in all_findings if f.get('severity') == 'critical')} critical finding(s)")
return 1
print(f"Review complete: {len(all_findings)} finding(s), none critical")
return 0
def _fetch_diff(self) -> str:
result = subprocess.run(
["gh", "api", f"repos/{self.repo}/pulls/{self.pr_number}",
"--header", "Accept: application/vnd.github.v3.diff"],
capture_output=True, text=True,
)
return result.stdout
def _fetch_pr_info(self) -> dict:
result = subprocess.run(
["gh", "api", f"repos/{self.repo}/pulls/{self.pr_number}"],
capture_output=True, text=True,
)
try:
return json.loads(result.stdout)
except json.JSONDecodeError:
return {}
async def _run_security_review(self) -> str:
output = []
async for msg in query(
prompt=f"""You are a security reviewer. Analyze this PR diff.
PR: {self.pr_info.get('title', '')}
{self.diff}
Check for:
1. SQL injection (string concatenation in queries, unsanitized input)
2. XSS (unescaped user input in HTML/JSX rendering)
3. Authentication bypass (missing auth middleware, broken access control)
4. Path traversal (user input in file paths without sanitization)
5. Secrets in code (hardcoded API keys, passwords, connection strings)
6. SSRF (user-controlled URLs in server-side HTTP requests)
7. Command injection (user input in shell commands or exec calls)
8. Insecure deserialization (pickle, eval, yaml.load without SafeLoader)
Output a JSON array. Each finding:
{{"severity":"critical|high|medium","file":"path","line":N,"issue":"description","fix":"suggestion"}}
If nothing found, output: []
Only report real issues with specific file and line references.""",
options=ClaudeCodeOptions(
allowed_tools=[],
max_tokens=4096,
),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)
async def _run_correctness_review(self) -> str:
output = []
async for msg in query(
prompt=f"""You are a correctness reviewer. Analyze this PR diff.
PR: {self.pr_info.get('title', '')}
Description: {self.pr_info.get('body', '')[:2000]}
{self.diff}
Check for:
1. Logic errors (wrong conditions, off-by-one, missing edge cases)
2. Error handling gaps (unhandled promise rejections, missing try/catch)
3. Race conditions in async/concurrent code
4. Null/undefined access without guards
5. Incorrect API contract (wrong HTTP methods, missing required fields)
6. Resource leaks (unclosed connections, file handles, event listeners)
7. Incorrect error propagation (swallowed errors, wrong error types)
Output a JSON array. Each finding:
{{"severity":"critical|high|medium","file":"path","line":N,"issue":"description","suggestion":"what to do"}}
If nothing found, output: []""",
options=ClaudeCodeOptions(
allowed_tools=[],
max_tokens=4096,
),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)
async def _run_architecture_review(self) -> str:
output = []
async for msg in query(
prompt=f"""You are an architecture reviewer. Analyze this PR diff for
maintainability and design issues. Do NOT nitpick style -- focus on
structural problems that will cause pain in 6 months.
PR: {self.pr_info.get('title', '')}
{self.diff}
Check for:
1. Functions over 50 lines that should be decomposed
2. Duplicated logic (same pattern appearing in multiple places)
3. Missing abstractions (hardcoded values that should be config)
4. Tight coupling (modules directly depending on implementation details)
5. Missing tests for new functionality
6. API contract changes without migration path
7. Dead code or unreachable branches
Output a JSON array. Each finding:
{{"severity":"medium|low","file":"path","line":N,"issue":"description","suggestion":"what to do"}}
If nothing found, output: []""",
options=ClaudeCodeOptions(
allowed_tools=[],
max_tokens=4096,
),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)
def _parse_findings(self, raw: str) -> list:
try:
start = raw.find("[")
end = raw.rfind("]") + 1
if start >= 0 and end > start:
return json.loads(raw[start:end])
except json.JSONDecodeError:
pass
return []
def _format_review(self, findings: list) -> str:
title = self.pr_info.get("title", f"PR #{self.pr_number}")
parts = [f"## AI Code Review: {title}\n"]
criticals = [f for f in findings if f.get("severity") == "critical"]
highs = [f for f in findings if f.get("severity") == "high"]
mediums = [f for f in findings if f.get("severity") == "medium"]
lows = [f for f in findings if f.get("severity") == "low"]
if criticals:
parts.append(
f"**BLOCKING**: {len(criticals)} critical issue(s) must be "
f"resolved before merge.\n"
)
elif not findings:
parts.append("No significant issues found. Looks good.\n")
else:
parts.append(f"Found {len(findings)} issue(s) to review.\n")
# Summary table
if findings:
parts.append("| Severity | Category | File | Issue |")
parts.append("|----------|----------|------|-------|")
for f in findings:
sev = f.get("severity", "?")
cat = f.get("category", "?")
file_ref = f"`{f.get('file', '?')}:{f.get('line', '?')}`"
issue = f.get("issue", "")
parts.append(f"| {sev} | {cat} | {file_ref} | {issue} |")
# Detailed suggestions
actionable = [f for f in findings if f.get("fix") or f.get("suggestion")]
if actionable:
parts.append("\n### Suggested Fixes\n")
for f in actionable:
fix = f.get("fix") or f.get("suggestion", "")
parts.append(
f"**{f.get('file', '?')}:{f.get('line', '?')}** "
f"({f.get('category', '')})\n{fix}\n"
)
parts.append("\n---\n*Review generated by AI Code Review Pipeline*")
return "\n".join(parts)
def _post_comment(self, body: str):
subprocess.run(
["gh", "api", f"repos/{self.repo}/issues/{self.pr_number}/comments",
"--method", "POST",
"--field", f"body={body}"],
)
async def main():
repo = os.environ.get("GITHUB_REPOSITORY", "")
pr_number = int(os.environ.get("PR_NUMBER", "0"))
if not repo or not pr_number:
print("GITHUB_REPOSITORY and PR_NUMBER environment variables required")
raise SystemExit(1)
orchestrator = ReviewOrchestrator(repo, pr_number)
exit_code = await orchestrator.run()
raise SystemExit(exit_code)
if __name__ == "__main__":
asyncio.run(main())步骤 2:创建 GitHub Actions 工作流
yaml
# .github/workflows/ai-review-pipeline.yml
name: AI Code Review Pipeline
on:
pull_request:
types: [opened, synchronize, reopened]
permissions:
contents: read
pull-requests: write
issues: write
jobs:
ai-review:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install dependencies
run: |
pip install claude-code-sdk
npm install -g @anthropic-ai/claude-code
- name: Run AI Review Pipeline
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GITHUB_REPOSITORY: ${{ github.repository }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: python review_pipeline/orchestrator.py步骤 3:添加质量门禁 Hook
创建 PreToolUse hook 验证审查输出质量:
python
# review_pipeline/hooks/validate_review.py
"""
Validates that review findings are specific enough to be actionable.
Rejects vague findings like "code could be improved" without file/line refs.
"""
import sys
import json
def validate(tool_input: str):
try:
data = json.loads(tool_input)
findings = data if isinstance(data, list) else []
for finding in findings:
# Every finding must have a file reference
if not finding.get("file"):
print("REJECTED: Finding missing file reference", file=sys.stderr)
sys.exit(2)
# Every finding must have a concrete issue description
issue = finding.get("issue", "")
if len(issue) < 20:
print(f"REJECTED: Issue too vague: '{issue}'", file=sys.stderr)
sys.exit(2)
sys.exit(0)
except (json.JSONDecodeError, TypeError):
sys.exit(0) # Not a findings output, allow
if __name__ == "__main__":
validate(sys.argv[1] if len(sys.argv) > 1 else "")阶段 4:Review
先在你自己的 PR 上运行流水线:
bash
# Test locally
GITHUB_REPOSITORY=yourorg/yourrepo PR_NUMBER=42 python review_pipeline/orchestrator.py检查审查输出:
- 发现是否具体(文件 + 行号 + 描述)?
- 严重性评级是否适当(不是所有都标为 critical)?
- 建议是否可操作(能根据描述实现修复)?
- 成本是否在预算内(每个 PR 低于 $0.20)?
阶段 5:Ship
- 提交流水线代码
- 用流水线本身创建 PR -- 流水线会审查自己的 PR
- 在分支保护设置中将安全门禁设为必需检查
- 监控前 10 个 PR 的误报率并调整审查提示词
扩展:基于角色的审查(gstack 模式)
gstack 的审查模式对不同利益相关者使用不同的"镜头":
python
# Additional reviewer for product-impacting changes
async def _run_product_review(self) -> str:
"""Reviews from a product perspective -- UX impact, feature completeness."""
output = []
async for msg in query(
prompt=f"""You are a product reviewer. This is NOT a code review.
Analyze this PR from a product perspective:
1. Does this change affect the user experience? How?
2. Are there accessibility implications?
3. Is the feature complete, or are there missing states (loading, error, empty)?
4. Are there backward compatibility concerns for existing users?
5. Should this be behind a feature flag?
Only report product-level concerns, not code issues.
Output JSON array or [] if no concerns.""",
options=ClaudeCodeOptions(allowed_tools=[], max_tokens=2048),
):
if msg.type == "text":
output.append(msg.text)
return "".join(output)验收标准
- [ ] 流水线通过 GitHub Actions 在 PR 事件上触发
- [ ] 三个 subagent 并行运行(安全、正确性、架构)
- [ ] 结果聚合为单个结构化 PR 评论
- [ ] Critical 发现阻止合并(exit code 1)
- [ ] 平均审查成本低于 $0.20
- [ ] 流水线成功审查了自己的 PR
- [ ] 质量门禁 hook 拒绝模糊发现
- [ ] 完整的 R-P-E-R-S 工作流记录在提交历史中
你学到了什么
完成这个大项目后,你已经展示了:
- Ch10:权限规则将审查员 agent 限定为只读访问
- Ch11:使用 VS Code 本地开发和调试流水线
- Ch12:GitHub Actions 集成及正确的密钥管理
- Ch13:Agent SDK 编排并行 subagent
- Ch14:可选的 Managed Agents 用于大型 PR 的长时间审查
- Ch15:R-P-E-R-S 工作流、成本优化和框架模式
你现在拥有在生产规模部署 Claude Code 的技能。