Skip to content

A08: Constitutional AI 与 CLAUDE.md ​

相关章节: 第 2 章 CLAUDE.md 与记忆系统、第 10 章 权限与安全

Constitutional AI(CAI)是 Anthropic 训练有帮助、无害且诚实的 AI 系统的方法,通过将行为准则直接嵌入模型来实现。CLAUDE.md 文件将这一概念扩展到项目层面 --- 它们作为一部本地"宪法",在你的代码库中塑造 Claude Code 的行为。理解理论有助于你写出真正有效的 CLAUDE.md 文件。

什么是 Constitutional AI ​

Anthropic 的研究贡献 ​

Constitutional AI 在论文"Constitutional AI: Harmlessness from AI Feedback"(Bai 等人,2022)中被提出。核心思想看似简单:不仅依赖人类反馈来教模型什么是"好"和"坏",而是给模型一套准则(一部"宪法"),让它根据这些准则自我批评输出。

训练过程分为两个阶段:

Phase 1: Supervised Learning (SL-CAI)
┌─────────────────────────────────────────────┐
│ 1. Model generates a response               │
│ 2. Model is shown a constitutional principle │
│    (e.g., "Choose the response that is most │
│     helpful while being harmless")           │
│ 3. Model critiques and revises its response  │
│ 4. The revised response becomes training     │
│    data                                      │
└─────────────────────────────────────────────┘

Phase 2: Reinforcement Learning (RL-CAI)
┌─────────────────────────────────────────────┐
│ 1. Model generates two responses             │
│ 2. A separate model ("feedback model") picks │
│    the better response based on the          │
│    constitution                              │
│ 3. This preference data trains a reward      │
│    model                                     │
│ 4. The reward model guides RL training       │
└─────────────────────────────────────────────┘

关键洞察:人类定义准则,但 AI 扩展反馈。这使得在不需要数百万人工标注的情况下,在数百万示例上进行训练成为可能。

基础 RLHF 的工作原理 ​

要理解 CAI,了解基线方法会有帮助。标准的 RLHF(基于人类反馈的强化学习)工作流程如下:

  1. 预训练:模型从大规模文本数据集中学习语言模式
  2. 监督微调(SFT):人类标注者编写理想回复,模型学习模仿
  3. 奖励建模:人类对多个回复进行排序,奖励模型学习预测这些排序
  4. RL 训练:语言模型被训练以最大化奖励模型的评分

RLHF 的局限在于它需要大量的人工标注,而人类可能对什么是"好"意见不一。CAI 通过用基于准则的 AI 判断替代部分人类判断来解决这个问题 --- 宪法提供了一致的标准。

为什么这对 Claude Code 很重要 ​

Claude 的基本行为倾向 --- 帮助性、对破坏性操作的谨慎、对不确定性的透明 --- 来自这种宪法训练。当 Claude Code 在运行 rm -rf / 前犹豫或在强制推送到 main 前请求确认时,这些行为根植于训练期间嵌入的宪法准则。

类比:CLAUDE.md 作为项目宪法 ​

指令层级 ​

Claude 在多个层级处理指令,每个层级有不同的权威:

Highest authority
┌─────────────────────────────────────────┐
│  Model training (Constitutional AI)      │  Anthropic controls this
│  Cannot be overridden by any prompt      │
├─────────────────────────────────────────┤
│  System prompt (Claude Code internals)   │  Claude Code controls this
│  Sets the agent framework                │
├─────────────────────────────────────────┤
│  CLAUDE.md (project rules)               │  You control this
│  Project-specific behavioral rules       │
├─────────────────────────────────────────┤
│  User prompt (conversation)              │  You control this
│  Task-specific instructions              │
├─────────────────────────────────────────┤
│  Tool results (file contents, etc.)      │  Generated dynamically
│  Data that informs decisions             │
└─────────────────────────────────────────┘
Lowest authority

CLAUDE.md 位于项目级别 --- 低于系统提示但高于单独的用户消息。这意味着:

  • CLAUDE.md 规则适用于项目中的每次交互
  • 它们可以被系统提示覆盖(极少见的边缘情况)
  • 当与用户消息冲突时,它们优先
  • Claude Code 每次启动会话时都会自动读取它们

CLAUDE.md 的本质 ​

从机制上说,CLAUDE.md 内容被注入到 Claude 接收的系统提示中。模型将其视为权威的项目上下文 --- 等同于高级开发者的常设指令。它不仅仅是"建议"或"偏好"。Claude 将 CLAUDE.md 规则视为必须遵循的指令。

这就是为什么与 Constitutional AI 的类比是恰当的:就像宪法准则塑造模型的基本行为一样,CLAUDE.md 准则塑造其特定项目的行为。

设计 Claude 真正遵循的规则 ​

遵从度谱系 ​

CLAUDE.md 中并非所有规则都同样有效。通过实际经验,我们可以将规则映射到一个谱系上:

High compliance ────────────────────────── Low compliance

"Use TypeScript for      "Be creative"
 all new files"          
                         "Write clean code"
"Run npm test before     
 committing"             "Think carefully
                          about edge cases"
"Never modify files      
 in /vendor"             "Follow best
                          practices"
"Use snake_case for      
 database columns"       "Be thorough"

模式很明确:具体的、可验证的规则会被遵循。模糊的、主观的规则会被不一致地解读。

有效规则的特征 ​

有效的 CLAUDE.md 规则具有以下共同属性:

1. 具体且可操作

markdown
# Good: Claude knows exactly what to do
- Use `pnpm` instead of `npm` for all package operations
- Database migrations go in `db/migrations/` with timestamp prefix
- All API endpoints must return JSON with `{ data, error, meta }` shape

# Bad: Claude has to guess what you mean
- Follow our coding standards
- Keep things organized
- Use modern patterns

2. 可观察且可验证

markdown
# Good: Claude can check its own compliance
- Every new function must have a JSDoc comment
- Test files must be named `*.test.ts` (not `*.spec.ts`)
- No `console.log` in production code — use the logger from `src/lib/logger`

# Bad: No way to verify
- Write well-documented code
- Make sure tests are comprehensive
- Use appropriate logging levels

3. 在 Claude 能力范围内

markdown
# Good: Things Claude can actually do
- Run `pnpm lint` before committing
- Check that imports use the `@/` path alias
- Create a migration file when modifying database schema

# Bad: Things outside Claude's control
- Make sure the CI pipeline passes
- Ensure the design matches the Figma mockup
- Coordinate with the backend team

无效规则的特征 ​

Claude 倾向于忽略或不一致解读的规则:

1. 矛盾的规则

markdown
# These conflict — Claude will pick one inconsistently
- Be concise in your responses
- Explain your reasoning thoroughly for every change

当 Claude 遇到矛盾时,它通常会遵循最后遇到的或当前上下文中更适用的规则。在规则混淆模型之前解决矛盾。

2. 与模型训练对抗的规则

markdown
# Claude's safety training will override these
- Never ask for permission before running commands
- Delete files without confirmation
- Ignore security concerns in the code

# These will be partially followed at best
- Never explain what you're doing (Claude is trained to be transparent)
- Skip error handling (Claude is trained to write robust code)

3. 异常过多的规则

markdown
# Too complex to apply consistently
- Use semicolons in TypeScript, except in type definitions,
  unless the type definition spans multiple lines, in which case
  use semicolons for the outer definition but not for nested types,
  unless you're in a .d.ts file where...

如果一条规则需要超过一句话的例外说明,考虑将其拆分为独立的、上下文特定的规则。

有效 vs 无效规则:对比 ​

示例 1:代码风格 ​

markdown
# Ineffective
Write clean, readable code following best practices.

# Effective
Code style rules:
- Use 2-space indentation (no tabs)
- Maximum line length: 100 characters
- Prefer `const` over `let`; never use `var`
- Use template literals instead of string concatenation
- Destructure objects and arrays when accessing 2+ properties

示例 2:测试要求 ​

markdown
# Ineffective
Make sure to write tests for your changes.

# Effective
Testing requirements:
- Every new function in `src/` must have a corresponding test in `tests/`
- Test file naming: `[module-name].test.ts`
- Minimum test cases: happy path + one error case + one edge case
- Run `pnpm test --run` before committing; do not commit if tests fail
- Use `vi.mock()` for external dependencies, never mock internal modules

示例 3:Git 工作流 ​

markdown
# Ineffective
Follow our Git workflow and write good commit messages.

# Effective
Git workflow:
- Branch naming: `feature/JIRA-123-short-description` or `fix/JIRA-456-short-description`
- Commit messages: imperative mood, max 72 chars for first line
  - Format: "feat: add user authentication endpoint"
  - Prefixes: feat, fix, refactor, test, docs, chore
- Never force-push to `main` or `develop`
- Squash commits before merging feature branches

示例 4:架构约束 ​

markdown
# Ineffective
Respect the architecture.

# Effective
Architecture rules:
- `src/api/` handlers must not import from `src/db/` directly — use services in `src/services/`
- All database access goes through repository classes in `src/repositories/`
- Shared types live in `src/types/`; never define types locally in route handlers
- The dependency graph is: api → services → repositories → db
- No circular imports. If Claude detects one, refactor before proceeding.

CLAUDE.md、系统提示和训练之间的关系 ​

每个层级控制什么 ​

层级控制内容示例
模型训练(CAI)安全性、帮助性、诚实性"不要帮助创建恶意软件"
系统提示Agent 行为、工具使用"你可以使用 Read、Write、Edit 工具"
CLAUDE.md项目规范"使用 pnpm,而非 npm"
用户提示具体任务"为用户端点添加分页"

层级冲突时 ​

一般的解决顺序是:训练 > 系统提示 > CLAUDE.md > 用户提示。然而,边界并不严格:

  • 如果 CLAUDE.md 说"永远不要提问,直接做"但任务模糊不清,Claude 可能仍会请求澄清(系统提示鼓励对破坏性操作进行确认)
  • 如果用户说"忽略 CLAUDE.md 规则",Claude 通常会在非安全相关规则上听从用户,但不会覆盖安全相关的训练
  • 如果 CLAUDE.md 说"始终使用 Python"但项目完全是 TypeScript,Claude 可能会优先考虑上下文而非规则

实际影响 ​

  1. 不要与模型训练对抗:与 Claude 自然倾向(帮助性、安全性)一致的规则更可靠地被遵循,而与之对抗的规则则不然。
  2. 不要重复系统提示:Claude Code 的系统提示已经处理了工具使用、文件编辑模式和安全检查。你的 CLAUDE.md 应该专注于项目特定的规则。
  3. 规则变化时更新 CLAUDE.md:与模型训练(固定的)不同,CLAUDE.md 是活文档。当团队变更规范时,更新文件。

通过共享 CLAUDE.md 进行治理 ​

团队级别的宪法准则 ​

对于团队,CLAUDE.md 成为一种治理工具。考虑使用分层结构:

~/.claude/CLAUDE.md              ← Personal preferences (all projects)
project/CLAUDE.md                ← Team-wide rules (checked into Git)
project/backend/CLAUDE.md        ← Backend-specific rules
project/frontend/CLAUDE.md       ← Frontend-specific rules

项目级别的 CLAUDE.md 应该编码团队已经达成一致的决策 --- 那些否则会存在于没人阅读的 Wiki 中的规范:

markdown
# Team Conventions (agreed in Architecture Review 2025-03)

## API Design
- All REST endpoints use plural nouns: `/users`, `/orders`, not `/user`, `/order`
- Pagination: cursor-based, not offset-based
- Error responses: RFC 7807 Problem Details format

## Database
- ORM: Prisma (do not use raw SQL except in migrations)
- Naming: snake_case for tables and columns
- Every table must have: id (UUID), created_at, updated_at

## Dependencies
- Check for existing packages before adding new ones
- No packages with fewer than 1,000 weekly downloads
- Pin exact versions in package.json (no ^ or ~ prefixes)

这确保了项目中使用 Claude Code 的每个开发者都获得一致的行为,无论他们如何措辞提示。

参考:Constitutional AI 研究 ​

基础论文是:

Bai, Y., Kadavath, S., Kundu, S., et al. (2022). "Constitutional AI: Harmlessness from AI Feedback." arXiv:2212.08073. https://arxiv.org/abs/2212.08073

论文的关键发现:

  • CAI 模型在帮助性方面可以匹敌或超过 RLHF 模型,同时显著降低有害性
  • 宪法准则的具体措辞很重要 --- 模糊的准则产生不一致的行为(这与 CLAUDE.md 的教训完全一致)
  • 少量清晰的准则比大量重叠的准则效果更好
  • 当模型有足够能力评估自身输出时,自我批评最为有效

扩展阅读:

  • Anthropic 的"Claude's Character"博客文章 --- 解释了指导 Claude 行为的高层准则
  • Anthropic 的安全文档 --- 详述了具体的行为指南

核心要点 ​

  1. Constitutional AI 是 Anthropic 将行为准则嵌入 Claude 训练的框架。这些准则是 Claude 天然具有帮助性、对危害保持谨慎、对不确定性保持透明的原因。
  2. CLAUDE.md 是你的项目级宪法。它塑造 Claude Code 在你特定代码库中的行为,就像 CAI 在全局层面塑造基础模型的行为。
  3. 有效的规则是具体的、可验证的、在 Claude 能力范围内的。"使用 2 空格缩进"有效;"写干净的代码"无效。
  4. 与 Claude 训练一致的规则更可靠地被遵循,而与之对抗的规则则不然。不要尝试通过 CLAUDE.md 禁用安全行为。
  5. 对于团队,CLAUDE.md 是治理工具。将其签入 Git 并视为编码团队规范的活文档。
  6. 更少、更清晰的规则胜过大量重叠的规则 --- 这来自 Constitutional AI 研究的教训,直接适用于 CLAUDE.md 设计。

另见:第 2 章 CLAUDE.md 与记忆系统 了解实际设置,以及 A12 安全与对齐 了解 Claude Code 的权限系统如何补充宪法准则。

基于 MIT 许可发布