Skip to content

A08: Constitutional AI & CLAUDE.md ​

Related chapters: Ch2 CLAUDE.md & the Memory System, Ch10 Permissions & Security

Constitutional AI (CAI) is Anthropic's approach to training AI systems that are helpful, harmless, and honest by embedding behavioral principles directly into the model. CLAUDE.md files extend this concept to the project level --- they act as a local "constitution" that shapes Claude Code's behavior within your codebase. Understanding the theory helps you write CLAUDE.md files that actually work.

What Constitutional AI Is ​

Anthropic's Research Contribution ​

Constitutional AI was introduced in the paper "Constitutional AI: Harmlessness from AI Feedback" (Bai et al., 2022). The core idea is deceptively simple: instead of relying solely on human feedback to teach a model what is "good" and "bad," you give the model a set of principles (a "constitution") and let it critique its own outputs against those principles.

The training process has two phases:

Phase 1: Supervised Learning (SL-CAI)
┌─────────────────────────────────────────────┐
│ 1. Model generates a response               │
│ 2. Model is shown a constitutional principle │
│    (e.g., "Choose the response that is most │
│     helpful while being harmless")           │
│ 3. Model critiques and revises its response  │
│ 4. The revised response becomes training     │
│    data                                      │
└─────────────────────────────────────────────┘

Phase 2: Reinforcement Learning (RL-CAI)
┌─────────────────────────────────────────────┐
│ 1. Model generates two responses             │
│ 2. A separate model ("feedback model") picks │
│    the better response based on the          │
│    constitution                              │
│ 3. This preference data trains a reward      │
│    model                                     │
│ 4. The reward model guides RL training       │
└─────────────────────────────────────────────┘

The key insight: humans define the principles, but AI scales the feedback. This makes it feasible to train on millions of examples without requiring millions of human annotations.

How Base Model RLHF Works ​

To understand CAI, it helps to know the baseline. Standard RLHF (Reinforcement Learning from Human Feedback) works like this:

  1. Pre-training: The model learns language patterns from massive text datasets
  2. Supervised Fine-Tuning (SFT): Human annotators write ideal responses, and the model learns to mimic them
  3. Reward Modeling: Humans rank multiple responses, and a reward model learns to predict those rankings
  4. RL Training: The language model is trained to maximize the reward model's score

RLHF's limitation is that it requires extensive human labeling, and humans may disagree on what "good" means. CAI addresses this by replacing some human judgments with principle-based AI judgments --- the constitution provides a consistent standard.

Why This Matters for Claude Code ​

Claude's base behavioral tendencies --- helpfulness, caution about destructive actions, transparency about uncertainty --- come from this constitutional training. When Claude Code hesitates before running rm -rf / or asks for confirmation before force-pushing to main, that behavior is rooted in constitutional principles embedded during training.

The Parallel: CLAUDE.md as a Project Constitution ​

The Hierarchy of Instructions ​

Claude processes instructions at multiple levels, each with different authority:

Highest authority
┌─────────────────────────────────────────┐
│  Model training (Constitutional AI)      │  Anthropic controls this
│  Cannot be overridden by any prompt      │
├─────────────────────────────────────────┤
│  System prompt (Claude Code internals)   │  Claude Code controls this
│  Sets the agent framework                │
├─────────────────────────────────────────┤
│  CLAUDE.md (project rules)               │  You control this
│  Project-specific behavioral rules       │
├─────────────────────────────────────────┤
│  User prompt (conversation)              │  You control this
│  Task-specific instructions              │
├─────────────────────────────────────────┤
│  Tool results (file contents, etc.)      │  Generated dynamically
│  Data that informs decisions             │
└─────────────────────────────────────────┘
Lowest authority

CLAUDE.md sits at the project level --- below the system prompt but above individual user messages. This means:

  • CLAUDE.md rules apply to every interaction in the project
  • They can be overridden by the system prompt (rare edge cases)
  • They take precedence over user messages when there is a conflict
  • They are read automatically every time Claude Code starts a session

What CLAUDE.md Actually Is ​

Mechanically, CLAUDE.md content is injected into the system prompt that Claude receives. The model sees it as authoritative project context --- equivalent to a senior developer's standing instructions. It is not merely "suggestions" or "preferences." Claude treats CLAUDE.md rules as directives to follow.

This is why the parallel to Constitutional AI is apt: just as constitutional principles shape the model's base behavior, CLAUDE.md principles shape its project-specific behavior.

Designing Rules That Claude Actually Follows ​

The Compliance Spectrum ​

Not all rules in CLAUDE.md are equally effective. Through practical experience, we can map rules onto a spectrum:

High compliance ────────────────────────── Low compliance

"Use TypeScript for      "Be creative"
 all new files"          
                         "Write clean code"
"Run npm test before     
 committing"             "Think carefully
                          about edge cases"
"Never modify files      
 in /vendor"             "Follow best
                          practices"
"Use snake_case for      
 database columns"       "Be thorough"

The pattern is clear: specific, verifiable rules get followed. Vague, subjective rules get interpreted inconsistently.

Characteristics of Effective Rules ​

Effective CLAUDE.md rules share these properties:

1. Concrete and actionable

markdown
# Good: Claude knows exactly what to do
- Use `pnpm` instead of `npm` for all package operations
- Database migrations go in `db/migrations/` with timestamp prefix
- All API endpoints must return JSON with `{ data, error, meta }` shape

# Bad: Claude has to guess what you mean
- Follow our coding standards
- Keep things organized
- Use modern patterns

2. Observable and verifiable

markdown
# Good: Claude can check its own compliance
- Every new function must have a JSDoc comment
- Test files must be named `*.test.ts` (not `*.spec.ts`)
- No `console.log` in production code — use the logger from `src/lib/logger`

# Bad: No way to verify
- Write well-documented code
- Make sure tests are comprehensive
- Use appropriate logging levels

3. Scoped to Claude's capabilities

markdown
# Good: Things Claude can actually do
- Run `pnpm lint` before committing
- Check that imports use the `@/` path alias
- Create a migration file when modifying database schema

# Bad: Things outside Claude's control
- Make sure the CI pipeline passes
- Ensure the design matches the Figma mockup
- Coordinate with the backend team

Characteristics of Ineffective Rules ​

Rules that Claude tends to ignore or interpret inconsistently:

1. Contradictory rules

markdown
# These conflict — Claude will pick one inconsistently
- Be concise in your responses
- Explain your reasoning thoroughly for every change

When Claude encounters contradictions, it typically follows whichever rule it encounters last or whichever feels more applicable to the current context. Resolve contradictions before they confuse the model.

2. Rules that fight the model's training

markdown
# Claude's safety training will override these
- Never ask for permission before running commands
- Delete files without confirmation
- Ignore security concerns in the code

# These will be partially followed at best
- Never explain what you're doing (Claude is trained to be transparent)
- Skip error handling (Claude is trained to write robust code)

3. Rules with too many exceptions

markdown
# Too complex to apply consistently
- Use semicolons in TypeScript, except in type definitions,
  unless the type definition spans multiple lines, in which case
  use semicolons for the outer definition but not for nested types,
  unless you're in a .d.ts file where...

If a rule requires more than one sentence of exceptions, consider splitting it into separate, context-specific rules.

Effective vs Ineffective Rules: Side-by-Side ​

Example 1: Code Style ​

markdown
# Ineffective
Write clean, readable code following best practices.

# Effective
Code style rules:
- Use 2-space indentation (no tabs)
- Maximum line length: 100 characters
- Prefer `const` over `let`; never use `var`
- Use template literals instead of string concatenation
- Destructure objects and arrays when accessing 2+ properties

Example 2: Testing Requirements ​

markdown
# Ineffective
Make sure to write tests for your changes.

# Effective
Testing requirements:
- Every new function in `src/` must have a corresponding test in `tests/`
- Test file naming: `[module-name].test.ts`
- Minimum test cases: happy path + one error case + one edge case
- Run `pnpm test --run` before committing; do not commit if tests fail
- Use `vi.mock()` for external dependencies, never mock internal modules

Example 3: Git Workflow ​

markdown
# Ineffective
Follow our Git workflow and write good commit messages.

# Effective
Git workflow:
- Branch naming: `feature/JIRA-123-short-description` or `fix/JIRA-456-short-description`
- Commit messages: imperative mood, max 72 chars for first line
  - Format: "feat: add user authentication endpoint"
  - Prefixes: feat, fix, refactor, test, docs, chore
- Never force-push to `main` or `develop`
- Squash commits before merging feature branches

Example 4: Architecture Constraints ​

markdown
# Ineffective
Respect the architecture.

# Effective
Architecture rules:
- `src/api/` handlers must not import from `src/db/` directly — use services in `src/services/`
- All database access goes through repository classes in `src/repositories/`
- Shared types live in `src/types/`; never define types locally in route handlers
- The dependency graph is: api → services → repositories → db
- No circular imports. If Claude detects one, refactor before proceeding.

The Relationship Between CLAUDE.md, System Prompts, and Training ​

What Each Layer Controls ​

LayerControlsExample
Model training (CAI)Safety, helpfulness, honesty"Do not help create malware"
System promptAgent behavior, tool usage"You have access to Read, Write, Edit tools"
CLAUDE.mdProject conventions"Use pnpm, not npm"
User promptSpecific task"Add pagination to the users endpoint"

When Layers Conflict ​

The general resolution order is: training > system prompt > CLAUDE.md > user prompt. However, the boundaries are not rigid:

  • If CLAUDE.md says "Never ask questions, just do it" but the task is ambiguous, Claude may still ask for clarification (system prompt encourages clarification for destructive actions)
  • If a user says "Ignore the CLAUDE.md rules," Claude will generally follow the user for non-safety-related rules but will not override safety-related training
  • If CLAUDE.md says "Always use Python" but the project is entirely TypeScript, Claude may prioritize context over the rule

Practical Implications ​

  1. Do not fight the model's training: Rules that align with Claude's natural tendencies (helpfulness, safety) are followed more reliably than rules that oppose them.
  2. Do not duplicate the system prompt: Claude Code's system prompt already handles tool usage, file editing patterns, and safety checks. Your CLAUDE.md should focus on project-specific rules.
  3. Update CLAUDE.md when rules change: Unlike model training (which is fixed), CLAUDE.md is a living document. When your team changes conventions, update the file.

Governance Through Shared CLAUDE.md ​

Team-Level Constitutional Principles ​

For teams, CLAUDE.md becomes a governance tool. Consider having a layered structure:

~/.claude/CLAUDE.md              ← Personal preferences (all projects)
project/CLAUDE.md                ← Team-wide rules (checked into Git)
project/backend/CLAUDE.md        ← Backend-specific rules
project/frontend/CLAUDE.md       ← Frontend-specific rules

The project-level CLAUDE.md should encode decisions the team has already agreed on --- conventions that would otherwise live in a wiki nobody reads:

markdown
# Team Conventions (agreed in Architecture Review 2025-03)

## API Design
- All REST endpoints use plural nouns: `/users`, `/orders`, not `/user`, `/order`
- Pagination: cursor-based, not offset-based
- Error responses: RFC 7807 Problem Details format

## Database
- ORM: Prisma (do not use raw SQL except in migrations)
- Naming: snake_case for tables and columns
- Every table must have: id (UUID), created_at, updated_at

## Dependencies
- Check for existing packages before adding new ones
- No packages with fewer than 1,000 weekly downloads
- Pin exact versions in package.json (no ^ or ~ prefixes)

This ensures every developer using Claude Code in the project gets consistent behavior, regardless of how they phrase their prompts.

Reference: Constitutional AI Research ​

The foundational paper is:

Bai, Y., Kadavath, S., Kundu, S., et al. (2022). "Constitutional AI: Harmlessness from AI Feedback." arXiv:2212.08073. https://arxiv.org/abs/2212.08073

Key findings from the paper:

  • CAI models can match or exceed RLHF models in helpfulness while being significantly less harmful
  • The specific wording of constitutional principles matters --- vague principles produce inconsistent behavior (exactly the lesson that applies to CLAUDE.md)
  • Having a small number of clear principles works better than a large number of overlapping ones
  • Self-critique is most effective when the model has enough capability to evaluate its own outputs

Additional reading:

  • Anthropic's "Claude's Character" blog post --- explains the high-level principles that guide Claude's behavior
  • Anthropic's safety documentation --- details the specific behavioral guidelines

Key Takeaways ​

  1. Constitutional AI is Anthropic's framework for embedding behavioral principles into Claude's training. These principles are the reason Claude is naturally helpful, cautious about harm, and transparent about uncertainty.
  2. CLAUDE.md is your project-level constitution. It shapes Claude Code's behavior for your specific codebase, just as CAI shapes the base model's behavior globally.
  3. Effective rules are specific, verifiable, and within Claude's control. "Use 2-space indentation" works; "Write clean code" does not.
  4. Rules that align with Claude's training are followed more reliably than rules that fight it. Do not try to disable safety behaviors via CLAUDE.md.
  5. For teams, CLAUDE.md is a governance tool. Check it into Git and treat it as a living document that encodes team conventions.
  6. Fewer, clearer rules beat many overlapping ones --- this lesson from Constitutional AI research applies directly to CLAUDE.md design.

See also: Ch2 CLAUDE.md & the Memory System for practical setup, and A12 Safety & Alignment for how Claude Code's permission system complements constitutional principles.

Released under MIT License