Skip to content

A12: Safety & Alignment ​

Related chapters: Ch10 Permission Management & Security

Claude Code is an autonomous agent that can read your files, write code, execute shell commands, and interact with external services. Making that kind of power safe is not optional -- it is the core engineering challenge. This appendix explains the multi-layered safety architecture that makes Claude Code trustworthy, why Claude sometimes refuses requests, and how to design agent workflows that are safe by default.

Why Claude Refuses Certain Requests ​

If you have used Claude Code for any length of time, you have encountered a refusal -- a moment where Claude declines to do what you asked. This is not a bug. It is one of the most important features of the system.

The Alignment Training Pipeline ​

Claude's safety behavior comes from a multi-stage training process:

Pre-training (raw capability)
       │
       ▼
RLHF (Reinforcement Learning from Human Feedback)
       │  Humans rate outputs for helpfulness AND harmlessness
       ▼
Constitutional AI (self-critique)
       │  Claude evaluates its own outputs against principles
       ▼
Tool-use specific training
       │  Additional training for safe tool calling behavior
       ▼
Deployed Claude (in Claude Code)
       │
       ▼
Runtime safety layers (permissions, sandboxing)

Each layer adds a different kind of safety:

  • RLHF teaches Claude general boundaries: do not help with malware, do not produce harmful content, do not deceive users
  • Constitutional AI gives Claude internalized principles it applies through self-critique (see A08 Constitutional AI & CLAUDE.md)
  • Tool-use training teaches Claude specific caution around destructive operations: rm -rf, git push --force, modifying system files
  • Runtime layers (permissions, sandboxing) provide defense in depth even if the model's judgment fails

Common Refusal Categories in Claude Code ​

SituationWhy Claude RefusesWhat to Do Instead
Writing malware or exploit codeSafety training prevents harmful code generationDescribe the defensive need: "Write a test that verifies our input sanitization blocks SQL injection"
Deleting system filesTool-use training flags destructive system operationsBe specific about what you want deleted and why
Running commands with sudoElevated privilege operations require extra cautionGrant specific permissions in settings or confirm interactively
Accessing credentials in plaintextSafety training around secrets handlingUse environment variables, vault references, or .env files with proper .gitignore
Generating deceptive contentConstitutional AI principles against deceptionReframe the request honestly

Refusals Are Calibrated, Not Binary ​

Claude does not have a simple blocklist. Its refusals are contextual:

  • Writing a function called encrypt_payload is fine in a security library
  • Writing a function called encrypt_payload that exfiltrates data to an external server will be refused
  • Explaining how SQL injection works is fine in an educational context
  • Generating a SQL injection payload targeting a specific production database will be refused

If Claude refuses something you believe is legitimate, provide more context about why you need it. The refusal often comes from ambiguity about intent.

The Permission System as a Safety Layer ​

Claude Code implements a three-tier permission system that acts as runtime defense in depth:

┌─────────────────────────────────────────────────┐
│              Permission Resolution               │
│                                                 │
│  1. Check DENY rules  ──► Blocked? → Refuse     │
│         │                                       │
│         ▼                                       │
│  2. Check ALLOW rules ──► Allowed? → Execute    │
│         │                                       │
│         ▼                                       │
│  3. Default: ASK USER ──► Prompt for approval   │
│                                                 │
└─────────────────────────────────────────────────┘

Defense in Depth ​

The permission system provides safety in addition to model alignment, not instead of it. This is the principle of defense in depth:

Layer 1: Model alignment (Claude's training)
  │  Claude's judgment about what is safe
  ▼
Layer 2: Permission rules (.claude/settings.json)
  │  Project-level rules about what is allowed
  ▼
Layer 3: Interactive approval (ask-user prompts)
  │  Human in the loop for unrecognized operations
  ▼
Layer 4: OS-level sandboxing
  │  Process isolation, filesystem restrictions
  ▼
Layer 5: Git safety net
     Checkpoints, undo capability, version control

If any single layer fails, the others still provide protection. A model alignment failure (Claude misjudges a command as safe) is caught by the permission system. A permission misconfiguration (too-broad ALLOW rule) is mitigated by the model's own judgment. This redundancy is intentional.

Configuring Permissions Responsibly ​

json
// .claude/settings.json -- recommended starting point
{
  "permissions": {
    "allow": [
      "Read",
      "Glob",
      "Grep",
      "Bash(npm test*)",
      "Bash(npm run lint*)",
      "Bash(git status)",
      "Bash(git diff*)",
      "Bash(git log*)"
    ],
    "deny": [
      "Bash(rm -rf /)*",
      "Bash(sudo *)",
      "Bash(curl * | bash)",
      "Bash(git push --force*)",
      "Bash(chmod 777*)"
    ]
  }
}

Principle: Allow read broadly, allow write narrowly, deny destructive explicitly.

Read operations (Read, Glob, Grep) are safe to allow globally -- they cannot modify your system. Write operations and shell commands should be allowed only for specific, known-safe patterns. Destructive operations should be explicitly denied.

The Danger of Over-Permissioning ​

It is tempting to add "Bash(*)" to the allow list to eliminate all permission prompts. This is dangerous:

json
// DO NOT DO THIS in production workflows
{
  "permissions": {
    "allow": ["Bash(*)"]  // Allows ANY shell command without asking
  }
}

With this configuration, if Claude misinterprets a task or hallucinates a command, there is no human checkpoint to catch it. The model's alignment is your only safety layer, and alignment is probabilistic, not absolute.

A safer approach for reducing prompts:

json
{
  "permissions": {
    "allow": [
      "Bash(npm *)",
      "Bash(node *)",
      "Bash(git add *)",
      "Bash(git commit *)",
      "Bash(npx jest*)",
      "Bash(npx tsc*)",
      "Write"
    ]
  }
}

This allows the common development commands while still requiring approval for anything unexpected.

The Brain-vs-Hands Principle ​

Chapter 14 introduces the brain-vs-hands principle for agent systems, which is central to safety:

┌──────────────────────────────────┐
│          BRAIN (Claude)          │
│                                  │
│  Decides WHAT to do              │
│  Reasons about approach          │
│  Plans multi-step actions        │
│  Evaluates results               │
│                                  │
│  ───────── boundary ──────────   │
│                                  │
│          HANDS (Tools)           │
│                                  │
│  Executes specific actions       │
│  Reads/writes files              │
│  Runs commands                   │
│  Reports results                 │
│                                  │
└──────────────────────────────────┘

Safety implication: The brain (model) should never directly execute actions. It should always go through the hands (tools), and those tools have safety constraints:

  • Read: No side effects, always safe
  • Write/Edit: Modifiable by git, reversible via checkpoints
  • Bash: Sandboxed, permission-gated, timeout-limited
  • WebFetch: Network access is restricted and auditable

When you design agent workflows, maintain this separation. The model reasons; the tools act. The tools are constrained; the model is not. This asymmetry is the foundation of safe agent design.

Designing Agent Systems That Are Safe by Default ​

When building automation with Claude Code (CI/CD pipelines, scheduled agents, batch processing), safety becomes even more critical because there is no human watching in real time.

Principle 1: Least Privilege ​

Give the agent only the permissions it needs for its specific task:

bash
# CI/CD code review agent -- read-only, no write access needed
claude --model claude-sonnet-4-6-20250514 \
  --permission-mode deny-all \
  --allowedTools "Read,Glob,Grep,Bash(git diff*),Bash(git log*)" \
  -p "Review the diff on this branch for security issues."

Principle 2: Immutable Inputs ​

Agent systems should not modify their own configuration:

json
// The agent should NOT be able to edit these files
// Add them to deny rules
{
  "permissions": {
    "deny": [
      "Edit(.claude/*)",
      "Write(.claude/*)",
      "Edit(CLAUDE.md)",
      "Write(CLAUDE.md)"
    ]
  }
}

If an agent can modify its own CLAUDE.md or settings, it can effectively reprogram itself -- removing safety constraints that were intentionally placed.

Principle 3: Bounded Execution ​

Set time and cost limits on autonomous agents:

bash
# Set a maximum session cost
export CLAUDE_MAX_COST=5.00  # Stop after $5 of API usage

# Set a maximum number of tool calls
export CLAUDE_MAX_TURNS=50   # Stop after 50 turns

An agent without bounds can enter an infinite loop, consuming tokens indefinitely. Bounded execution is a safety net for both cost and behavior.

Principle 4: Audit Trail ​

Every action Claude Code takes is logged. In team environments, ensure these logs are preserved:

bash
# Enable verbose logging for CI/CD agents
claude --verbose --output-format json \
  -p "Run the deployment checklist" \
  2>&1 | tee /var/log/claude-agent/$(date +%Y%m%d-%H%M%S).json

The audit trail serves two purposes: debugging (what went wrong?) and accountability (who authorized this action?).

Responsible Agent Deployment Checklist ​

Before deploying any Claude Code agent in production or CI/CD, verify:

Permissions ​

  • [ ] Read operations are explicitly allowed (avoid over-prompting)
  • [ ] Write operations are scoped to specific directories
  • [ ] Shell commands are restricted to known-safe patterns
  • [ ] Destructive operations (rm -rf, git push --force) are explicitly denied
  • [ ] Network access is limited to necessary endpoints

Boundaries ​

  • [ ] Maximum cost per session is configured
  • [ ] Maximum turns per session is configured
  • [ ] Timeout is set for individual tool calls
  • [ ] Session timeout is set for the overall run

Secrets ​

  • [ ] API keys are in environment variables, not in files
  • [ ] .env files are in .gitignore
  • [ ] The agent cannot read credential stores or keychains
  • [ ] Output logs are scrubbed of sensitive data

Recovery ​

  • [ ] Git checkpoints are enabled (can undo agent changes)
  • [ ] The agent operates on a branch, not on main
  • [ ] There is a human review step before merging agent output
  • [ ] Rollback procedure is documented and tested

Monitoring ​

  • [ ] Agent runs are logged with timestamps and cost
  • [ ] Anomalous behavior triggers alerts (unexpected file modifications, high token usage)
  • [ ] Regular audits of permission configurations

Common Safety Pitfalls ​

Pitfall 1: Trusting Agent Output Without Review ​

bash
# DANGEROUS: Agent generates and deploys without human review
claude -p "Fix the production bug and deploy" --allow-all

Even a correct fix might have unintended side effects. Always include a human review step between "fix" and "deploy."

Pitfall 2: Allowing Agents to Install Packages ​

bash
# An agent that can run `npm install` can introduce supply-chain vulnerabilities
claude -p "Add the library we need for PDF generation"

If the agent installs a malicious or compromised package, it runs in your environment with your permissions. Lock down package installation to human-approved dependencies, or use a lockfile verification step.

Pitfall 3: Sharing Credentials via CLAUDE.md ​

markdown
<!-- DO NOT DO THIS in CLAUDE.md -->
## Database Access
Use the following connection string:
postgres://admin:s3cret_passw0rd@prod-db.example.com:5432/myapp

CLAUDE.md is checked into version control. Never put credentials in it. Use environment variable references instead:

markdown
## Database Access
The database connection string is in the $DATABASE_URL environment variable.

Pitfall 4: Ignoring the Sandbox in CI/CD ​

In CI/CD mode, Claude Code runs with reduced interactive safety (no human to approve prompts). If you skip the sandbox configuration, every shell command runs with full CI runner permissions:

yaml
# GitHub Actions -- configure sandboxing explicitly
- name: Run Claude Code Review
  run: |
    claude --permission-mode deny-all \
      --allowedTools "Read,Glob,Grep" \
      -p "Review the changes in this PR"

Pitfall 5: Recursive Self-Improvement ​

An agent that can modify its own prompts, instructions, or tool definitions can escape its safety constraints:

Agent reads CLAUDE.md → modifies CLAUDE.md to remove restrictions → 
reads new CLAUDE.md → operates without restrictions

Prevent this by denying write access to configuration files (see Principle 2 above).

Anthropic's Responsible Scaling Policy ​

Anthropic publishes a Responsible Scaling Policy that governs how Claude models are developed and deployed. Key points relevant to Claude Code users:

  1. ASL (AI Safety Level) framework: Models are evaluated for dangerous capabilities before deployment. Claude Code only ships with models that pass Anthropic's safety evaluations.

  2. Red teaming: Before each model release, dedicated teams attempt to elicit harmful behavior. Claude's refusals are informed by these exercises.

  3. Deployment safeguards: The permission system, sandboxing, and audit logging in Claude Code are deployment-level safeguards that complement model-level training.

  4. Iterative deployment: Anthropic releases capabilities incrementally, monitoring for misuse and adjusting safety measures. New tool types or permissions are added gradually.

As a Claude Code user, you benefit from this pipeline automatically. Your responsibility is to configure the runtime safety layers (permissions, sandboxing, access controls) appropriately for your environment.

The Balance Between Safety and Productivity ​

Safety and productivity are not opposed -- they are complementary when designed well:

Safety MeasureProductivity ImpactNet Effect
Read permissions allowed globallyEliminates prompts for safe operationsPositive
Destructive commands deniedPrevents catastrophic mistakesPositive
Per-command approval for unknownsSlight friction, high protectionPositive
All commands deniedUnusable -- constant promptingNegative
All commands allowedFast but dangerousNegative

The goal is to find the configuration that maximizes the area under both curves. For most teams, this means: allow reads, allow known-safe writes and commands, deny known-dangerous operations, and ask about everything else.

Key Takeaways ​

  1. Claude's refusals are a feature, not a bug. They come from multiple training stages designed to prevent harmful actions. Provide more context if a refusal seems incorrect.
  2. Defense in depth works because each layer covers different failure modes. Model alignment, permissions, sandboxing, and git checkpoints each catch different categories of errors.
  3. Least privilege is the foundation of safe agent design. Give agents only the permissions they need for their specific task.
  4. Never let agents modify their own configuration. Self-modifying agents can escape safety constraints.
  5. Human review before deployment is non-negotiable. Even correct code can have unintended side effects.
  6. Audit everything. In team and CI/CD environments, logs are your safety net for both debugging and accountability.

See also: A08 Constitutional AI & CLAUDE.md for how CLAUDE.md implements project-level AI governance, and A04 Agent Architecture Patterns for how safety integrates with agent design.

Released under MIT License