A15: Claude Code Internals
Related chapters: Ch3 Context Window & Session Management, Ch14 Advanced Automation & CI/CD
Claude Code is not a thin wrapper around the Claude API. It is a sophisticated runtime -- a harness -- that manages the agentic loop, orchestrates tool execution, maintains conversation state, performs context compaction, and coordinates between you and the model. Understanding these internals helps you debug unexpected behaviors, optimize your workflows, and use Claude Code at its full potential.
The Harness Loop
At its core, Claude Code runs a loop. Every interaction follows the same cycle:
┌──────────────────────────────────────────────────────┐
│ The Harness Loop │
│ │
│ ┌─────────┐ │
│ │ User │ │
│ │ Message │ │
│ └────┬────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────┐ │
│ │ Assemble Full Prompt │ │
│ │ • System prompt (Claude Code's identity) │ │
│ │ • CLAUDE.md content │ │
│ │ • Tool definitions │ │
│ │ • Conversation history │ │
│ │ • System reminders │ │
│ │ • User message │ │
│ └────────────────────┬────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────┐ │
│ │ Send to Claude API │ │
│ └────────────────────┬────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────┐ │
│ │ Parse Response │ │
│ │ • Text content → display to user │ │
│ │ • Tool calls → execute tools │ │
│ │ • Stop reason → check if done │ │
│ └────────────────────┬────────────────────────┘ │
│ │ │
│ ┌───────┴───────┐ │
│ │ │ │
│ Tool calls? end_turn? │
│ │ │ │
│ ▼ ▼ │
│ Execute tools Display final text │
│ Collect results Wait for next user message │
│ │ │
│ ▼ │
│ Append tool results to conversation │
│ │ │
│ └──────── Loop back to API call │
│ │
└──────────────────────────────────────────────────────┘One "Turn" in Detail
A single user message often triggers multiple API calls. Here is a concrete example:
User: "Fix the failing test in auth.test.ts"
API Call 1:
Claude responds: "Let me look at the failing test."
Tool call: Read("src/tests/auth.test.ts")
→ Harness executes Read, returns file contents
API Call 2:
Claude responds: "I see the issue. The mock is outdated."
Tool call: Read("src/auth/service.ts")
→ Harness executes Read, returns file contents
API Call 3:
Claude responds: "The function signature changed. Let me update the test."
Tool call: Edit("src/tests/auth.test.ts", old_string, new_string)
→ Harness executes Edit, returns success
API Call 4:
Claude responds: "Now let me run the test to verify."
Tool call: Bash("npm test -- --grep auth")
→ Harness executes Bash, returns test output
API Call 5:
Claude responds: "The test passes now. Here's what I changed..."
Stop reason: end_turn
→ Harness displays final response, waits for userEach API call in this sequence includes the full conversation history up to that point -- including all previous tool calls and their results. This is how Claude maintains context across the loop.
Why This Architecture Matters
Understanding the loop explains several Claude Code behaviors:
- Why long sessions get expensive: Each API call sends the full history. As the conversation grows, every call costs more.
- Why Claude sometimes re-reads files: If the conversation was compacted, Claude may have lost the file contents and needs to re-read.
- Why tool execution is sequential: The harness executes one tool at a time (though Claude can request multiple tools in a single response, which are then batched).
- Why interrupting Claude works: You can press Ctrl+C to stop the current API call. The harness simply does not send the next request.
Session Management and Persistence
How Sessions Work
A session is a conversation between you and Claude Code. It has:
- A unique session ID
- A conversation history (all messages, tool calls, and results)
- Associated metadata (model, start time, working directory)
~/.claude/sessions/
├── abc123-session-id/
│ ├── conversation.json # Full message history
│ ├── metadata.json # Session configuration
│ └── summary.json # Compacted summary (if compaction occurred)How --resume Works
When you run claude --resume, the harness:
- Looks up the most recent session (or a specific session ID if provided)
- Loads the conversation history from disk
- Reconstructs the full prompt as if the conversation had been continuous
- Sends the next user message with the full history
# Resume the most recent session
claude --resume
# Resume a specific session by ID
claude --resume abc123-session-id
# List recent sessions to find the one you want
claude --resume # Shows a picker if multiple sessions existWhat is preserved: All messages, tool calls, tool results, and any compacted summaries.
What is NOT preserved: The actual state of your filesystem. If you modified files between sessions, Claude will not know about those changes unless it re-reads the files.
Practical implication: After resuming, if the codebase has changed, tell Claude:
I've made some changes since our last session. Please re-read src/auth/
before continuing.How /compact Actually Works
When a conversation grows long, it consumes more tokens per API call and eventually approaches the context window limit. The /compact command triggers compaction -- a process that summarizes the conversation to free up context space.
The Compaction Process
Before compaction:
┌──────────────────────────────────────┐
│ System prompt (~2K tokens) │
│ CLAUDE.md (~1K tokens) │
│ Turn 1: message + tools (~5K tokens) │
│ Turn 2: message + tools (~8K tokens) │
│ Turn 3: message + tools (~12K tokens)│
│ Turn 4: message + tools (~6K tokens) │
│ Turn 5: message + tools (~9K tokens) │
│ ─────────────────────────────────────│
│ Total: ~43K tokens │
└──────────────────────────────────────┘
After compaction:
┌──────────────────────────────────────┐
│ System prompt (~2K tokens) │
│ CLAUDE.md (~1K tokens) │
│ [Summary of turns 1-4] (~2K tokens) │
│ Turn 5: message + tools (~9K tokens) │
│ ─────────────────────────────────────│
│ Total: ~14K tokens │
└──────────────────────────────────────┘What Happens During Compaction
- The harness identifies which turns to summarize (typically older turns, preserving the most recent)
- It sends the turns to be summarized to the Claude API with a summarization prompt
- The API returns a condensed summary capturing the key decisions, changes made, and current state
- The harness replaces the original turns with the summary in the conversation history
- Subsequent API calls use the compacted history
What Survives Compaction
The summary preserves:
- What files were modified and what the changes were (conceptually)
- Key decisions made during the conversation
- Current state of the task (what is done, what remains)
- Errors encountered and how they were resolved
The summary loses:
- Exact file contents that were read (Claude will need to re-read)
- Verbatim tool outputs (command outputs, test results)
- Nuance in reasoning (the detailed thought process is compressed)
Automatic vs Manual Compaction
# Manual compaction (you decide when)
> /compact
# Compaction with a custom instruction
> /compact Focus on preserving the database migration decisions
# Automatic compaction happens when context approaches the limit
# The harness triggers it transparentlyTip: If you are working on a task with critical context (e.g., a complex debugging session where the exact error messages matter), use /compact manually with an instruction to preserve specific details:
/compact Preserve the exact error messages from the test failures
and the database connection pooling configuration values.The Checkpoint System
Claude Code uses git to create safety checkpoints that allow you to undo changes.
How Checkpoints Work
Before Claude makes changes:
1. Harness creates a git stash or commit of current state
2. Records the checkpoint ID
Claude makes changes:
3. Edit/Write/Bash tool calls modify files
If you want to undo:
4. /undo restores the checkpoint
5. Files return to pre-change stateCheckpoint Internals
Checkpoints are implemented using git's internal mechanisms:
# What the harness does internally (simplified):
# Before changes: snapshot current state
git stash push -m "claude-checkpoint-$(date +%s)" --include-untracked
# After changes: if user wants to undo
git stash pop # Restores the saved stateFor more complex scenarios (changes across multiple tool calls), the harness may create lightweight commits on a temporary ref:
refs/claude/checkpoints/
├── cp-001 (state before first edit)
├── cp-002 (state before second edit)
└── cp-003 (state before bash command)Important: Checkpoints only cover tracked and untracked files in the git repository. Files outside the repo, environment variables, database state, and running processes are not covered.
The /undo Command
> /undo # Undo the most recent change
> /undo 3 # Undo the last 3 changesEach /undo restores the filesystem to the state before the corresponding checkpoint. This is why Claude Code requires a git repository -- git is the undo mechanism.
The Extension Architecture
Claude Code runs in multiple environments, each with a different extension:
┌─────────────────────────────────────────────────────────┐
│ Claude Code Core │
│ (harness loop, tool execution, session management) │
└──────────────┬──────────────────────────────────────────┘
│
┌──────────┼──────────┬──────────┬──────────┐
│ │ │ │ │
▼ ▼ ▼ ▼ ▼
┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐
│ CLI │ │VS Code │ │JetBrains│ │Desktop │ │ Web │
│Terminal│ │Extension│ │ Plugin │ │ App │ │ App │
└────────┘ └────────┘ └────────┘ └────────┘ └────────┘| Extension | Unique Capabilities |
|---|---|
| CLI (Terminal) | Full terminal control, pipe integration with Unix tools (echo "..." | claude -p), scriptable for CI/CD |
| VS Code Extension | Inline diff view, per-change accept/reject, file tree showing agent activity |
| JetBrains Plugin | Integration with IntelliJ/PyCharm/WebStorm diff viewer and run configurations |
| Desktop App | Standalone Electron app, no terminal or IDE required, simplified setup |
| Web Interface | Browser-based, no local installation, runs against a remote environment |
All extensions share the same core engine. The differences are in the presentation layer (how results are displayed) and the input layer (how you interact). A session started in VS Code can be resumed in the CLI because the conversation state is the same format.
How CI/CD Mode Differs from Interactive Mode
Claude Code has two fundamental modes of operation:
| Aspect | Interactive | CI/CD (-p flag) |
|---|---|---|
| Permission handling | Ask user | Pre-configured only |
| Clarifying questions | Supported | Not possible |
| Session length | Open-ended | Single task |
| Output format | Human-readable | Machine-parseable (JSON option) |
| Error recovery | Human can guide | Must self-recover or fail |
CI/CD mode is triggered by the -p flag or piped input:
# CI/CD mode: pre-configure permissions, no interactive prompts
claude -p "Review the diff" \
--model claude-sonnet-4-6-20250514 \
--output-format json \
--permission-mode deny-all \
--allowedTools "Read,Glob,Grep,Bash(git diff*)"The --permission-mode deny-all flag is critical: it ensures no interactive prompts. A permission prompt in a CI pipeline means a hung job.
Environment Variables and Their Effects
Claude Code's behavior is influenced by several environment variables:
| Variable | Purpose | Example |
|---|---|---|
ANTHROPIC_MODEL | Select the model | claude-sonnet-4-6-20250514 |
ANTHROPIC_API_KEY | API authentication | sk-ant-... |
ANTHROPIC_BASE_URL | Custom API endpoint/proxy | https://api.anthropic.com |
CLAUDE_MAX_COST | Maximum session cost (safety limit) | 10.00 |
CLAUDE_MAX_TURNS | Maximum turns before stopping | 100 |
CLAUDE_CODE_DISABLE_TELEMETRY | Disable telemetry | 1 |
CLAUDE_WORKING_DIR | Override working directory | /path/to/project |
CLAUDE_CODE_VERBOSE | Enable verbose logging | 1 |
GITHUB_TOKEN | GitHub token for PR operations | ghp_... |
CI | Indicate running in CI | true |
NO_COLOR | Disable color output | 1 |
How Variables Are Resolved
Environment variables are resolved in this order (later overrides earlier):
1. System environment (lowest priority)
2. .env file in project root
3. ~/.claude/.env (user-level)
4. CLI flags (highest priority)The Difference Between claude-code (CLI) and claude-code-sdk (Library)
Claude Code ships as two packages that serve different purposes:
claude-code (CLI)
# Install
npm install -g @anthropic-ai/claude-code
# Use interactively
claude
# Use in scripts
claude -p "Do something"The CLI is a complete application: it includes the harness loop, tool implementations, permission system, session management, and the terminal UI. It is what you use directly.
claude-code-sdk (Library)
// Install
// npm install @anthropic-ai/claude-code-sdk
import { ClaudeCode } from "@anthropic-ai/claude-code-sdk";
// Use programmatically
const claude = new ClaudeCode({
model: "claude-sonnet-4-6-20250514",
workingDirectory: "/path/to/project",
});
const result = await claude.run("Review this codebase for security issues");
console.log(result.text);The SDK is a library for embedding Claude Code in your own applications. It provides the same core capabilities (harness loop, tools, permissions) but without the terminal UI. You control:
- When sessions start and stop
- How output is displayed
- How permissions are handled
- How errors are surfaced
When to Use Which
| Use Case | CLI | SDK |
|---|---|---|
| Daily development work | Yes | No |
| CI/CD pipelines (simple) | Yes (-p flag) | No |
| CI/CD pipelines (complex orchestration) | No | Yes |
| Custom IDE extensions | No | Yes |
| Building AI-powered dev tools | No | Yes |
| Multi-agent systems | Yes (multiple processes) | Yes (programmatic control) |
| Batch processing | Yes (scripting) | Yes (better control) |
The SDK gives you programmatic control over every aspect of Claude Code's behavior. The CLI gives you a ready-to-use tool. Most developers use the CLI; tool builders use the SDK.
Configuration Resolution
Claude Code reads configuration from multiple sources, merged in a specific order:
1. Default values (built into Claude Code)
│
▼
2. Enterprise settings (managed by organization admin)
│ ~/.claude/settings.json (enterprise-managed)
▼
3. User settings
│ ~/.claude/settings.json (user-managed section)
▼
4. Project settings
│ .claude/settings.json (in the repo)
▼
5. CLAUDE.md files (project root, then nested)
│
▼
6. CLI flags and environment variables
│ (highest priority -- overrides everything)
▼
Final resolved configurationConflict resolution: For most settings, the later source wins. For permissions, the rules are merged (deny from any level takes precedence over allow).
// Enterprise policy: deny destructive operations
// ~/.claude/settings.json (enterprise)
{
"permissions": {
"deny": ["Bash(rm -rf*)", "Bash(git push --force*)"]
}
}
// Project settings: allow project-specific commands
// .claude/settings.json
{
"permissions": {
"allow": ["Bash(npm test*)", "Bash(npm run build*)"]
}
}
// Resolved: deny rules from enterprise + allow rules from project
// Enterprise deny rules CANNOT be overridden by project settingsThis hierarchy ensures that organizational security policies are enforced even when individual projects configure more permissive settings.
Performance Optimization
The dominant latency in any Claude Code interaction is model inference (1-30 seconds depending on model and complexity). Prompt assembly, network, and tool execution are secondary. To optimize:
- Reduce input tokens: Keep CLAUDE.md concise, use
/compactbefore context grows large, point Claude to specific files instead of letting it search - Reduce round trips: Give Claude enough context to act in fewer turns, pre-grant permissions for common operations
- Use prompt caching: The system prompt and CLAUDE.md are cached across turns, so longer CLAUDE.md files benefit MORE from caching (see A10 Prompt Caching)
Key Takeaways
- The harness loop (prompt assembly, API call, tool execution, repeat) is the fundamental architecture. Understanding it explains most Claude Code behaviors.
- Sessions are persistent and can be resumed. But filesystem state is not tracked -- tell Claude about external changes.
- Compaction preserves decisions and state but loses exact file contents and command outputs. Use manual compaction with preservation instructions for critical context.
- Checkpoints are git-based. The
/undocommand is powered by git stash/refs, which is why a git repo is required. - CI/CD mode removes all interactive elements. Pre-configure permissions and use bounded execution to prevent hung pipelines.
- The CLI is for using Claude Code; the SDK is for building with it. Most developers use the CLI; tool builders use the SDK.
- Configuration is layered: enterprise > user > project > CLI flags, with deny rules taking precedence at all levels.
See also: A05 Tool Calling Internals for how Claude decides which tools to call, and A10 Prompt Caching for optimizing the repeated parts of the harness loop.