A02: Context Engineering
Related chapters: Ch3 Context Window Management, Ch15 Production Workflows
Prompt engineering asks "what should I say to the model?" Context engineering asks a bigger question: "what should the model see when it generates a response?" This appendix covers the paradigm shift from prompt engineering to context engineering, and provides practical strategies for managing all inputs to Claude Code -- not just the text you type.
The Paradigm Shift: Prompt Engineering to Context Engineering
The Prompt Engineering Era (2022-2024)
When ChatGPT launched, the world discovered prompt engineering. Users learned that phrasing mattered: "think step by step" outperformed "give me the answer." An entire discipline emerged around crafting the perfect instruction.
Prompt engineering treated the model as a text-in, text-out function. Your prompt was the only lever you could pull. Techniques like few-shot examples, role assignment ("you are a senior engineer"), and chain-of-thought all focused on optimizing that single text input.
This worked well for chat-style interactions. But it broke down when models became agents.
The Context Engineering Era (2025-2026)
When Claude Code and similar agentic tools arrived, the rules changed. Your typed prompt became a small fraction of what the model actually sees. In a typical Claude Code session, the model's input includes:
Token budget breakdown for a typical Claude Code session:
┌──────────────────────────────────────────────────────┐
│ System prompt .............. ~3,000 tokens (fixed) │
│ Tool definitions ........... ~5,000 tokens (fixed) │
│ CLAUDE.md files ............ ~2,000 tokens (fixed) │
│ Memory files ............... ~500 tokens (fixed) │
│ Conversation history ....... ~varies (grows) │
│ Tool results (file reads) .. ~varies (grows) │
│ Your current prompt ........ ~100 tokens (tiny) │
│ │
│ Your prompt is typically < 1% of total context │
└──────────────────────────────────────────────────────┘Context engineering recognizes this reality. Instead of optimizing one input, you architect the entire information environment around the model. This includes:
- What instructions persist across sessions (CLAUDE.md, memory)
- What information enters the context (which files Claude reads, which commands it runs)
- What information exits the context (compaction, summarization)
- When to reset the context (new sessions,
/clear)
Anthropic's own research team has emphasized this shift. As Toby Jia-Jun Li (Anthropic researcher) noted in 2025: the most impactful improvements to agent performance come not from better prompts, but from better context management.
The Context Stack: Layers of Input
Claude Code assembles its context from multiple layers, each with different characteristics:
Layer 1: System Prompt (You Don't Control This)
The system prompt is injected by Claude Code's harness before your conversation begins. It contains:
- Claude's core behavioral instructions
- Tool definitions with JSON schemas
- Permission rules and safety constraints
- The current date and environment info
This layer is fixed -- you cannot modify it. But understanding it helps you avoid redundant instructions. For example, you do not need to tell Claude to "use the Read tool to read files" because the system prompt already establishes tool-calling behavior.
Token cost: approximately 3,000-8,000 tokens depending on tool definitions and configuration.
Layer 2: CLAUDE.md Files (You Control This)
CLAUDE.md files are your primary lever for persistent context engineering. They are loaded at session start and remain in context throughout the conversation.
Resolution order (all loaded, later overrides earlier):
1. ~/.claude/CLAUDE.md (global preferences)
2. ~/project/CLAUDE.md (project root)
3. ~/project/.claude/CLAUDE.md (project config dir)
4. ~/project/src/CLAUDE.md (directory-level)Design principle: CLAUDE.md is instruction context -- rules, patterns, and constraints that Claude should always follow. It is NOT a place to dump documentation or code examples. Keep it lean.
# Good CLAUDE.md (focused, ~500 tokens)
## Code Style
- Use TypeScript strict mode
- Prefer functional patterns; avoid classes
- Error handling: return Result<T, Error>, never throw
## Testing
- Every new function needs a test in __tests__/
- Use vitest, not jest
- Mock external APIs with msw
## Architecture
- Routes in src/routes/, services in src/services/
- Database access only through src/db/ layer# Bad CLAUDE.md (bloated, ~5,000 tokens)
## Full API Documentation
[entire API spec pasted here]
## Database Schema
[full SQL schema pasted here]
## Meeting Notes
[notes from last sprint planning]Token cost: aim for 500-2,000 tokens. Every token in CLAUDE.md is loaded on every turn, so bloated CLAUDE.md files compound costs quadratically over a session.
Layer 3: Memory Files (Claude Manages This)
Memory files store Claude's learned preferences about you and your project. They persist between sessions and are loaded automatically.
Memory is useful for preferences that Claude discovers during work -- your preferred variable naming, how you like commit messages formatted, which test runner you use. But memory is not a substitute for CLAUDE.md. CLAUDE.md is for project rules that all developers on the team should follow. Memory is for your personal preferences.
Token cost: typically 200-800 tokens.
Layer 4: Conversation History (Both of You Build This)
This is where context engineering becomes critical. Conversation history grows with every exchange:
Turn 1: You ask a question +50 tokens
Claude reads 3 files +3,000 tokens
Claude responds +500 tokens
Turn 2: You ask a follow-up +30 tokens
Claude reads 2 more files +2,000 tokens
Claude edits a file +400 tokens
Claude runs tests +1,500 tokens
Claude responds +300 tokens
Turn 3: Already at ~8,000 tokens of history
...
Turn 20: Context is 60-80% fullThe conversation history is the fastest growing layer and the one you have the most control over through your workflow choices.
Layer 5: Tool Results (Generated by Actions)
Every Read, Bash, Grep, or other tool call produces results that enter the context. A single cat of a large file can inject thousands of tokens. A test suite's output can inject tens of thousands.
This is often the most wasteful layer. Claude reads files it does not need, runs verbose commands when terse alternatives exist, and explores broadly when a targeted search would suffice.
The Context Budget: Allocating Tokens Wisely
Think of the 200K token context window as a budget. Every token spent on one thing is a token unavailable for another.
The Golden Ratio: Instruction Context vs Working Context
Context engineering divides the window into two categories:
- Instruction context: system prompt, CLAUDE.md, memory, tool definitions. This tells Claude how to behave.
- Working context: conversation history, tool results, the current task. This is the material Claude works with.
The golden ratio:
┌──────────────────────────────────────────────────────┐
│ Instruction context: 10-20% of window │
│ ████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ │
│ │
│ Working context: 30-50% of window │
│ ░░░░░░░░██████████████████████████░░░░░░░░░░░░░░░ │
│ │
│ Reserved headroom: 30-50% of window │
│ ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░████████████████ │
│ │
│ Total used: ideally 40-60% (the golden zone) │
└──────────────────────────────────────────────────────┘Why reserve headroom? Claude needs room to generate its response (output tokens) and to call tools that produce results (which re-enter the context). If the window is 90% full, Claude cannot read another file without triggering compaction or losing information.
Budgeting by Task Type
Different tasks have different context profiles:
| Task Type | Instruction Needs | Working Context Needs | Recommended Approach |
|---|---|---|---|
| Simple edit | Low (just file location) | Low (one file) | Direct prompt, single session |
| Bug fix | Medium (symptoms, constraints) | Medium (a few files + test output) | Investigation-first prompt |
| Feature build | Medium (requirements, patterns) | High (many files, tests) | Checkpoint frequently |
| Architecture review | Low (just "review") | Very high (many files to read) | Split across sessions |
| Large refactor | High (constraints, safety) | Very high (many files to edit) | Multi-session with plan |
Practical Strategies
Strategy 1: Summarize Before Compact
When Claude's automatic compaction triggers, it summarizes the conversation to fit the window. But automatic compaction is lossy -- it decides what to keep based on recency and apparent relevance, which may not match your priorities.
The pattern: manually summarize before compaction would trigger:
You (at ~50% context usage):
"Before we continue, summarize what we've done so far and what remains.
Include: (1) the architecture decisions we made, (2) the files we've
changed, (3) the remaining tasks. Then I'll /clear and we'll continue
with your summary."Then start a fresh session:
You (new session):
"Continuing a refactor. Here's where we left off:
[paste Claude's summary]
Next task: implement the database migration for the new schema."This gives you control over what survives the context reset, rather than leaving it to automatic compaction.
Strategy 2: Checkpoint Before Explore
Before embarking on an exploration that might consume lots of context (reading many files, running extensive tests), create a checkpoint:
You: "We're about to investigate the payment processing pipeline.
Before we start, confirm: we've completed the auth refactor (files:
src/auth/service.ts, src/auth/middleware.ts, src/auth/types.ts), all
auth tests pass, and the remaining work is the payments module. Correct?"
Claude: "Yes, that's correct. [confirms details]"
You: "Good. Now investigate the payments pipeline. Start with
src/payments/index.ts."The checkpoint serves two purposes:
- It creates a summarized anchor point in the conversation history that compaction is likely to preserve
- It verifies shared understanding before the context gets cluttered with new information
Strategy 3: One Task Per Session
The most powerful context engineering strategy is also the simplest: do one thing per session.
# Session 1: Fix the auth bug
claude "Fix the 401 error in token rotation. The issue is in src/auth/."
# Session 2: Add rate limiting
claude "Add rate limiting to /api/upload. Max 10/minute per user. Use Redis."
# Session 3: Update tests
claude "The auth middleware changed. Update tests in tests/auth/ to match."Each session starts with a clean context window. Claude has maximum headroom for working context. There is no risk of context rot from earlier tasks polluting later ones.
When single-task sessions are not practical (for example, multi-step features), use the summarize-before-compact pattern to create natural session boundaries within a longer conversation.
Strategy 4: Front-Load Critical Context
The attention mechanism in Transformer models is not uniform across the context window. Research (including Anthropic's own studies on long-context retrieval) shows that models attend most strongly to:
- The very beginning of the context (system prompt, CLAUDE.md)
- The very end of the context (your most recent prompt)
- Distinctive or structured content (headers, code blocks, formatted lists)
Content in the middle of a long context receives less attention. This is sometimes called the "lost in the middle" effect (Liu et al., 2023).
Practical implications:
- Put your most important rules in CLAUDE.md (always at the top of context)
- Repeat critical constraints in your current prompt (always at the bottom of context)
- Use structured formatting (headers, bullet points) for instructions that must not be missed
- Do not rely on a casual mention from turn 5 of a 30-turn conversation being remembered
# Effective: critical constraint repeated at point of action
You (turn 15): "Now implement the payments endpoint. IMPORTANT: do NOT
modify the existing /api/users endpoint -- it's in production and any
change will break the mobile app."Strategy 5: Control Tool Result Size
Tool results are often the largest and most wasteful entries in the context. You can influence their size:
# Bad: reads entire file when you only need the function signature
You: "Read the auth module and find the validateToken function."
# Claude reads the entire 500-line file (+3,000 tokens)
# Good: directs Claude to search, not read
You: "Find the validateToken function signature in src/auth/.
Use grep, don't read the whole file."
# Claude greps for the function (+200 tokens)# Bad: runs verbose test output
You: "Run the tests."
# Claude runs pytest with full output (+5,000 tokens)
# Good: controls output verbosity
You: "Run pytest with -q flag. Only show failures."
# Claude runs pytest -q (+500 tokens)This may seem like micro-optimization, but over a 20-turn session, controlling tool result size can be the difference between staying in the golden zone and triggering context rot.
Context Rot: Why Quality Degrades
What Context Rot Looks Like
Context rot is the progressive degradation of Claude's output quality as the context window fills up. It manifests as:
- Forgotten requirements: Claude ignores constraints you stated earlier
- Contradictory outputs: Claude does the opposite of what it agreed to in a previous turn
- Hallucinated context: Claude references files or functions that do not exist (confusing earlier file contents with current state)
- Decreasing specificity: Claude's responses become more generic and less tailored to your codebase
- Increased repetition: Claude re-reads files it already read, or re-explains things it already explained
Why It Happens
Context rot has multiple interacting causes:
Attention dilution: As the context grows, each token competes with more tokens for attention weight. Important instructions from turn 2 get less attention when they are surrounded by 100K tokens of file contents and test output.
Signal-to-noise ratio: Early in a session, the context is mostly signal (your instructions, relevant code). Late in a session, it is mostly noise (intermediate results, abandoned approaches, verbose output). The model struggles to extract the signal.
Positional bias: The "lost in the middle" effect means that information from the middle turns of a long conversation is most likely to be overlooked. Your critical constraint from turn 5 may be invisible by turn 20.
Compaction artifacts: When automatic compaction triggers, it summarizes earlier turns. These summaries inevitably lose nuance. After multiple compaction cycles, the model's understanding of the original requirements may have drifted significantly.
How to Combat Context Rot
| Strategy | When to Use | How It Helps |
|---|---|---|
| One task per session | Always, when feasible | Prevents rot entirely |
/clear between tasks | When doing multiple tasks | Resets context between unrelated work |
| Manual summarization | At ~50% context usage | You control what survives |
| Repeat critical constraints | Before each major action | Refreshes attention to important rules |
| Short sessions | For high-precision work | Less time for rot to accumulate |
Pipe mode (| claude -p) | For independent subtasks | Each invocation gets clean context |
Advanced: Context Engineering for Teams
Shared CLAUDE.md as Team Context
When a team shares a CLAUDE.md file in the repository, they create a shared context engineering layer. Every team member's Claude Code session starts with the same instructions, patterns, and constraints.
This is powerful but requires discipline:
# Team CLAUDE.md best practices
## DO include:
- Code style rules that apply to ALL files
- Architecture patterns that ALL code must follow
- Testing requirements that ALL changes must meet
- File naming and organization conventions
## DO NOT include:
- Individual preferences (use personal ~/.claude/CLAUDE.md)
- Temporary instructions ("until the migration is done")
- Documentation that belongs in docs/ or README
- Rules for code that no longer existsContext Engineering in CI/CD
When Claude Code runs in CI/CD (pipe mode), context engineering takes a different form. There is no interactive conversation -- you have exactly one prompt. Every token must count:
# CI/CD context engineering: front-load everything
git diff main..HEAD | claude -p "$(cat <<'EOF'
You are reviewing a pull request. The diff is piped as input.
Project rules:
- TypeScript strict mode, no any types
- All public functions need JSDoc
- No console.log in production code (use the logger)
- SQL queries must use parameterized statements
Review for:
1. Rule violations from the list above
2. Logic bugs
3. Security issues (OWASP Top 10)
Output format: one finding per line, with severity (HIGH/MEDIUM/LOW),
file path, line number, and description.
EOF
)"In this pattern, the system prompt and CLAUDE.md are replaced by an explicit context block in the pipe prompt. The diff provides the working context. There is no conversation history to manage.
Context Engineering Checklist
Use this checklist when starting a new project or optimizing an existing workflow:
- [ ] CLAUDE.md is under 2,000 tokens and contains only active rules
- [ ] No documentation or code examples are pasted into CLAUDE.md
- [ ] Complex tasks are broken into single-session units
- [ ] Critical constraints are repeated in the prompt at the point of action
- [ ] Tool results are controlled (grep before read, quiet flags on tests)
- [ ] Context usage is monitored with
/costor/context - [ ] Manual summarization happens before the 60% threshold
- [ ] Team members share a reviewed CLAUDE.md in the repository
- [ ] Personal preferences are in
~/.claude/CLAUDE.md, not the project file
Key Takeaways
- Context engineering is about all inputs, not just the prompt. Your prompt is less than 1% of what Claude sees. Optimize the other 99%.
- The golden zone is 40-60% context usage. Past 60%, quality degrades measurably. Past 80%, it degrades severely.
- One task per session is the most powerful strategy. Clean context beats clever context management every time.
- CLAUDE.md is your persistent instruction layer. Keep it lean, focused, and maintained. Treat it like code -- review it, version it, prune it.
- Control tool result size aggressively. Use grep instead of read, quiet flags on commands, and targeted searches instead of broad exploration.
- Context rot is real and predictable. Plan for it by checkpointing, summarizing, and resetting at natural boundaries.
See also: A01 Prompt Engineering for optimizing the prompt layer specifically, and A07 Token Economics for understanding the cost implications of context budget decisions.