Chapter 3: Context Window Management
Learning Objectives
- Understand why Claude gets worse over long sessions (it's not a bug, it's context)
- Monitor consumption with
/context,/usage, and/cost - Know exactly when to
/compactand what you lose - Use
--resumeand session management to pick up where you left off - Internalize the 40-60% golden zone for optimal output quality
Appendix links: A03 LLM Fundamentals explains the Transformer architecture and attention mechanism that underlie everything in this chapter. A07 Token Economics covers BPE tokenization, cost models, and budget optimization. A10 Prompt Caching details how Claude Code reuses cached prefixes to reduce cost and latency.
Concepts
The Context Window is Claude's Working Memory
Everything in your conversation -- your prompts, Claude's responses, every file it reads, every command output -- lives in a single memory buffer called the context window. When it fills up, Claude starts forgetting things or producing lower quality output.
┌──────────────────────────────────────────────────┐
│ Context Window (200K tokens) │
│ │
│ ┌──────────┐ ┌──────────┐ ┌───────────┐ │
│ │ System │ │ CLAUDE.md│ │ Memory │ Fixed │
│ │ Prompt │ │ │ │ Files │ cost │
│ └──────────┘ └──────────┘ └───────────┘ │
│ │
│ ┌────────────────────────────────────────┐ │
│ │ Conversation History + Tool Results │ │
│ │ │ │
│ │ You: "Analyze the routing module" │ │
│ │ Claude: [Read 5 files] -> 2000 lines │ │
│ │ Claude: "Here's how routing works..." │ │
│ │ You: "Now add tests" │ │
│ │ Claude: [Write test file] │ Grows │
│ │ Claude: [Bash] pytest -> output │ │
│ │ ...keeps growing... │ │
│ └────────────────────────────────────────┘ │
│ │
│ ████████████████████░░░░░░░ 68% used │
└──────────────────────────────────────────────────┘Why Context Windows Have Limits
Context windows are not arbitrary product restrictions -- they are a fundamental consequence of how Transformer models work. The core of every LLM is the attention mechanism (Attention), which lets each token "look at" every other token in the sequence. This creates quadratic scaling: processing N tokens requires roughly N squared operations. Double the context from 100K to 200K tokens, and the computation cost roughly quadruples.
This is why even with 200K or 1M token windows, there are practical limits. The model needs increasingly more memory and compute as the context grows. Hardware constraints impose a ceiling.
Deep dive: A03 LLM Fundamentals covers the Transformer architecture and attention mechanism in detail, including recent advances like sliding window attention and sparse attention that help extend context lengths.
What Are Tokens?
Before we discuss context budgets, you need to understand what tokens actually are. Tokens are not words and not characters -- they are subword units produced by Byte Pair Encoding (BPE) tokenization. The tokenizer splits text into the most statistically frequent subword pieces from its training data.
A rough rule of thumb: 1 token is approximately 3-4 characters in English, or about 0.75 words. For code, the ratio varies:
"hello world" → 2 tokens (common words = 1 token each)
"implementation" → 1 token (common long word)
"XMLHttpRequest" → 4 tokens (camelCase splits)
"const x = arr.map()" → 8 tokens (code has more splits)
"你好世界" → 4 tokens (Chinese: ~1-2 tokens per character)This matters for context management because code is token-expensive: a 200-line Python file might be 3000 tokens, not the 1000 you might guess from word count.
Deep dive: A07 Token Economics explains BPE tokenization in full, including how to estimate costs before running a session.
Context Rot: The Silent Killer
GSD (50k stars) coined a term that every Claude Code user should know: Context Rot.
Their research found that once context usage passes ~50%, Claude's output quality starts degrading. By 70%, it's noticeably worse. By 80%, you're getting outputs that miss requirements, introduce bugs, and forget earlier instructions.
Quality vs. Context Usage:
100% |████
|████████
|████████████
|████████████████
|████████████████████
|██████████████████████████
|████████████████████████████████
|██████████████████████████████████████████
0% |____________________________________________
0% 20% 40% 60% 80% 100%
Context usage -->
The sweet spot is 40-60%. Past that, quality drops fast.This is not theoretical. HumanLayer's research across thousands of sessions confirmed the 40-60% golden zone -- Claude performs best when the context window is less than 60% full. Past that, you're fighting degrading returns.
Important caveat: The 40-60% golden zone is a widely reported observation from the Claude Code community (GSD, HumanLayer, and many practitioners), not a peer-reviewed academic finding. Your mileage may vary depending on task complexity and content type. The key insight -- that quality degrades well before the window is full -- is consistent across reports.
The "Needle in a Haystack" Effect
Research on long-context models (the "needle in a haystack" tests) shows that model performance is not uniform across the context window. Models tend to attend most strongly to:
- The beginning of the context (system prompt, CLAUDE.md) -- highest attention
- The most recent turns -- strong recency bias
- The middle of the context -- weakest attention, details here are most likely to be "forgotten"
This is called the "lost in the middle" phenomenon. It means that a critical decision you made 30 turns ago, buried in the middle of the conversation, is the most likely thing Claude will lose track of -- even if the context window is not full yet.
Practical implication: if you have important constraints or decisions, put them in CLAUDE.md (always at the top of context) or reiterate them in your current prompt, rather than relying on Claude remembering something from turn 12 of a 40-turn session.
What Eats Tokens?
| Content | Approximate Cost | Impact |
|---|---|---|
| 1 line of code | ~10-15 tokens | Reads add up fast |
| 1 file (200 lines) | ~2,000-3,000 tokens | A few files = 5-10% of context |
| 1 conversation turn | ~100-500 tokens | Depends on prompt length |
| Bash command output | Varies wildly | npm install output can burn 5,000+ tokens |
| CLAUDE.md | ~200-500 tokens | Loaded every turn |
| Memory files | ~100-300 tokens | Loaded every turn |
The biggest token sink is usually reading files. Every time Claude reads a file to answer your question, that entire file content goes into the context and stays there.
Context Sizes by Plan
| Plan | Context Window | Good For |
|---|---|---|
| Standard (Sonnet) | 200K tokens | Most tasks, moderate sessions |
| Max/Team/Enterprise + Opus 4.6 | 1M tokens | Huge codebases, extended analysis |
What Happens at the Limit
At ~83.5% context usage, Claude Code triggers auto-compaction: it compresses the conversation history into a summary, preserving CLAUDE.md and memory files but losing details from earlier in the conversation.
Here is what auto-compaction looks like in your terminal:
$ [auto-compaction triggered]
⚠ Context usage reached 83.5% (167,000 / 200,000 tokens)
Compacting conversation history...
Preserved:
✓ CLAUDE.md (full)
✓ Memory files (full)
✓ Last 2-3 conversation turns (full)
Compressed:
↓ 47 earlier turns → summary (2,100 tokens)
↓ 12 file reads → key findings only
↓ 8 bash outputs → result summaries
Context after compaction: 24,800 / 200,000 tokens (12.4%)
ℹ Tip: Use /compact proactively at 40-50% for better summaries.The problem? By the time auto-compaction kicks in, you've already been in the degraded zone for a while. That's why proactive management matters.
Demo 6: Watch Context Fill Up on a Real Codebase
Goal
Clone Flask, have Claude analyze its routing system, and watch the context window fill up in real time. Then compact and see the difference.
Steps
1. Set Up
We'll use pallets/flask (68k+ stars). Flask's routing system is complex enough to burn through tokens quickly.
mkdir -p ~/claude-demos/demo-06 && cd ~/claude-demos/demo-06
git clone --depth 1 https://github.com/pallets/flask.git
cd flask
cat > CLAUDE.md << 'EOF'
# Flask Context Demo
- Focus on the routing and request handling subsystems
- Keep explanations concise
EOF2. Baseline Measurement
claude/contextHere is what the /context output looks like at the start of a session:
$ /context
Context usage: 9,400 / 200,000 tokens (4.7%)
Breakdown:
System prompt: 4,200 tokens
CLAUDE.md: 180 tokens
Memory files: 120 tokens
Conversation: 4,900 tokens
░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ 4.7%
Status: Healthy — plenty of room.Check your quota/cost too:
/usage(Or /cost if you're on an API key.)
3. Start Burning Context
Ask Claude to read and analyze Flask's routing:
Read the core routing files in src/flask/ -- I want to understand how Flask registers routes and dispatches requests. Start with app.py and find the route() decorator implementation./contextYou should see something like this now:
$ /context
Context usage: 36,200 / 200,000 tokens (18.1%)
Breakdown:
System prompt: 4,200 tokens
CLAUDE.md: 180 tokens
Memory files: 120 tokens
Conversation: 31,700 tokens ← jumped from file reads + analysis
███░░░░░░░░░░░░░░░░░░░░░░░░░░░ 18.1%
Status: Healthy — good headroom remaining.Record the new percentage. Reading a few source files probably jumped you 5-15%.
4. Go Deeper
Now trace what happens when a request comes in. How does Flask match the URL to a route? Show me the code path from the WSGI entry point to the view function being called./contextYou should be climbing. Claude had to read more files, produce longer explanations, and the context is filling up.
5. One More Push
Analyze Flask's blueprint system. How do blueprints register their routes? What happens when there's a URL conflict between a blueprint route and an app route?/context6. The Same Task at 30% vs. 70%
This is the critical experiment. If you're above 50%, now do this:
/compact keep the routing analysis and blueprint explanationHere is what /compact output looks like:
$ /compact keep the routing analysis and blueprint explanation
Compacting with guidance: "keep the routing analysis and blueprint explanation"
Before: 104,600 / 200,000 tokens (52.3%)
After: 22,400 / 200,000 tokens (11.2%)
Preserved (per your instructions):
✓ Routing analysis (route() decorator, URL rule registration)
✓ Blueprint explanation (blueprint routes, conflict resolution)
✓ CLAUDE.md (always preserved)
✓ Memory files (always preserved)
Dropped:
✗ Raw file contents from app.py, blueprints.py, scaffold.py
✗ Intermediate reasoning and exploratory reads
✗ Bash command outputs
ℹ Context freed: 82,200 tokens (41.1%)/contextContext should drop significantly. Now ask:
Explain Flask's error handling system. How does @app.errorhandler work, and where in the request lifecycle do error handlers execute?Compare the quality of this answer to the responses you got when the context was fuller. You should notice the post-compact response is more focused, more structured, and catches more nuance.
What Just Happened?
Here is the tool-call sequence Claude used during the routing analysis, and why it chose each tool:
Context consumption timeline:
Start: [System + CLAUDE.md + Memory] = ~5%
Read routing files: = ~18%
Trace request handling: = ~35%
Analyze blueprints: = ~52%
/compact: = ~12%
Analyze error handling (fresh context): = ~20%The error handling analysis at 20% was almost certainly better than if you'd asked the same question at 65%. That's Context Rot in action.
Key lesson from GSD: "Long-running sessions are the enemy." Split your work into focused tasks, compact proactively, or start fresh sessions.
Demo 7: Session Management & Resuming
Goal
Learn to exit and resume sessions without losing context. This is essential for work that spans hours or days.
Steps
1. Create a Session Worth Resuming
cd ~/claude-demos/demo-06/flask
claudeHelp me design a middleware system inspired by Flask's approach. I want:
1. A middleware registration API
2. Before-request and after-request hooks
3. Error-handling middleware
Just produce the design doc, no code yet.After Claude delivers the design, note a few key details (module names, hook ordering), then exit:
Ctrl+C
2. Resume Where You Left Off
claude --resumeVerify it worked:
What was the hook ordering we decided on for the middleware system?Claude should recall everything from the previous session. --resume loads the full session log, not a compressed summary.
3. Browse Session History
Start a new session to see your history:
claude/sessionsThis lists recent sessions with timestamps and summaries. You can resume any of them.
4. Resume a Specific Session
# Each session has an ID shown in /sessions
claude --resume <session-id>What Just Happened?
Here is the tool-call sequence Claude used when designing the middleware system:
Key insight: Even though Claude is "just" producing a design document and not writing code, it still uses tools to ground its design in real implementation patterns. This is why tool calls consume context -- and why the design task itself uses significant tokens.
Key Difference
| Action | What Loads | Detail Level |
|---|---|---|
| New session | CLAUDE.md + Memory only | Fresh start, no prior conversation |
--resume | Full session log | Everything, as if you never left |
After /compact | Compressed summary + CLAUDE.md + Memory | Key points preserved, details lost |
When to use each:
- New session: Switching tasks, starting something unrelated
--resume: Picking up after a break, interrupted by a meeting/compact: Staying in the same session but running low on context
When Things Go Wrong
Context management is not always smooth. Here are the most common failure modes and how to handle them.
Context Overflow Mid-Task
Symptom: Claude is halfway through implementing a feature. Auto-compaction fires. Claude's next response misses requirements it was previously tracking, or it re-reads files it already analyzed.
⚠ Context usage reached 83.5%
Compacting conversation history...
...
Context after compaction: 24,800 / 200,000 tokens (12.4%)Then Claude says something like:
I'll continue implementing the feature. Let me re-read the main file
to understand the current state...It has lost track of what it was doing.
Fix: Before starting a large task, tell Claude your plan upfront so it is in the recent turns (which survive compaction). If you are already deep in a task and see context climbing past 50%, compact proactively with explicit instructions:
/compact keep: 1) the API design we agreed on 2) the three files I need modified 3) the test planCompact Losing Critical Context
Symptom: After compacting, Claude no longer remembers a key architectural decision, a constraint you explained, or a tricky edge case you discussed.
Fix: Pre-summarize important decisions before compacting. Type a message like:
Before we compact, let me summarize our key decisions for reference:
1. We chose event-based middleware over chain-of-responsibility because...
2. The hook ordering is: auth -> rate-limit -> logging -> handler
3. Error handlers must NOT swallow exceptions, only transform them
Now: /compact keep the above summary and the current implementation planBy putting the summary in a recent turn, it survives compaction with high fidelity.
Even better: put persistent decisions in CLAUDE.md or a memory file. Those are always preserved in full, regardless of compaction.
Session Too Long -- When to Start Fresh vs. Compact
Rule of thumb:
| Situation | Action |
|---|---|
| Same task, context at 40-50% | /compact with guidance |
| Same task, already compacted twice | Start a fresh session. Write key context into CLAUDE.md first. |
| Switching to a different task | New session (do not compact, just leave) |
| Need to reference work from yesterday | --resume if you need full detail, new session with CLAUDE.md notes if you just need the conclusions |
The compaction-of-compaction problem: Each /compact pass is lossy. Compacting a summary that was already a compaction of earlier work loses progressively more detail. After two compactions in the same session, you are working from a summary of a summary. Start fresh instead.
Going Deeper
Context Management Strategies
| Scenario | Strategy |
|---|---|
| Quick task (<10 min) | Just do it, don't overthink context |
| Medium task (10-30 min) | Compact proactively around 40-50% |
| Long task (>30 min) | Split into sub-tasks. Fresh session per sub-task. |
| Massive codebase analysis | Use sub-agents to isolate context (Ch8) |
| Reading a large file | Tell Claude which lines/functions you care about |
From GSD's "Wave Execution" model: treat each task as a wave. Each wave gets a fresh context. Results from one wave feed into the next as concise input, not raw conversation history.
Reducing Token Consumption
# 1. Be specific about what to read
# Bad: "Read main.py" (might be 2000 lines)
# Good: "Read the handle_request function in main.py, around line 150"
# 2. Limit command output
# Bad: "Run npm install" (hundreds of lines of output)
# Good: "Run npm install and just tell me if it succeeded"
# 3. Search instead of read
# Bad: "Read all Python files and find database code"
# Good: "Grep for 'database' across all .py files"
# 4. Compact with instructions
/compact keep the API design and the test results
# Claude preserves what you tell it to prioritizeAuto-Compaction Details
When context hits ~83.5%, Claude auto-compacts. After compaction:
- CLAUDE.md: preserved in full
- Memory files: preserved in full
- Recent turns (last 2-3): preserved
- Earlier conversation: compressed to summary
- File contents that were read: lost (only the summary of what was found remains)
- Long Bash outputs: lost
This is why proactive compaction at 50% beats waiting for auto-compaction. At 50%, you're still in the golden zone and Claude can write a better summary. At 83.5%, it's already degraded.
Prompt Caching and Context Costs
Claude Code uses prompt caching to reduce the cost of repeated context. The system prompt, CLAUDE.md, and memory files are sent as a cached prefix -- subsequent turns reuse this cached prefix rather than re-processing it from scratch. This means:
- The fixed overhead (system prompt + CLAUDE.md + memory) is cheap after the first turn
- But conversation history grows linearly and is not cached
- Compaction resets the conversation portion, but the cached prefix stays warm
Deep dive: A10 Prompt Caching explains the caching mechanism, prefix matching rules, and how this affects your costs.
Knowledge Check
Exercise: Context Management Under Pressure
Task
- Start a session in the Flask repo (from Demo 6)
- Ask Claude to analyze three different subsystems, checking
/contextafter each:- The template rendering system (Jinja2 integration)
- The session/cookie handling
- The testing utilities (Flask's test client)
- When context hits 40-50%, run
/compact keep the analysis of all three subsystems - After compaction, ask Claude to compare the three subsystems' design patterns
- Exit and resume with
--resume, then ask a follow-up about one of the subsystems - Record your context percentages at each stage
Success Criteria
- [ ] Recorded
/contextoutput at 3+ different points - [ ] Compacted at the right moment (40-50%, not 80%+)
- [ ] Claude's post-compact output was coherent and useful
- [ ] Successfully resumed with
--resume - [ ] Can explain the difference between the 40-60% golden zone and the 80%+ danger zone
Reference Data
Typical context consumption (200K window):
- System prompt + CLAUDE.md: ~3-5%
- Reading one 200-line file: ~1-2%
- One conversation turn (prompt + response): ~0.5-1%
- Generating a 100-line code file: ~1.5-2%
- After compaction: back to ~8-12%
Chapter Summary
- The context window is Claude's working memory. Everything competes for space inside it.
- Context limits exist because of quadratic scaling in the attention mechanism -- not arbitrary product restrictions.
- Tokens are BPE subword units (~3-4 chars each). Code is more token-expensive than prose.
- Context Rot is real: quality degrades past 50% usage. The 40-60% range is the golden zone (community-observed, widely confirmed).
- The "lost in the middle" effect means Claude naturally attends less to information buried in mid-conversation. Put critical context in CLAUDE.md or recent turns.
- Monitor with
/context. Subscribers check/usage, API users check/cost. - Compact proactively at 40-50%, don't wait for auto-compaction at 83.5%.
/compactpreserves CLAUDE.md and memory but drops earlier details. Give it hints about what to keep.--resumerestores full session context for work that spans breaks.- When things go wrong: pre-summarize before compacting, avoid double-compaction, and start fresh when a session is too far gone.
- GSD's principle: "Long-running sessions are the enemy." Break work into focused waves.
Further reading: A03 LLM Fundamentals (attention mechanism, why context has limits) | A07 Token Economics (BPE tokenization, cost estimation) | A10 Prompt Caching (how Claude Code reduces cost with cached prefixes)
Next up: Chapter 4: Git Workflows & PR Automation. You've been running git commands yourself this whole time. That's about to change. Claude can handle your entire branch-commit-PR workflow, and with the checkpoint system, you can let it try risky refactors without fear. The /rewind command is going to become your best friend.