Skip to content

Chapter 3: Context Window Management ​

Learning Objectives ​

  • Understand why Claude gets worse over long sessions (it's not a bug, it's context)
  • Monitor consumption with /context, /usage, and /cost
  • Know exactly when to /compact and what you lose
  • Use --resume and session management to pick up where you left off
  • Internalize the 40-60% golden zone for optimal output quality

Appendix links: A03 LLM Fundamentals explains the Transformer architecture and attention mechanism that underlie everything in this chapter. A07 Token Economics covers BPE tokenization, cost models, and budget optimization. A10 Prompt Caching details how Claude Code reuses cached prefixes to reduce cost and latency.

Concepts ​

The Context Window is Claude's Working Memory ​

Everything in your conversation -- your prompts, Claude's responses, every file it reads, every command output -- lives in a single memory buffer called the context window. When it fills up, Claude starts forgetting things or producing lower quality output.

┌──────────────────────────────────────────────────┐
│          Context Window (200K tokens)              │
│                                                    │
│  ┌──────────┐ ┌──────────┐ ┌───────────┐         │
│  │ System    │ │ CLAUDE.md│ │ Memory    │  Fixed   │
│  │ Prompt    │ │          │ │ Files     │  cost    │
│  └──────────┘ └──────────┘ └───────────┘         │
│                                                    │
│  ┌────────────────────────────────────────┐       │
│  │  Conversation History + Tool Results    │       │
│  │                                        │       │
│  │  You: "Analyze the routing module"     │       │
│  │  Claude: [Read 5 files] -> 2000 lines  │       │
│  │  Claude: "Here's how routing works..." │       │
│  │  You: "Now add tests"                  │       │
│  │  Claude: [Write test file]             │ Grows  │
│  │  Claude: [Bash] pytest -> output       │       │
│  │  ...keeps growing...                   │       │
│  └────────────────────────────────────────┘       │
│                                                    │
│  ████████████████████░░░░░░░  68% used             │
└──────────────────────────────────────────────────┘

Why Context Windows Have Limits ​

Context windows are not arbitrary product restrictions -- they are a fundamental consequence of how Transformer models work. The core of every LLM is the attention mechanism (Attention), which lets each token "look at" every other token in the sequence. This creates quadratic scaling: processing N tokens requires roughly N squared operations. Double the context from 100K to 200K tokens, and the computation cost roughly quadruples.

This is why even with 200K or 1M token windows, there are practical limits. The model needs increasingly more memory and compute as the context grows. Hardware constraints impose a ceiling.

Deep dive: A03 LLM Fundamentals covers the Transformer architecture and attention mechanism in detail, including recent advances like sliding window attention and sparse attention that help extend context lengths.

What Are Tokens? ​

Before we discuss context budgets, you need to understand what tokens actually are. Tokens are not words and not characters -- they are subword units produced by Byte Pair Encoding (BPE) tokenization. The tokenizer splits text into the most statistically frequent subword pieces from its training data.

A rough rule of thumb: 1 token is approximately 3-4 characters in English, or about 0.75 words. For code, the ratio varies:

"hello world"           → 2 tokens   (common words = 1 token each)
"implementation"        → 1 token    (common long word)
"XMLHttpRequest"        → 4 tokens   (camelCase splits)
"const x = arr.map()"  → 8 tokens   (code has more splits)
"你好世界"              → 4 tokens   (Chinese: ~1-2 tokens per character)

This matters for context management because code is token-expensive: a 200-line Python file might be 3000 tokens, not the 1000 you might guess from word count.

Deep dive: A07 Token Economics explains BPE tokenization in full, including how to estimate costs before running a session.

Context Rot: The Silent Killer ​

GSD (50k stars) coined a term that every Claude Code user should know: Context Rot.

Their research found that once context usage passes ~50%, Claude's output quality starts degrading. By 70%, it's noticeably worse. By 80%, you're getting outputs that miss requirements, introduce bugs, and forget earlier instructions.

Quality vs. Context Usage:

100% |████
     |████████
     |████████████
     |████████████████
     |████████████████████
     |██████████████████████████
     |████████████████████████████████
     |██████████████████████████████████████████
  0% |____________________________________________
     0%    20%    40%    60%    80%    100%
           Context usage -->

     The sweet spot is 40-60%. Past that, quality drops fast.

This is not theoretical. HumanLayer's research across thousands of sessions confirmed the 40-60% golden zone -- Claude performs best when the context window is less than 60% full. Past that, you're fighting degrading returns.

Important caveat: The 40-60% golden zone is a widely reported observation from the Claude Code community (GSD, HumanLayer, and many practitioners), not a peer-reviewed academic finding. Your mileage may vary depending on task complexity and content type. The key insight -- that quality degrades well before the window is full -- is consistent across reports.

The "Needle in a Haystack" Effect ​

Research on long-context models (the "needle in a haystack" tests) shows that model performance is not uniform across the context window. Models tend to attend most strongly to:

  1. The beginning of the context (system prompt, CLAUDE.md) -- highest attention
  2. The most recent turns -- strong recency bias
  3. The middle of the context -- weakest attention, details here are most likely to be "forgotten"

This is called the "lost in the middle" phenomenon. It means that a critical decision you made 30 turns ago, buried in the middle of the conversation, is the most likely thing Claude will lose track of -- even if the context window is not full yet.

Practical implication: if you have important constraints or decisions, put them in CLAUDE.md (always at the top of context) or reiterate them in your current prompt, rather than relying on Claude remembering something from turn 12 of a 40-turn session.

What Eats Tokens? ​

ContentApproximate CostImpact
1 line of code~10-15 tokensReads add up fast
1 file (200 lines)~2,000-3,000 tokensA few files = 5-10% of context
1 conversation turn~100-500 tokensDepends on prompt length
Bash command outputVaries wildlynpm install output can burn 5,000+ tokens
CLAUDE.md~200-500 tokensLoaded every turn
Memory files~100-300 tokensLoaded every turn

The biggest token sink is usually reading files. Every time Claude reads a file to answer your question, that entire file content goes into the context and stays there.

Context Sizes by Plan ​

PlanContext WindowGood For
Standard (Sonnet)200K tokensMost tasks, moderate sessions
Max/Team/Enterprise + Opus 4.61M tokensHuge codebases, extended analysis

What Happens at the Limit ​

At ~83.5% context usage, Claude Code triggers auto-compaction: it compresses the conversation history into a summary, preserving CLAUDE.md and memory files but losing details from earlier in the conversation.

Here is what auto-compaction looks like in your terminal:

terminal
$ [auto-compaction triggered]

  ⚠ Context usage reached 83.5% (167,000 / 200,000 tokens)
  Compacting conversation history...

  Preserved:
    ✓ CLAUDE.md (full)
    ✓ Memory files (full)
    ✓ Last 2-3 conversation turns (full)

  Compressed:
    ↓ 47 earlier turns → summary (2,100 tokens)
    ↓ 12 file reads → key findings only
    ↓ 8 bash outputs → result summaries

  Context after compaction: 24,800 / 200,000 tokens (12.4%)

  ℹ Tip: Use /compact proactively at 40-50% for better summaries.

The problem? By the time auto-compaction kicks in, you've already been in the degraded zone for a while. That's why proactive management matters.


Demo 6: Watch Context Fill Up on a Real Codebase ​

6
Watch Context Fill Up on a Real Codebase
Beginner~15 min

Goal ​

Clone Flask, have Claude analyze its routing system, and watch the context window fill up in real time. Then compact and see the difference.

Steps ​

1. Set Up ​

We'll use pallets/flask (68k+ stars). Flask's routing system is complex enough to burn through tokens quickly.

bash
mkdir -p ~/claude-demos/demo-06 && cd ~/claude-demos/demo-06
git clone --depth 1 https://github.com/pallets/flask.git
cd flask

cat > CLAUDE.md << 'EOF'
# Flask Context Demo
- Focus on the routing and request handling subsystems
- Keep explanations concise
EOF

2. Baseline Measurement ​

bash
claude
/context

Here is what the /context output looks like at the start of a session:

terminal
$ /context

  Context usage: 9,400 / 200,000 tokens (4.7%)

  Breakdown:
    System prompt:     4,200 tokens
    CLAUDE.md:           180 tokens
    Memory files:        120 tokens
    Conversation:      4,900 tokens

  ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░  4.7%

  Status: Healthy — plenty of room.

Check your quota/cost too:

/usage

(Or /cost if you're on an API key.)

3. Start Burning Context ​

Ask Claude to read and analyze Flask's routing:

Read the core routing files in src/flask/ -- I want to understand how Flask registers routes and dispatches requests. Start with app.py and find the route() decorator implementation.
/context

You should see something like this now:

terminal
$ /context

  Context usage: 36,200 / 200,000 tokens (18.1%)

  Breakdown:
    System prompt:     4,200 tokens
    CLAUDE.md:           180 tokens
    Memory files:        120 tokens
    Conversation:     31,700 tokens  ← jumped from file reads + analysis

  ███░░░░░░░░░░░░░░░░░░░░░░░░░░░  18.1%

  Status: Healthy — good headroom remaining.

Record the new percentage. Reading a few source files probably jumped you 5-15%.

4. Go Deeper ​

Now trace what happens when a request comes in. How does Flask match the URL to a route? Show me the code path from the WSGI entry point to the view function being called.
/context

You should be climbing. Claude had to read more files, produce longer explanations, and the context is filling up.

5. One More Push ​

Analyze Flask's blueprint system. How do blueprints register their routes? What happens when there's a URL conflict between a blueprint route and an app route?
/context

6. The Same Task at 30% vs. 70% ​

This is the critical experiment. If you're above 50%, now do this:

/compact keep the routing analysis and blueprint explanation

Here is what /compact output looks like:

terminal
$ /compact keep the routing analysis and blueprint explanation

  Compacting with guidance: "keep the routing analysis and blueprint explanation"

  Before: 104,600 / 200,000 tokens (52.3%)
  After:   22,400 / 200,000 tokens (11.2%)

  Preserved (per your instructions):
    ✓ Routing analysis (route() decorator, URL rule registration)
    ✓ Blueprint explanation (blueprint routes, conflict resolution)
    ✓ CLAUDE.md (always preserved)
    ✓ Memory files (always preserved)

  Dropped:
    ✗ Raw file contents from app.py, blueprints.py, scaffold.py
    ✗ Intermediate reasoning and exploratory reads
    ✗ Bash command outputs

  ℹ Context freed: 82,200 tokens (41.1%)
/context

Context should drop significantly. Now ask:

Explain Flask's error handling system. How does @app.errorhandler work, and where in the request lifecycle do error handlers execute?

Compare the quality of this answer to the responses you got when the context was fuller. You should notice the post-compact response is more focused, more structured, and catches more nuance.

What Just Happened? ​

Here is the tool-call sequence Claude used during the routing analysis, and why it chose each tool:

1
Glob
src/flask/*.py
↓
2
Read
src/flask/app.py
↓
3
Grep
add_url_rule across src/flask/
↓
4
Read
src/flask/scaffold.py
↓
5
Read
src/flask/blueprints.py

Context consumption timeline:

Start:    [System + CLAUDE.md + Memory]           = ~5%
Read routing files:                                = ~18%
Trace request handling:                            = ~35%
Analyze blueprints:                                = ~52%
/compact:                                          = ~12%
Analyze error handling (fresh context):            = ~20%

The error handling analysis at 20% was almost certainly better than if you'd asked the same question at 65%. That's Context Rot in action.

Key lesson from GSD: "Long-running sessions are the enemy." Split your work into focused tasks, compact proactively, or start fresh sessions.


Demo 7: Session Management & Resuming ​

7
Session Management & Resuming
Beginner~10 min

Goal ​

Learn to exit and resume sessions without losing context. This is essential for work that spans hours or days.

Steps ​

1. Create a Session Worth Resuming ​

bash
cd ~/claude-demos/demo-06/flask
claude
Help me design a middleware system inspired by Flask's approach. I want:
1. A middleware registration API
2. Before-request and after-request hooks
3. Error-handling middleware

Just produce the design doc, no code yet.

After Claude delivers the design, note a few key details (module names, hook ordering), then exit:

Ctrl+C

2. Resume Where You Left Off ​

bash
claude --resume

Verify it worked:

What was the hook ordering we decided on for the middleware system?

Claude should recall everything from the previous session. --resume loads the full session log, not a compressed summary.

3. Browse Session History ​

Start a new session to see your history:

bash
claude
/sessions

This lists recent sessions with timestamps and summaries. You can resume any of them.

4. Resume a Specific Session ​

bash
# Each session has an ID shown in /sessions
claude --resume <session-id>

What Just Happened? ​

Here is the tool-call sequence Claude used when designing the middleware system:

1
Read
src/flask/app.py
↓
2
Grep
before_request across src/flask/
↓
3
Read
src/flask/wrappers.py

Key insight: Even though Claude is "just" producing a design document and not writing code, it still uses tools to ground its design in real implementation patterns. This is why tool calls consume context -- and why the design task itself uses significant tokens.

Key Difference ​

ActionWhat LoadsDetail Level
New sessionCLAUDE.md + Memory onlyFresh start, no prior conversation
--resumeFull session logEverything, as if you never left
After /compactCompressed summary + CLAUDE.md + MemoryKey points preserved, details lost

When to use each:

  • New session: Switching tasks, starting something unrelated
  • --resume: Picking up after a break, interrupted by a meeting
  • /compact: Staying in the same session but running low on context

When Things Go Wrong ​

Context management is not always smooth. Here are the most common failure modes and how to handle them.

Context Overflow Mid-Task ​

Symptom: Claude is halfway through implementing a feature. Auto-compaction fires. Claude's next response misses requirements it was previously tracking, or it re-reads files it already analyzed.

terminal
  ⚠ Context usage reached 83.5%
  Compacting conversation history...
  ...
  Context after compaction: 24,800 / 200,000 tokens (12.4%)

Then Claude says something like:

I'll continue implementing the feature. Let me re-read the main file
to understand the current state...

It has lost track of what it was doing.

Fix: Before starting a large task, tell Claude your plan upfront so it is in the recent turns (which survive compaction). If you are already deep in a task and see context climbing past 50%, compact proactively with explicit instructions:

/compact keep: 1) the API design we agreed on 2) the three files I need modified 3) the test plan

Compact Losing Critical Context ​

Symptom: After compacting, Claude no longer remembers a key architectural decision, a constraint you explained, or a tricky edge case you discussed.

Fix: Pre-summarize important decisions before compacting. Type a message like:

Before we compact, let me summarize our key decisions for reference:
1. We chose event-based middleware over chain-of-responsibility because...
2. The hook ordering is: auth -> rate-limit -> logging -> handler
3. Error handlers must NOT swallow exceptions, only transform them

Now: /compact keep the above summary and the current implementation plan

By putting the summary in a recent turn, it survives compaction with high fidelity.

Even better: put persistent decisions in CLAUDE.md or a memory file. Those are always preserved in full, regardless of compaction.

Session Too Long -- When to Start Fresh vs. Compact ​

Rule of thumb:

SituationAction
Same task, context at 40-50%/compact with guidance
Same task, already compacted twiceStart a fresh session. Write key context into CLAUDE.md first.
Switching to a different taskNew session (do not compact, just leave)
Need to reference work from yesterday--resume if you need full detail, new session with CLAUDE.md notes if you just need the conclusions

The compaction-of-compaction problem: Each /compact pass is lossy. Compacting a summary that was already a compaction of earlier work loses progressively more detail. After two compactions in the same session, you are working from a summary of a summary. Start fresh instead.


Going Deeper ​

Context Management Strategies ​

ScenarioStrategy
Quick task (<10 min)Just do it, don't overthink context
Medium task (10-30 min)Compact proactively around 40-50%
Long task (>30 min)Split into sub-tasks. Fresh session per sub-task.
Massive codebase analysisUse sub-agents to isolate context (Ch8)
Reading a large fileTell Claude which lines/functions you care about

From GSD's "Wave Execution" model: treat each task as a wave. Each wave gets a fresh context. Results from one wave feed into the next as concise input, not raw conversation history.

Reducing Token Consumption ​

bash
# 1. Be specific about what to read
# Bad: "Read main.py" (might be 2000 lines)
# Good: "Read the handle_request function in main.py, around line 150"

# 2. Limit command output
# Bad: "Run npm install" (hundreds of lines of output)
# Good: "Run npm install and just tell me if it succeeded"

# 3. Search instead of read
# Bad: "Read all Python files and find database code"
# Good: "Grep for 'database' across all .py files"

# 4. Compact with instructions
/compact keep the API design and the test results
# Claude preserves what you tell it to prioritize

Auto-Compaction Details ​

When context hits ~83.5%, Claude auto-compacts. After compaction:

  • CLAUDE.md: preserved in full
  • Memory files: preserved in full
  • Recent turns (last 2-3): preserved
  • Earlier conversation: compressed to summary
  • File contents that were read: lost (only the summary of what was found remains)
  • Long Bash outputs: lost

This is why proactive compaction at 50% beats waiting for auto-compaction. At 50%, you're still in the golden zone and Claude can write a better summary. At 83.5%, it's already degraded.

Prompt Caching and Context Costs ​

Claude Code uses prompt caching to reduce the cost of repeated context. The system prompt, CLAUDE.md, and memory files are sent as a cached prefix -- subsequent turns reuse this cached prefix rather than re-processing it from scratch. This means:

  • The fixed overhead (system prompt + CLAUDE.md + memory) is cheap after the first turn
  • But conversation history grows linearly and is not cached
  • Compaction resets the conversation portion, but the cached prefix stays warm

Deep dive: A10 Prompt Caching explains the caching mechanism, prefix matching rules, and how this affects your costs.


Knowledge Check ​

Your /context shows 55% usage. You still have 3 files to analyze for your current task. What should you do?
Keep going -- 55% is fine, auto-compaction will handle it at 83.5%
Compact now with guidance about what to keep, then continue the analysis
Start a completely new session immediately
Switch to a model with a larger context window
After /compact, Claude no longer remembers a key design decision from earlier in the conversation. How should you prevent this in the future?
Never use /compact
Put critical decisions in CLAUDE.md or summarize them in a recent message before compacting
Use --resume instead of /compact
Copy-paste the decision into every prompt
Why does Claude tend to 'forget' information from the middle of a long conversation, even when the context window is not full?
Claude deliberately ignores older messages to save compute
The attention mechanism attends most strongly to the beginning and end of context, with weaker attention to the middle
Middle messages are automatically deleted after 10 turns
This only happens when using /compact
You compacted your session, then compacted again 20 minutes later. The context is now at 15%. Is this a good state to keep working from?
Yes -- 15% means maximum quality
No -- double-compacted context is a summary of a summary, with significant detail loss. Start a fresh session instead.
It depends on which model you are using
Yes, as long as you used /compact with keep instructions both times

Exercise: Context Management Under Pressure ​

Task ​

  1. Start a session in the Flask repo (from Demo 6)
  2. Ask Claude to analyze three different subsystems, checking /context after each:
    • The template rendering system (Jinja2 integration)
    • The session/cookie handling
    • The testing utilities (Flask's test client)
  3. When context hits 40-50%, run /compact keep the analysis of all three subsystems
  4. After compaction, ask Claude to compare the three subsystems' design patterns
  5. Exit and resume with --resume, then ask a follow-up about one of the subsystems
  6. Record your context percentages at each stage

Success Criteria ​

  • [ ] Recorded /context output at 3+ different points
  • [ ] Compacted at the right moment (40-50%, not 80%+)
  • [ ] Claude's post-compact output was coherent and useful
  • [ ] Successfully resumed with --resume
  • [ ] Can explain the difference between the 40-60% golden zone and the 80%+ danger zone
Reference Data

Typical context consumption (200K window):

  • System prompt + CLAUDE.md: ~3-5%
  • Reading one 200-line file: ~1-2%
  • One conversation turn (prompt + response): ~0.5-1%
  • Generating a 100-line code file: ~1.5-2%
  • After compaction: back to ~8-12%

Chapter Summary ​

  • The context window is Claude's working memory. Everything competes for space inside it.
  • Context limits exist because of quadratic scaling in the attention mechanism -- not arbitrary product restrictions.
  • Tokens are BPE subword units (~3-4 chars each). Code is more token-expensive than prose.
  • Context Rot is real: quality degrades past 50% usage. The 40-60% range is the golden zone (community-observed, widely confirmed).
  • The "lost in the middle" effect means Claude naturally attends less to information buried in mid-conversation. Put critical context in CLAUDE.md or recent turns.
  • Monitor with /context. Subscribers check /usage, API users check /cost.
  • Compact proactively at 40-50%, don't wait for auto-compaction at 83.5%.
  • /compact preserves CLAUDE.md and memory but drops earlier details. Give it hints about what to keep.
  • --resume restores full session context for work that spans breaks.
  • When things go wrong: pre-summarize before compacting, avoid double-compaction, and start fresh when a session is too far gone.
  • GSD's principle: "Long-running sessions are the enemy." Break work into focused waves.

Further reading: A03 LLM Fundamentals (attention mechanism, why context has limits) | A07 Token Economics (BPE tokenization, cost estimation) | A10 Prompt Caching (how Claude Code reduces cost with cached prefixes)

Next up: Chapter 4: Git Workflows & PR Automation. You've been running git commands yourself this whole time. That's about to change. Claude can handle your entire branch-commit-PR workflow, and with the checkpoint system, you can let it try risky refactors without fear. The /rewind command is going to become your best friend.

Released under MIT License