Skip to content

A10: Prompt Caching ​

Related chapters: Ch3 Context Window Management, Ch13 Agent SDK, A07 Token Economics

Prompt caching is an optimization that lets the API skip reprocessing tokens it has already seen. When the beginning of your prompt matches a recently cached prompt, those tokens are served from cache at 90% lower cost and significantly reduced latency. For Claude Code, where the system prompt and CLAUDE.md content are identical across every turn, caching delivers automatic and substantial savings.

How Prompt Caching Works ​

Prefix Matching ​

Prompt caching uses exact prefix matching. The API compares the beginning of your current prompt with recently cached prompts. If a contiguous prefix matches exactly (byte-for-byte), those tokens are served from cache.

Turn 1 (cache miss — nothing cached yet):
┌─────────────────────────────────────────────────┐
│ System prompt       [3,000 tokens]  PROCESSED   │
│ CLAUDE.md           [2,000 tokens]  PROCESSED   │
│ Tool definitions    [1,500 tokens]  PROCESSED   │
│ User message        [  200 tokens]  PROCESSED   │
└─────────────────────────────────────────────────┘
  Total: 6,700 tokens processed at full price
  Cache: prefix of 6,500 tokens stored

Turn 2 (cache hit on prefix):
┌─────────────────────────────────────────────────┐
│ System prompt       [3,000 tokens]  CACHED ✓    │
│ CLAUDE.md           [2,000 tokens]  CACHED ✓    │
│ Tool definitions    [1,500 tokens]  CACHED ✓    │
│ Turn 1 history      [  600 tokens]  PROCESSED   │
│ User message        [  150 tokens]  PROCESSED   │
└─────────────────────────────────────────────────┘
  Cached: 6,500 tokens at 90% discount
  Processed: 750 tokens at full price

The critical insight: caching works on the prefix only. If any byte in the middle of the prefix differs, the cache breaks at that point and everything after it is processed at full price.

What Gets Cached ​

The cache stores the intermediate computational state (the key-value attention cache, internally) for the matched prefix. This means the model does not need to re-attend to those tokens --- it can jump directly to processing the new tokens.

Without caching:
  Prompt [A B C D E] → Process A, then B, then C, then D, then E

With caching (prefix A B C is cached):
  Prompt [A B C D E] → Load cached state for A B C, process only D and E

Minimum Prefix Length ​

Caching requires a minimum prefix length to be effective. Anthropic requires at least 1,024 tokens in the cached prefix for a cache entry to be created. Short prompts below this threshold will not benefit from caching.

In practice, Claude Code sessions always exceed this threshold because the system prompt alone is approximately 3,000 tokens.

Cache Hit Requirements ​

For a cache hit to occur, the following must match exactly:

1. Same Model ​

Cache entries are per-model. A prompt cached for claude-sonnet-4-20250514 will not produce a hit for claude-opus-4-20250514.

2. Exact Prefix Bytes ​

Every byte in the prefix must match. This includes:

  • System prompt content
  • CLAUDE.md content (any edit breaks the cache)
  • Tool definitions (if tools change, cache breaks)
  • Message ordering and content
  • Whitespace, newlines, and formatting
# These are DIFFERENT prefixes (no cache hit):

"System: You are a helpful assistant."    # period
"System: You are a helpful assistant"     # no period

"Use pnpm for packages"                  # original
"Use pnpm  for packages"                 # extra space (!)

3. Same API Parameters (for certain fields) ​

Some API parameters are part of the cache key. If you change model, the cache breaks. Parameters like max_tokens and temperature do not affect the cache key.

What Breaks the Cache ​

Common scenarios that invalidate the cache prefix:

Cache-breaking changes:
├── Editing CLAUDE.md (any character change)
├── Adding/removing MCP servers (changes tool definitions)
├── Switching models mid-session
├── Different system prompt versions (Claude Code updates)
└── Modifying the message history (editing a past message)

Non-breaking changes (cache preserved):
├── New user messages (appended to the end)
├── New assistant messages (appended to the end)
├── Tool results (appended to the end)
├── Changing max_tokens
└── Changing temperature

Impact on Latency and Cost ​

Time to First Token (TTFT) ​

Caching dramatically reduces the time to first token because the model skips processing the cached prefix. The improvement scales with prefix length:

Prefix size    │  TTFT (no cache)  │  TTFT (cached)  │  Speedup
──────────────┼──────────────────┼────────────────┼──────────
5,000 tokens   │  ~1.2 seconds    │  ~0.3 seconds  │  4x
20,000 tokens  │  ~3.5 seconds    │  ~0.5 seconds  │  7x
50,000 tokens  │  ~7.0 seconds    │  ~0.8 seconds  │  9x
100,000 tokens │  ~12.0 seconds   │  ~1.0 seconds  │  12x

For long Claude Code sessions where the conversation history grows to 50K+ tokens, caching makes the difference between a snappy tool and one that feels sluggish.

Cost Reduction ​

Cached input tokens are priced at 10% of the standard input rate:

ModelStandard Input (per 1M)Cached Input (per 1M)Savings
Claude Opus 4$15.00$1.5090%
Claude Sonnet 4$3.00$0.3090%
Claude Haiku 3.5$0.80$0.0890%

Note: there is also a small cache write cost when a new cache entry is created (25% surcharge on the first request). This is amortized over subsequent cache hits.

Example: 20-turn session with 5,000-token stable prefix (Sonnet)

Without caching:
  20 turns × 5,000 prefix tokens × $3.00/1M = $0.30

With caching:
  Turn 1 (cache write): 5,000 × $3.75/1M = $0.01875
  Turns 2-20 (cache hit): 19 × 5,000 × $0.30/1M = $0.0285
  Total: $0.047  (84% savings on the prefix)

The 5-Minute TTL ​

How It Works ​

Cache entries have a 5-minute time-to-live (TTL). If 5 minutes pass without a matching request, the cache entry expires and the next request will be a cache miss (full processing + cache write).

Timeline:
  0:00  Request 1 → Cache miss (cache created)
  0:30  Request 2 → Cache hit ✓ (TTL resets to 5 min)
  1:00  Request 3 → Cache hit ✓ (TTL resets)
  
  ... 6 minutes of inactivity ...
  
  7:00  Request 4 → Cache miss (expired, new cache created)
  7:15  Request 5 → Cache hit ✓ (TTL resets)

Each cache hit resets the TTL. As long as you keep making requests within 5 minutes of each other, the cache stays warm.

Implications for Claude Code Usage ​

  • Active sessions benefit most: If you are actively coding (sending prompts every few minutes), caching stays warm for the entire session.
  • Breaks longer than 5 minutes cause a cache miss on the next request. The first request after a break is slower and more expensive, but subsequent requests benefit again.
  • Idle sessions: If you step away for lunch, expect the first request back to be a cache miss.
  • CI/CD pipelines: If pipeline runs are spaced more than 5 minutes apart, each run starts cold. For frequent pipelines (e.g., running on every commit), caching helps significantly.

Strategies to Keep the Cache Warm ​

For scenarios where cache warmth matters (e.g., shared API-based tools):

python
# Ping approach: send a minimal request to refresh the TTL
# (Only useful for API integrations, not for Claude Code CLI)
import time
import threading

def keep_cache_warm(client, system_prompt, interval=240):
    """Send a minimal request every 4 minutes to keep cache alive."""
    def ping():
        while True:
            client.messages.create(
                model="claude-sonnet-4-20250514",
                max_tokens=1,
                system=system_prompt,
                messages=[{"role": "user", "content": "ping"}]
            )
            time.sleep(interval)
    thread = threading.Thread(target=ping, daemon=True)
    thread.start()

This is rarely necessary for interactive Claude Code usage but can be valuable for API-based integrations where the system prompt is large and expensive to reprocess.

How to Structure Prompts for Maximum Cache Hits ​

Principle: Stable Content First, Variable Content Last ​

Since caching works on prefixes, put the most stable content at the beginning of your context:

Optimal ordering (maximizes cache prefix):
┌────────────────────────────────────────┐
│ 1. System prompt        (never changes) │  ← Always cached
│ 2. CLAUDE.md            (rarely changes)│  ← Almost always cached
│ 3. Tool definitions     (rarely changes)│  ← Almost always cached
│ 4. Conversation history (grows linearly)│  ← Partially cached
│ 5. Current user message (always new)    │  ← Never cached
└────────────────────────────────────────┘

Poor ordering (breaks cache early):
┌────────────────────────────────────────┐
│ 1. Current timestamp    (always changes)│  ← Breaks cache!
│ 2. System prompt                        │  ← Not cached
│ 3. Everything else                      │  ← Not cached
└────────────────────────────────────────┘

Claude Code already structures its prompts this way. You benefit automatically.

CLAUDE.md Stability Matters ​

Since CLAUDE.md is part of the cached prefix, editing it mid-session breaks the cache:

Session timeline:
  Turn 1-5:  Cache building, prefix grows, cost decreasing
  Turn 6:    User edits CLAUDE.md (adds one line)
  Turn 7:    Cache MISS on the entire prefix (CLAUDE.md changed)
  Turn 8+:   New cache builds from the updated prefix

Practical advice: Make CLAUDE.md edits between sessions, not during them. If you must edit mid-session, do it early to minimize cache waste.

Conversation History and Caching ​

As your conversation grows, the cached prefix grows with it:

Turn 1:  Prefix = system + CLAUDE.md + tools (5,000 tk) → miss
Turn 2:  Prefix = above + turn 1 (6,000 tk) → 5,000 tk cached
Turn 5:  Prefix = above + turns 2-4 (12,000 tk) → 11,000 tk cached
Turn 10: Prefix = above + turns 5-9 (25,000 tk) → 24,000 tk cached
Turn 20: Prefix = above + turns 10-19 (50,000 tk) → 49,000 tk cached

Longer sessions mean more cache hits because the growing conversation history is always a prefix of the next request.

However, when Claude Code performs compaction (summarizing old messages to fit the context window), the compacted content differs from the original, breaking the cache for the compacted portion. This is an unavoidable trade-off between context window management and cache efficiency.

CLAUDE.md and Prompt Caching Synergy ​

CLAUDE.md is uniquely well-suited for prompt caching because:

  1. It is injected early in the prompt (part of the system prompt prefix)
  2. It rarely changes during a session
  3. It is identical across all turns in a session
  4. It applies to every request (no conditional inclusion)

This means a well-crafted CLAUDE.md is effectively "free" after the first turn. A 3,000-token CLAUDE.md costs:

Turn 1 (cache write):  3,000 tokens × $3.75/1M = $0.011  (Sonnet)
Turn 2+ (cache read):  3,000 tokens × $0.30/1M = $0.0009 per turn

Over a 20-turn session:
  Without caching: 20 × 3,000 × $3.00/1M = $0.18
  With caching:    $0.011 + 19 × $0.0009  = $0.028
  Savings: 84%

Do not skimp on CLAUDE.md content to "save tokens." The caching system means that CLAUDE.md content is nearly free after the first turn. Invest in a thorough CLAUDE.md --- it pays for itself in better Claude behavior with minimal cost impact.

Practical Measurement: Detecting Cache Hits ​

API Response Headers ​

When using the Anthropic API directly, cache information is returned in the response:

json
{
  "usage": {
    "input_tokens": 2500,
    "output_tokens": 800,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 6500
  }
}

The key fields:

  • cache_creation_input_tokens: Tokens written to cache (first request or after cache miss). Billed at 1.25x standard input rate.
  • cache_read_input_tokens: Tokens served from cache. Billed at 0.1x standard input rate.
  • input_tokens: Tokens processed normally (not cached).

Calculating Your Cache Hit Rate ​

python
# From API response usage data
def cache_hit_rate(usage):
    total_input = (
        usage["input_tokens"] + 
        usage["cache_creation_input_tokens"] + 
        usage["cache_read_input_tokens"]
    )
    if total_input == 0:
        return 0
    return usage["cache_read_input_tokens"] / total_input

# Example
usage = {
    "input_tokens": 2500,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 6500
}
print(f"Cache hit rate: {cache_hit_rate(usage):.1%}")  # 72.2%

What Good Cache Rates Look Like ​

Scenario                          │  Expected Cache Hit Rate
──────────────────────────────────┼─────────────────────────
Active multi-turn session         │  60-85%
Long session (20+ turns)          │  75-90%
Session after 5+ min break        │  0% (first turn), then 60%+
CI/CD with frequent runs (<5 min) │  50-70%
CI/CD with infrequent runs        │  0-10%
First turn of any session         │  0% (always a miss)

Claude Code CLI Observation ​

In Claude Code, you do not see cache metrics directly in the UI. However, you can observe caching effects indirectly:

  • First turn of a session is noticeably slower than subsequent turns
  • After editing CLAUDE.md, the next response is slower (cache rebuilt)
  • After a long break, the first response is slower (cache expired)
  • Token usage reported in the status bar reflects effective (post-caching) costs

Interaction with Extended Thinking ​

Extended thinking (see A09) generates output tokens, which are never cached. Only input tokens benefit from prompt caching. This means:

With thinking enabled:
  Input tokens  → Can be cached (90% savings possible)
  Thinking tokens → Output, never cached, always full price
  Response tokens → Output, never cached, always full price

Caching and thinking are complementary optimizations that address different cost components:

  • Caching reduces the cost of processing context (input tokens)
  • Thinking budget management reduces the cost of reasoning (output tokens)

Reference: Anthropic's Documentation ​

For the latest details on prompt caching, refer to:

  • Anthropic Docs: Prompt Caching --- official documentation with current pricing and API details
  • Anthropic Cookbook: Prompt caching examples with Python and TypeScript SDKs
  • API Reference: The usage object in the Messages API response

Caching behavior and pricing may change. Always check the official documentation for the latest information.

Key Takeaways ​

  1. Prompt caching reduces input costs by 90% for tokens that match a previously cached prefix. Claude Code benefits automatically because the system prompt and CLAUDE.md form a stable prefix.
  2. Cache entries expire after 5 minutes of inactivity. Active sessions stay warm; breaks longer than 5 minutes cause a cold start on the next request.
  3. Stable prefixes maximize caching. Put unchanging content (system prompt, CLAUDE.md) at the beginning. Avoid editing CLAUDE.md mid-session.
  4. A thorough CLAUDE.md is nearly free after the first turn. Do not skimp on project instructions to save tokens --- caching makes the per-turn cost negligible.
  5. Cache hit rate of 60-85% is typical for active multi-turn sessions. First turns are always cache misses.
  6. Caching and extended thinking are complementary: caching reduces input costs, thinking budget management reduces output costs.

See also: A07 Token Economics for full cost analysis, Ch3 Context Window Management for context strategies, and A09 Extended Thinking for output token optimization.

Released under MIT License