A07: Token Economics
Related chapters: Ch3 Context Window Management, Ch15 Production Workflow Design, A10 Prompt Caching
Every interaction with Claude Code consumes tokens, and tokens cost money. Understanding token economics lets you get maximum value from your usage --- whether you are an individual developer on a subscription or a team lead managing API costs across an engineering organization.
How Tokens Work
BPE Tokenization
Claude uses Byte Pair Encoding (BPE) to convert text into tokens. BPE works by learning the most frequent character sequences in training data and assigning them single token IDs. The result is that common words become one token, while rare or technical terms get split into multiple tokens.
Common patterns:
"function" → 1 token (common English/programming word)
"the" → 1 token (most common English word)
"authentication" → 2 tokens ("authentic" + "ation")
"XMLHttpRequest" → 4 tokens (camelCase splits + unusual combo)
"дефрагментация" → 6 tokens (non-Latin text uses more tokens)Claude-specific examples for code:
"return null;" → 3 tokens (return + null + ;)
"console.log(" → 3 tokens (console + .log + ()
"if (err !== null)" → 6 tokens
"import { useState } from 'react';" → 9 tokensRules of Thumb
- 1 token is roughly 3-4 characters in English prose
- 1 token is roughly 2-3 characters in code (more punctuation and structure)
- 100 tokens is roughly 75 English words
- A typical 200-line source file is 1,500-3,000 tokens
- A 1,000-line file can be 8,000-15,000 tokens depending on language
What Counts as Tokens
In a Claude Code session, tokens are consumed by every part of the context:
Token budget breakdown for one turn:
┌──────────────────────────────────────────────────────┐
│ System prompt (Claude Code internals) ~3,000 tk │
│ CLAUDE.md content ~500-5,000 │
│ Tool definitions ~2,000 tk │
│ Conversation history (all prior turns) variable │
│ Current user prompt ~100-500 │
│ Tool results (file contents, etc.) variable │
│ ─────────────────────────────────────────────────── │
│ = Total input tokens (you pay for all of this) │
│ │
│ Model response + tool calls │
│ = Total output tokens (priced differently) │
└──────────────────────────────────────────────────────┘Cost per Token Across Models
Anthropic offers three model tiers. Prices shown are as of early 2025 (check Anthropic's pricing page for current rates):
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Cache Read (per 1M) | Best For |
|---|---|---|---|---|
| Claude Opus 4 | $15.00 | $75.00 | $1.50 | Complex reasoning, architecture |
| Claude Sonnet 4 | $3.00 | $15.00 | $0.30 | Daily development, balanced |
| Claude Haiku 3.5 | $0.80 | $4.00 | $0.08 | High-throughput, simple tasks |
Input vs Output Token Pricing
Output tokens are 5x more expensive than input tokens across all models. This is because generating tokens requires more computation than processing them.
This asymmetry has practical implications:
Scenario A: Claude reads 5,000 tokens, writes 200 tokens
Input cost: 5,000 / 1M × $3.00 = $0.015
Output cost: 200 / 1M × $15.00 = $0.003
Total: $0.018
Scenario B: Claude reads 1,000 tokens, writes 2,000 tokens
Input cost: 1,000 / 1M × $3.00 = $0.003
Output cost: 2,000 / 1M × $15.00 = $0.030
Total: $0.033Scenario B is almost twice as expensive despite using fewer total tokens, because output tokens dominate the cost. This means that asking Claude to generate verbose output (detailed explanations, long code blocks) is the single biggest cost driver.
Prompt Caching Impact on Cost
When Claude Code sends a request where the prefix matches a recently cached prompt, the cached portion is charged at 90% less than the standard input rate (for Sonnet: $0.30/1M instead of $3.00/1M).
In a typical Claude Code session, the system prompt and CLAUDE.md content remain identical across turns. This prefix gets cached automatically:
Turn 1 (cache miss):
System prompt + CLAUDE.md: 5,000 tokens × $3.00/1M = $0.015
New content: 2,000 tokens × $3.00/1M = $0.006
Total input cost: $0.021
Turn 2 (cache hit on prefix):
Cached prefix: 5,000 tokens × $0.30/1M = $0.0015
New content: 3,000 tokens × $3.00/1M = $0.009
Total input cost: $0.0105 (50% savings)
Turn 10 (most of conversation is cached):
Cached prefix: 30,000 tokens × $0.30/1M = $0.009
New content: 2,000 tokens × $3.00/1M = $0.006
Total input cost: $0.015 (vs $0.096 without caching = 84% savings)See A10 Prompt Caching for the full mechanics.
Batch API Pricing Advantages
For non-interactive use cases (CI/CD pipelines, bulk code review, documentation generation), the Batch API offers 50% discount on both input and output tokens. The trade-off is that results are not streamed --- you submit a batch and receive results within 24 hours.
| Scenario | Standard API | Batch API | Savings |
|---|---|---|---|
| Review 50 PRs (Sonnet) | ~$7.50 | ~$3.75 | 50% |
| Generate docs for 200 files (Haiku) | ~$3.20 | ~$1.60 | 50% |
| Architecture analysis (Opus) | ~$22.50 | ~$11.25 | 50% |
Batch API is ideal for CI/CD integration (see Ch12) where you do not need real-time responses.
Cost Optimization Strategies
Strategy 1: Right-Size Your Model
Not every task needs the most powerful model. A practical rule:
┌─────────────────────────────────────────────────────┐
│ Task Complexity │ Recommended Model │
│──────────────────────────│──────────────────────────│
│ Formatting, renaming, │ Haiku (cheapest) │
│ simple edits │ │
│──────────────────────────│──────────────────────────│
│ Feature implementation, │ Sonnet (balanced) │
│ debugging, refactoring │ │
│──────────────────────────│──────────────────────────│
│ Architecture decisions, │ Opus (most capable) │
│ complex multi-file │ │
│ reasoning │ │
└─────────────────────────────────────────────────────┘In Claude Code, switch models with /model or by pressing the model indicator in the status bar.
Strategy 2: Reduce Input Token Waste
The biggest source of token waste is reading unnecessary files. Techniques to reduce it:
- Point Claude to the right files: "Fix the bug in
src/auth/jwt.ts" is cheaper than "Fix the auth bug somewhere in the codebase." - Use CLAUDE.md to describe project structure: Claude reads CLAUDE.md once (cached after the first turn) and navigates directly to the right files instead of exploring.
- Use
/clearbetween unrelated tasks: Each turn in a conversation carries the full history. A 50-turn session means turn 50 sends all 49 prior turns as input. - Keep CLAUDE.md concise: Every token in CLAUDE.md is sent with every request. A 5,000-token CLAUDE.md costs ~$0.015 per turn (Sonnet, uncached) or ~$0.0015 per turn (cached).
Strategy 3: Control Output Verbosity
Since output tokens cost 5x more than input:
- Add "Be concise" or "No explanations needed" when you want just the code change.
- Use "Reply with only the file path" for search tasks.
- In CLAUDE.md, add rules like: "When making code changes, show only the changed lines, not the full file."
Strategy 4: Prompt Caching by Design
Structure your workflow to maximize cache hits:
- Keep CLAUDE.md content stable (do not edit it frequently during a session)
- Put the most stable context at the beginning of your prompts
- Use long sessions for related tasks (the growing conversation history stays cached)
Strategy 5: Use Subagents for Parallel Work
When Claude spawns subagents (see Ch8), each subagent has its own context. This means:
- The main agent's context stays small
- Subagents run in parallel, reducing wall-clock time
- Each subagent reads only the files it needs
- Total token usage may be similar or slightly higher, but time-to-completion improves
Strategy 6: CI/CD Optimization
For automated pipelines:
- Use Haiku for lint checks and simple validations
- Use Sonnet for PR reviews
- Use the Batch API for non-urgent tasks (overnight documentation, bulk analysis)
- Set
--max-turnsto limit runaway sessions - Monitor per-run costs and set budget alerts
Monthly Cost Estimates
These estimates assume Sonnet 4 with prompt caching active.
Individual Developer
Usage pattern: 40 interactive sessions/week, ~15 turns each
Input tokens: ~2M tokens/day (with caching: effective cost ~0.6M)
Output tokens: ~400K tokens/day
Daily cost: ~$8-12
Monthly cost: ~$180-270
With Max plan ($200/month subscription): Included in planSmall Team (5 developers)
Usage pattern: Each developer uses Claude Code for ~4 hours/day
Combined daily tokens: ~10M input, ~2M output
Daily cost: ~$45-60
Monthly cost: ~$1,000-1,300
Cost per developer: ~$200-260/monthCI/CD Pipeline (API usage)
Usage pattern: 50 PRs/week, each reviewed by Claude
Per PR: ~50K input tokens, ~5K output tokens
Weekly: 2.5M input + 250K output
Monthly (Sonnet, standard API): ~$50-70
Monthly (Sonnet, Batch API): ~$25-35
Monthly (Haiku, Batch API): ~$7-12Large Organization (50 developers + CI/CD)
Developer usage: 50 devs × ~$250/month = ~$12,500
CI/CD pipeline: 200 PRs/week = ~$200-300/month
Batch jobs: Weekly docs generation = ~$50-100/month
Total monthly: ~$13,000-15,000
Per developer: ~$260-300/month (including CI/CD overhead)Budget Management Guidance
Setting Budget Alerts
If you are using the Anthropic API directly, set spending limits in the Anthropic Console:
- Navigate to Settings > Billing > Usage Limits
- Set a monthly hard limit (requests will be rejected beyond this)
- Set a soft limit (you will receive email notifications)
- Monitor the usage dashboard weekly
Tracking Cost Per Task Type
Categorize your Claude Code usage to identify optimization targets:
Category | % of Spend | Optimization Potential
─────────────────┼────────────┼───────────────────────
PR Reviews | 25% | Switch to Haiku or Batch API
Feature Dev | 40% | Already optimal with Sonnet
Debugging | 20% | Point to files to reduce search
Docs Generation | 10% | Batch API, run overnight
Exploration | 5% | Use /clear often, keep sessions shortThe "Token Budget" Mindset
Think of your token budget like a computing budget:
- Do not optimize prematurely --- first understand where tokens go
- The biggest savings come from architectural decisions, not prompt tweaking (e.g., using subagents vs. one massive context)
- Caching is automatic and free --- just structure your workflow to enable it
- Time is money too --- a $0.50 session that saves you 2 hours of manual work is a bargain
Anatomy of a Claude Code Session Cost
To make token economics concrete, here is a detailed breakdown of a realistic 15-turn session where a developer asks Claude to implement a feature.
Session: "Add pagination to the users API endpoint"
Turn │ Action │ Input tk │ Output tk │ Cost (Sonnet)
──────┼──────────────────────────────┼──────────┼───────────┼─────────────
1 │ User prompt + system init │ 6,500 │ 300 │ $0.024
2 │ Read package.json │ 7,200 │ 150 │ $0.005 *
3 │ Read src/api/users.ts │ 9,800 │ 200 │ $0.006 *
4 │ Read src/db/queries.ts │ 12,100 │ 180 │ $0.006 *
5 │ Plan + explain approach │ 13,000 │ 1,200 │ $0.022
6 │ Edit src/api/users.ts │ 14,500 │ 800 │ $0.015
7 │ Edit src/db/queries.ts │ 16,200 │ 600 │ $0.012
8 │ Create test file │ 17,500 │ 1,500 │ $0.027
9 │ Run tests (bash) │ 19,800 │ 100 │ $0.007 *
10 │ Read test output │ 21,000 │ 400 │ $0.009
11 │ Fix failing test │ 22,500 │ 500 │ $0.011
12 │ Rerun tests (bash) │ 24,000 │ 100 │ $0.008 *
13 │ Read success output │ 25,200 │ 300 │ $0.009
14 │ Summary response │ 26,000 │ 800 │ $0.017
15 │ User says "thanks" │ 26,800 │ 150 │ $0.010
──────┼──────────────────────────────┼──────────┼───────────┼─────────────
│ Total │ │ │ $0.188
│ Without prompt caching │ │ │ $0.410
* Turns marked with * have very low output (tool calls only)Key observations from this breakdown:
- 70% of input tokens are cached after turn 2, cutting input costs dramatically
- Output-heavy turns (turns 5, 8) are the most expensive despite having fewer input tokens
- Tool-only turns (reads, bash commands) are cheap because they produce minimal output
- The entire session costs about $0.19 --- roughly 1,000 such sessions per $200 monthly budget
Where the Money Goes
Cost distribution in a typical session:
┌──────────────────────────────────────────────────┐
│ │
│ Output tokens (responses + edits) 52% │
│ ██████████████████████████ │
│ │
│ Input tokens (uncached new content) 28% │
│ ██████████████ │
│ │
│ Input tokens (cached prefix) 12% │
│ ██████ │
│ │
│ Cache write cost (first turn) 8% │
│ ████ │
│ │
└──────────────────────────────────────────────────┘This confirms that output tokens dominate costs even with prompt caching active. The most impactful optimization is controlling output length, not reducing input size.
Token Counting for Common File Types
Understanding how different file types map to tokens helps you estimate costs before a session:
File type │ Avg tokens per 100 lines │ Notes
─────────────────────┼─────────────────────────┼──────────────────────
Python │ 500-700 │ Whitespace-efficient
TypeScript │ 600-900 │ Type annotations add tokens
HTML │ 700-1,100 │ Verbose tags and attributes
JSON (config) │ 400-600 │ Repetitive structure
JSON (data) │ 800-1,200 │ Depends on string content
CSS/SCSS │ 500-800 │ Property-value pairs
Markdown │ 300-500 │ Mostly prose
SQL │ 400-600 │ Keyword-heavy
YAML │ 350-550 │ Indentation is whitespace
Go │ 500-700 │ Concise syntax
Rust │ 600-900 │ Lifetime annotations, genericsPractical example: If Claude reads a 500-line TypeScript file, that is roughly 3,000-4,500 input tokens. At Sonnet rates (uncached), that is about $0.009-0.014. If cached, it drops to $0.0009-0.0014.
Monitoring and Alerting
API Usage Dashboard
If you are using the Anthropic API, the console provides:
- Real-time usage graphs showing input/output tokens per hour
- Cost breakdowns by model and by API key
- Rate limit monitoring showing how close you are to limits
Custom Monitoring with the SDK
For teams that want granular tracking, extract usage from each API response:
import anthropic
from datetime import datetime
client = anthropic.Anthropic()
def tracked_request(messages, model="claude-sonnet-4-20250514"):
"""Make an API request and log usage metrics."""
response = client.messages.create(
model=model,
max_tokens=4096,
messages=messages
)
usage = response.usage
log_entry = {
"timestamp": datetime.now().isoformat(),
"model": model,
"input_tokens": usage.input_tokens,
"output_tokens": usage.output_tokens,
"cache_read": getattr(usage, "cache_read_input_tokens", 0),
"cache_write": getattr(usage, "cache_creation_input_tokens", 0),
}
# Calculate cost
rates = {
"claude-sonnet-4-20250514": {"input": 3.0, "output": 15.0, "cache": 0.30},
"claude-opus-4-20250514": {"input": 15.0, "output": 75.0, "cache": 1.50},
}
r = rates.get(model, rates["claude-sonnet-4-20250514"])
cost = (
(usage.input_tokens * r["input"] / 1_000_000) +
(usage.output_tokens * r["output"] / 1_000_000) +
(log_entry["cache_read"] * r["cache"] / 1_000_000)
)
log_entry["cost_usd"] = round(cost, 6)
# Send to your monitoring system
print(f"Cost: ${cost:.4f} | "
f"In: {usage.input_tokens} | "
f"Out: {usage.output_tokens} | "
f"Cached: {log_entry['cache_read']}")
return response
# Usage
response = tracked_request([
{"role": "user", "content": "Explain the adapter pattern in 2 sentences."}
])Team-Level Cost Dashboards
For organizations, aggregate per-developer and per-project costs:
Weekly cost report — Engineering Team
──────────────────────────────────────────────────────
Developer │ Sessions │ Total Cost │ Avg/Session
─────────────────┼──────────┼────────────┼────────────
Alice (backend) │ 45 │ $38.20 │ $0.85
Bob (frontend) │ 62 │ $29.10 │ $0.47
Carol (infra) │ 28 │ $52.40 │ $1.87 ←
Dave (mobile) │ 35 │ $22.80 │ $0.65
──────────────────────────────────────────────────────
Team total │ 170 │ $142.50 │ $0.84
Carol's avg/session is 2x the team average.
→ Investigation: using Opus for routine tasks.
→ Action: switch to Sonnet for non-architecture work.Key Takeaways
- Output tokens cost 5x more than input tokens --- controlling output verbosity is the single highest-impact cost optimization.
- Prompt caching reduces input costs by up to 90% in multi-turn sessions where the prefix stays stable.
- Model selection matters: Haiku is 4x cheaper than Sonnet, which is 5x cheaper than Opus. Match the model to the task complexity.
- For CI/CD, use the Batch API for 50% savings on non-interactive tasks.
- Point Claude to the right files to avoid expensive exploratory reads. A well-structured CLAUDE.md pays for itself in reduced token waste.
- Budget at the team level: ~$200-300/developer/month for active usage is a reasonable baseline with Sonnet.
See also: A10 Prompt Caching for cache mechanics, A11 Model Selection Guide for choosing the right model, and Ch3 Context Window Management for managing context size.