A07: Token 经济学
与 Claude Code 的每次交互都会消耗 Token,而 Token 需要花钱。理解 Token 经济学让你能从使用中获得最大价值 --- 无论你是使用订阅计划的个人开发者,还是管理工程组织 API 成本的团队负责人。
Token 的工作原理
BPE 分词
Claude 使用字节对编码(Byte Pair Encoding,BPE)将文本转换为 Token。BPE 的工作方式是学习训练数据中最频繁的字符序列,并为它们分配单个 Token ID。结果是常见词汇变成一个 Token,而罕见或专业术语被拆分成多个 Token。
常见模式:
"function" → 1 token (common English/programming word)
"the" → 1 token (most common English word)
"authentication" → 2 tokens ("authentic" + "ation")
"XMLHttpRequest" → 4 tokens (camelCase splits + unusual combo)
"дефрагментация" → 6 tokens (non-Latin text uses more tokens)Claude 代码相关的具体示例:
"return null;" → 3 tokens (return + null + ;)
"console.log(" → 3 tokens (console + .log + ()
"if (err !== null)" → 6 tokens
"import { useState } from 'react';" → 9 tokens经验法则
- 英文散文中 1 个 Token 大约 3-4 个字符
- 代码中 1 个 Token 大约 2-3 个字符(更多标点和结构)
- 100 个 Token 大约 75 个英文单词
- 一个典型的 200 行源文件是 1,500-3,000 个 Token
- 一个 1,000 行文件可能是 8,000-15,000 个 Token,取决于语言
什么被计为 Token
在 Claude Code 会话中,上下文的每一部分都会消耗 Token:
Token budget breakdown for one turn:
┌──────────────────────────────────────────────────────┐
│ System prompt (Claude Code internals) ~3,000 tk │
│ CLAUDE.md content ~500-5,000 │
│ Tool definitions ~2,000 tk │
│ Conversation history (all prior turns) variable │
│ Current user prompt ~100-500 │
│ Tool results (file contents, etc.) variable │
│ ─────────────────────────────────────────────────── │
│ = Total input tokens (you pay for all of this) │
│ │
│ Model response + tool calls │
│ = Total output tokens (priced differently) │
└──────────────────────────────────────────────────────┘各模型的 Token 单价
Anthropic 提供三个模型层级。以下价格截至 2025 年初(请查阅 Anthropic 定价页面获取最新费率):
| 模型 | 输入(每 1M Token) | 输出(每 1M Token) | 缓存读取(每 1M) | 最适合 |
|---|---|---|---|---|
| Claude Opus 4 | $15.00 | $75.00 | $1.50 | 复杂推理、架构设计 |
| Claude Sonnet 4 | $3.00 | $15.00 | $0.30 | 日常开发、均衡选择 |
| Claude Haiku 3.5 | $0.80 | $4.00 | $0.08 | 高吞吐量、简单任务 |
输入 vs 输出 Token 定价
所有模型的输出 Token 都比输入 Token 贵 5 倍。这是因为生成 Token 比处理 Token 需要更多计算。
这种不对称性有实际影响:
Scenario A: Claude reads 5,000 tokens, writes 200 tokens
Input cost: 5,000 / 1M × $3.00 = $0.015
Output cost: 200 / 1M × $15.00 = $0.003
Total: $0.018
Scenario B: Claude reads 1,000 tokens, writes 2,000 tokens
Input cost: 1,000 / 1M × $3.00 = $0.003
Output cost: 2,000 / 1M × $15.00 = $0.030
Total: $0.033场景 B 虽然总 Token 数更少,但成本几乎是场景 A 的两倍,因为输出 Token 主导了成本。这意味着让 Claude 生成冗余输出(详细解释、长代码块)是最大的成本驱动因素。
Prompt Caching 对成本的影响
当 Claude Code 发送的请求前缀与最近缓存的提示匹配时,缓存部分的费率比标准输入费率低 90%(Sonnet 为 $0.30/1M 而非 $3.00/1M)。
在典型的 Claude Code 会话中,系统提示和 CLAUDE.md 内容在各轮之间保持不变。这个前缀会自动被缓存:
Turn 1 (cache miss):
System prompt + CLAUDE.md: 5,000 tokens × $3.00/1M = $0.015
New content: 2,000 tokens × $3.00/1M = $0.006
Total input cost: $0.021
Turn 2 (cache hit on prefix):
Cached prefix: 5,000 tokens × $0.30/1M = $0.0015
New content: 3,000 tokens × $3.00/1M = $0.009
Total input cost: $0.0105 (50% savings)
Turn 10 (most of conversation is cached):
Cached prefix: 30,000 tokens × $0.30/1M = $0.009
New content: 2,000 tokens × $3.00/1M = $0.006
Total input cost: $0.015 (vs $0.096 without caching = 84% savings)详见 A10 Prompt Caching 了解完整机制。
Batch API 定价优势
对于非交互式用例(CI/CD 流水线、批量代码审查、文档生成),Batch API 提供输入和输出 Token 50% 的折扣。权衡是结果不会流式传输 --- 你提交一个批次,在 24 小时内收到结果。
| 场景 | 标准 API | Batch API | 节省 |
|---|---|---|---|
| 审查 50 个 PR(Sonnet) | ~$7.50 | ~$3.75 | 50% |
| 为 200 个文件生成文档(Haiku) | ~$3.20 | ~$1.60 | 50% |
| 架构分析(Opus) | ~$22.50 | ~$11.25 | 50% |
Batch API 非常适合不需要实时响应的 CI/CD 集成(参见第 12 章)。
成本优化策略
策略 1:选择合适的模型
不是每个任务都需要最强大的模型。实用的规则:
┌─────────────────────────────────────────────────────┐
│ Task Complexity │ Recommended Model │
│──────────────────────────│──────────────────────────│
│ Formatting, renaming, │ Haiku (cheapest) │
│ simple edits │ │
│──────────────────────────│──────────────────────────│
│ Feature implementation, │ Sonnet (balanced) │
│ debugging, refactoring │ │
│──────────────────────────│──────────────────────────│
│ Architecture decisions, │ Opus (most capable) │
│ complex multi-file │ │
│ reasoning │ │
└─────────────────────────────────────────────────────┘在 Claude Code 中,使用 /model 或点击状态栏中的模型指示器来切换模型。
策略 2:减少输入 Token 浪费
Token 浪费的最大来源是读取不必要的文件。减少浪费的技巧:
- 将 Claude 指向正确的文件:"修复
src/auth/jwt.ts中的 bug" 比 "修复代码库中某处的认证 bug" 更便宜。 - 使用 CLAUDE.md 描述项目结构:Claude 读取一次 CLAUDE.md(第一轮后被缓存),然后直接导航到正确的文件而非探索。
- 在不相关的任务之间使用
/clear:对话中的每一轮都携带完整历史。一个 50 轮的会话意味着第 50 轮发送了之前所有 49 轮的内容作为输入。 - 保持 CLAUDE.md 简洁:CLAUDE.md 中的每个 Token 都会随每次请求发送。一个 5,000 Token 的 CLAUDE.md 每轮花费约 $0.015(Sonnet,未缓存)或约 $0.0015(已缓存)。
策略 3:控制输出详细度
由于输出 Token 比输入贵 5 倍:
- 当你只想要代码变更时,添加"简洁回复"或"不需要解释"。
- 搜索任务时使用"只回复文件路径"。
- 在 CLAUDE.md 中添加规则如:"修改代码时,只显示更改的行,不要显示完整文件。"
策略 4:面向缓存的提示设计
组织你的工作流以最大化缓存命中:
- 保持 CLAUDE.md 内容稳定(会话期间不要频繁编辑)
- 将最稳定的上下文放在提示的开头
- 对相关任务使用长会话(增长的对话历史会保持缓存)
策略 5:使用子 Agent 进行并行工作
当 Claude 生成子 Agent 时(参见第 8 章),每个子 Agent 有自己的上下文。这意味着:
- 主 Agent 的上下文保持较小
- 子 Agent 并行运行,减少实际耗时
- 每个子 Agent 只读取它需要的文件
- 总 Token 使用量可能相似或略高,但完成时间缩短
策略 6:CI/CD 优化
对于自动化流水线:
- 对 lint 检查和简单验证使用 Haiku
- 对 PR 审查使用 Sonnet
- 对非紧急任务使用 Batch API(过夜文档、批量分析)
- 设置
--max-turns以限制失控的会话 - 监控每次运行的成本并设置预算告警
月度成本估算
以下估算假设使用 Sonnet 4 并启用 Prompt Caching。
个人开发者
Usage pattern: 40 interactive sessions/week, ~15 turns each
Input tokens: ~2M tokens/day (with caching: effective cost ~0.6M)
Output tokens: ~400K tokens/day
Daily cost: ~$8-12
Monthly cost: ~$180-270
With Max plan ($200/month subscription): Included in plan小团队(5 名开发者)
Usage pattern: Each developer uses Claude Code for ~4 hours/day
Combined daily tokens: ~10M input, ~2M output
Daily cost: ~$45-60
Monthly cost: ~$1,000-1,300
Cost per developer: ~$200-260/monthCI/CD 流水线(API 使用)
Usage pattern: 50 PRs/week, each reviewed by Claude
Per PR: ~50K input tokens, ~5K output tokens
Weekly: 2.5M input + 250K output
Monthly (Sonnet, standard API): ~$50-70
Monthly (Sonnet, Batch API): ~$25-35
Monthly (Haiku, Batch API): ~$7-12大型组织(50 名开发者 + CI/CD)
Developer usage: 50 devs × ~$250/month = ~$12,500
CI/CD pipeline: 200 PRs/week = ~$200-300/month
Batch jobs: Weekly docs generation = ~$50-100/month
Total monthly: ~$13,000-15,000
Per developer: ~$260-300/month (including CI/CD overhead)预算管理指南
设置预算告警
如果你直接使用 Anthropic API,在 Anthropic 控制台中设置支出限制:
- 导航到 Settings > Billing > Usage Limits
- 设置月度硬性限制(超出后请求将被拒绝)
- 设置软限制(你将收到邮件通知)
- 每周监控使用量仪表板
按任务类型追踪成本
对你的 Claude Code 使用进行分类以确定优化目标:
Category | % of Spend | Optimization Potential
─────────────────┼────────────┼───────────────────────
PR Reviews | 25% | Switch to Haiku or Batch API
Feature Dev | 40% | Already optimal with Sonnet
Debugging | 20% | Point to files to reduce search
Docs Generation | 10% | Batch API, run overnight
Exploration | 5% | Use /clear often, keep sessions short"Token 预算"思维
将你的 Token 预算看作计算预算:
- 不要过早优化 --- 先了解 Token 花在哪里
- 最大的节省来自架构决策,而非提示微调(例如使用子 Agent vs. 一个巨大上下文)
- 缓存是自动且免费的 --- 只需组织你的工作流使其生效
- 时间也是金钱 --- 一个花费 $0.50 但节省你 2 小时手动工作的会话是非常划算的
Claude Code 会话成本解剖
为了让 Token 经济学更具体,以下是一个现实的 15 轮会话的详细分解,其中开发者要求 Claude 实现一个功能。
会话:"为用户 API 端点添加分页"
Turn │ Action │ Input tk │ Output tk │ Cost (Sonnet)
──────┼──────────────────────────────┼──────────┼───────────┼─────────────
1 │ User prompt + system init │ 6,500 │ 300 │ $0.024
2 │ Read package.json │ 7,200 │ 150 │ $0.005 *
3 │ Read src/api/users.ts │ 9,800 │ 200 │ $0.006 *
4 │ Read src/db/queries.ts │ 12,100 │ 180 │ $0.006 *
5 │ Plan + explain approach │ 13,000 │ 1,200 │ $0.022
6 │ Edit src/api/users.ts │ 14,500 │ 800 │ $0.015
7 │ Edit src/db/queries.ts │ 16,200 │ 600 │ $0.012
8 │ Create test file │ 17,500 │ 1,500 │ $0.027
9 │ Run tests (bash) │ 19,800 │ 100 │ $0.007 *
10 │ Read test output │ 21,000 │ 400 │ $0.009
11 │ Fix failing test │ 22,500 │ 500 │ $0.011
12 │ Rerun tests (bash) │ 24,000 │ 100 │ $0.008 *
13 │ Read success output │ 25,200 │ 300 │ $0.009
14 │ Summary response │ 26,000 │ 800 │ $0.017
15 │ User says "thanks" │ 26,800 │ 150 │ $0.010
──────┼──────────────────────────────┼──────────┼───────────┼─────────────
│ Total │ │ │ $0.188
│ Without prompt caching │ │ │ $0.410
* Turns marked with * have very low output (tool calls only)从这个分解中可以得出的关键观察:
- 第 2 轮之后,70% 的输入 Token 被缓存,大幅降低了输入成本
- 输出密集型轮次(第 5、8 轮)尽管输入 Token 更少,却是最贵的
- 纯工具轮次(读取、bash 命令)因为产生的输出很少而很便宜
- 整个会话花费约 $0.19 --- 大约 1,000 个这样的会话对应 $200 月度预算
钱花在了哪里
Cost distribution in a typical session:
┌──────────────────────────────────────────────────┐
│ │
│ Output tokens (responses + edits) 52% │
│ ██████████████████████████ │
│ │
│ Input tokens (uncached new content) 28% │
│ ██████████████ │
│ │
│ Input tokens (cached prefix) 12% │
│ ██████ │
│ │
│ Cache write cost (first turn) 8% │
│ ████ │
│ │
└──────────────────────────────────────────────────┘这证实了即使在 Prompt Caching 启用的情况下,输出 Token 仍然主导成本。最有影响的优化是控制输出长度,而非减少输入大小。
常见文件类型的 Token 计数
了解不同文件类型如何映射到 Token 有助于你在会话前估算成本:
File type │ Avg tokens per 100 lines │ Notes
─────────────────────┼─────────────────────────┼──────────────────────
Python │ 500-700 │ Whitespace-efficient
TypeScript │ 600-900 │ Type annotations add tokens
HTML │ 700-1,100 │ Verbose tags and attributes
JSON (config) │ 400-600 │ Repetitive structure
JSON (data) │ 800-1,200 │ Depends on string content
CSS/SCSS │ 500-800 │ Property-value pairs
Markdown │ 300-500 │ Mostly prose
SQL │ 400-600 │ Keyword-heavy
YAML │ 350-550 │ Indentation is whitespace
Go │ 500-700 │ Concise syntax
Rust │ 600-900 │ Lifetime annotations, generics实际示例:如果 Claude 读取一个 500 行的 TypeScript 文件,大约是 3,000-4,500 个输入 Token。按 Sonnet 费率(未缓存),大约是 $0.009-0.014。如果缓存了,降至 $0.0009-0.0014。
监控与告警
API 使用量仪表板
如果你使用 Anthropic API,控制台提供:
- 实时使用量图表,按小时显示输入/输出 Token
- 成本分解,按模型和 API 密钥
- 速率限制监控,显示你离限制有多远
使用 SDK 自定义监控
对于想要细粒度追踪的团队,可以从每个 API 响应中提取使用数据:
import anthropic
from datetime import datetime
client = anthropic.Anthropic()
def tracked_request(messages, model="claude-sonnet-4-20250514"):
"""Make an API request and log usage metrics."""
response = client.messages.create(
model=model,
max_tokens=4096,
messages=messages
)
usage = response.usage
log_entry = {
"timestamp": datetime.now().isoformat(),
"model": model,
"input_tokens": usage.input_tokens,
"output_tokens": usage.output_tokens,
"cache_read": getattr(usage, "cache_read_input_tokens", 0),
"cache_write": getattr(usage, "cache_creation_input_tokens", 0),
}
# Calculate cost
rates = {
"claude-sonnet-4-20250514": {"input": 3.0, "output": 15.0, "cache": 0.30},
"claude-opus-4-20250514": {"input": 15.0, "output": 75.0, "cache": 1.50},
}
r = rates.get(model, rates["claude-sonnet-4-20250514"])
cost = (
(usage.input_tokens * r["input"] / 1_000_000) +
(usage.output_tokens * r["output"] / 1_000_000) +
(log_entry["cache_read"] * r["cache"] / 1_000_000)
)
log_entry["cost_usd"] = round(cost, 6)
# Send to your monitoring system
print(f"Cost: ${cost:.4f} | "
f"In: {usage.input_tokens} | "
f"Out: {usage.output_tokens} | "
f"Cached: {log_entry['cache_read']}")
return response
# Usage
response = tracked_request([
{"role": "user", "content": "Explain the adapter pattern in 2 sentences."}
])团队级别成本仪表板
对于组织,汇总每个开发者和每个项目的成本:
Weekly cost report — Engineering Team
──────────────────────────────────────────────────────
Developer │ Sessions │ Total Cost │ Avg/Session
─────────────────┼──────────┼────────────┼────────────
Alice (backend) │ 45 │ $38.20 │ $0.85
Bob (frontend) │ 62 │ $29.10 │ $0.47
Carol (infra) │ 28 │ $52.40 │ $1.87 ←
Dave (mobile) │ 35 │ $22.80 │ $0.65
──────────────────────────────────────────────────────
Team total │ 170 │ $142.50 │ $0.84
Carol's avg/session is 2x the team average.
→ Investigation: using Opus for routine tasks.
→ Action: switch to Sonnet for non-architecture work.核心要点
- 输出 Token 比输入 Token 贵 5 倍 --- 控制输出详细度是影响最大的单一成本优化。
- Prompt Caching 在多轮会话中可降低高达 90% 的输入成本,前提是前缀保持稳定。
- 模型选择很重要:Haiku 比 Sonnet 便宜 4 倍,Sonnet 比 Opus 便宜 5 倍。将模型匹配到任务复杂度。
- CI/CD 用 Batch API 可对非交互式任务节省 50%。
- 将 Claude 指向正确的文件以避免昂贵的探索性读取。一个结构良好的 CLAUDE.md 在减少 Token 浪费方面能自我回本。
- 按团队级别做预算:活跃使用 Sonnet 时,每个开发者每月约 $200-300 是合理的基线。
另见:A10 Prompt Caching 了解缓存机制,A11 模型选择指南 了解如何选择合适的模型,以及第 3 章 上下文窗口管理 了解上下文大小管理。