A02: 上下文工程
相关章节: 第 3 章 上下文窗口管理、第 15 章 生产工作流
提示工程(Prompt Engineering)问的是"我应该对模型说什么?"上下文工程(Context Engineering)问的是一个更大的问题:"当模型生成回复时,它应该看到什么?"本附录介绍了从提示工程到上下文工程的范式转变,并提供了管理 Claude Code 所有输入(不仅仅是你输入的文本)的实用策略。
范式转变:从提示工程到上下文工程
提示工程时代 (2022-2024)
当 ChatGPT 发布时,世界发现了提示工程。用户了解到措辞很重要:"think step by step(逐步思考)"比"give me the answer(给我答案)"效果更好。围绕精心设计完美指令,一整个学科应运而生。
提示工程将模型视为一个文本输入、文本输出的函数。你的提示是唯一可以调节的杠杆。少样本示例(Few-shot)、角色赋予("你是一位资深工程师")和思维链(Chain-of-Thought)等技术都专注于优化这单一文本输入。
这对聊天式交互效果很好。但当模型变成 Agent 后就行不通了。
上下文工程时代 (2025-2026)
当 Claude Code 和类似的 Agent 工具出现后,规则改变了。你输入的提示变成了模型实际看到内容的一小部分。在典型的 Claude Code 会话中,模型的输入包括:
Token budget breakdown for a typical Claude Code session:
┌──────────────────────────────────────────────────────┐
│ System prompt .............. ~3,000 tokens (fixed) │
│ Tool definitions ........... ~5,000 tokens (fixed) │
│ CLAUDE.md files ............ ~2,000 tokens (fixed) │
│ Memory files ............... ~500 tokens (fixed) │
│ Conversation history ....... ~varies (grows) │
│ Tool results (file reads) .. ~varies (grows) │
│ Your current prompt ........ ~100 tokens (tiny) │
│ │
│ Your prompt is typically < 1% of total context │
└──────────────────────────────────────────────────────┘上下文工程正视这一现实。它不再优化单一输入,而是架构模型周围的整个信息环境。这包括:
- 哪些指令持久存在于会话之间(CLAUDE.md、记忆文件)
- 哪些信息进入上下文(Claude 读取的文件、运行的命令)
- 哪些信息退出上下文(压缩、摘要)
- 何时重置上下文(新会话、
/clear)
Anthropic 自己的研究团队也强调了这一转变。正如 Toby Jia-Jun Li(Anthropic 研究员)在 2025 年指出的:对 Agent 性能最大的改进不是来自更好的提示,而是来自更好的上下文管理。
上下文栈:输入的层次
Claude Code 从多个层次组装其上下文,每层有不同的特性:
第 1 层:系统提示(你无法控制)
系统提示在对话开始前由 Claude Code 的运行框架注入。它包含:
- Claude 的核心行为指令
- 带有 JSON Schema 的工具定义
- 权限规则和安全约束
- 当前日期和环境信息
这一层是固定的 -- 你无法修改它。但理解它有助于你避免冗余指令。例如,你不需要告诉 Claude"使用 Read 工具来读取文件",因为系统提示已经建立了工具调用行为。
Token 成本:大约 3,000-8,000 Token,取决于工具定义和配置。
第 2 层:CLAUDE.md 文件(你控制)
CLAUDE.md 文件是你进行持久化上下文工程的主要杠杆。它们在会话开始时加载,并在整个对话中保持在上下文中。
Resolution order (all loaded, later overrides earlier):
1. ~/.claude/CLAUDE.md (global preferences)
2. ~/project/CLAUDE.md (project root)
3. ~/project/.claude/CLAUDE.md (project config dir)
4. ~/project/src/CLAUDE.md (directory-level)设计原则:CLAUDE.md 是指令上下文 -- Claude 应始终遵循的规则、模式和约束。它不是用来堆放文档或代码示例的地方。保持精简。
# Good CLAUDE.md (focused, ~500 tokens)
## Code Style
- Use TypeScript strict mode
- Prefer functional patterns; avoid classes
- Error handling: return Result<T, Error>, never throw
## Testing
- Every new function needs a test in __tests__/
- Use vitest, not jest
- Mock external APIs with msw
## Architecture
- Routes in src/routes/, services in src/services/
- Database access only through src/db/ layer# Bad CLAUDE.md (bloated, ~5,000 tokens)
## Full API Documentation
[entire API spec pasted here]
## Database Schema
[full SQL schema pasted here]
## Meeting Notes
[notes from last sprint planning]Token 成本:目标在 500-2,000 Token。CLAUDE.md 中的每个 Token 在每个回合都会被加载,因此臃肿的 CLAUDE.md 会在整个会话中导致成本二次方增长。
第 3 层:记忆文件(Claude 管理)
记忆文件存储 Claude 了解到的关于你和项目的偏好。它们在会话之间持久化,并自动加载。
记忆对于 Claude 在工作中发现的偏好很有用 -- 你偏好的变量命名、提交消息格式、使用的测试运行器等。但记忆不是 CLAUDE.md 的替代品。CLAUDE.md 用于所有开发者都应遵循的项目规则。记忆用于你个人的偏好。
Token 成本:通常 200-800 Token。
第 4 层:对话历史(你们共同构建)
这是上下文工程变得关键的地方。对话历史随每次交互增长:
Turn 1: You ask a question +50 tokens
Claude reads 3 files +3,000 tokens
Claude responds +500 tokens
Turn 2: You ask a follow-up +30 tokens
Claude reads 2 more files +2,000 tokens
Claude edits a file +400 tokens
Claude runs tests +1,500 tokens
Claude responds +300 tokens
Turn 3: Already at ~8,000 tokens of history
...
Turn 20: Context is 60-80% full对话历史是增长最快的层,也是你通过工作流选择拥有最大控制权的层。
第 5 层:工具结果(由操作生成)
每次 Read、Bash、Grep 或其他工具调用都会产生进入上下文的结果。单个大文件的 cat 可以注入数千 Token。一个测试套件的输出可以注入数万 Token。
这通常是最浪费的层。Claude 读取不需要的文件,在有简洁替代方案时运行冗长的命令,在有针对性的搜索就足够时进行广泛探索。
上下文预算:明智分配 Token
把 200K Token 的上下文窗口看作一个预算。花在一件事上的每个 Token 就是另一件事上不可用的 Token。
黄金比例:指令上下文 vs 工作上下文
上下文工程将窗口分为两类:
- 指令上下文:系统提示、CLAUDE.md、记忆文件、工具定义。告诉 Claude 如何行为。
- 工作上下文:对话历史、工具结果、当前任务。Claude 工作的素材。
The golden ratio:
┌──────────────────────────────────────────────────────┐
│ Instruction context: 10-20% of window │
│ ████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ │
│ │
│ Working context: 30-50% of window │
│ ░░░░░░░░██████████████████████████░░░░░░░░░░░░░░░ │
│ │
│ Reserved headroom: 30-50% of window │
│ ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░████████████████ │
│ │
│ Total used: ideally 40-60% (the golden zone) │
└──────────────────────────────────────────────────────┘为什么要预留空间? Claude 需要空间来生成回复(输出 Token)和调用产生结果的工具(这些结果会重新进入上下文)。如果窗口已满 90%,Claude 无法再读取一个文件而不触发压缩或丢失信息。
按任务类型分配预算
不同任务有不同的上下文需求:
| 任务类型 | 指令需求 | 工作上下文需求 | 推荐方法 |
|---|---|---|---|
| 简单编辑 | 低(只需文件位置) | 低(一个文件) | 直接提示,单次会话 |
| Bug 修复 | 中(症状、约束) | 中(几个文件 + 测试输出) | 先调查再行动 |
| 功能构建 | 中(需求、模式) | 高(多个文件、测试) | 频繁检查点 |
| 架构审查 | 低(只需"审查") | 非常高(需读取多个文件) | 分拆到多个会话 |
| 大型重构 | 高(约束、安全) | 非常高(需编辑多个文件) | 多会话 + 计划 |
实用策略
策略 1:在压缩前主动总结
当 Claude 的自动压缩(Compaction)触发时,它会摘要对话以适应窗口。但自动压缩是有损的 -- 它根据时效性和表面相关性决定保留什么,这可能与你的优先级不匹配。
模式:在压缩触发前手动总结:
You (at ~50% context usage):
"Before we continue, summarize what we've done so far and what remains.
Include: (1) the architecture decisions we made, (2) the files we've
changed, (3) the remaining tasks. Then I'll /clear and we'll continue
with your summary."然后开始新会话:
You (new session):
"Continuing a refactor. Here's where we left off:
[paste Claude's summary]
Next task: implement the database migration for the new schema."这让你控制什么信息存活于上下文重置,而不是交给自动压缩决定。
策略 2:探索前设置检查点
在进行可能消耗大量上下文的探索(读取多个文件、运行大量测试)之前,创建一个检查点:
You: "We're about to investigate the payment processing pipeline.
Before we start, confirm: we've completed the auth refactor (files:
src/auth/service.ts, src/auth/middleware.ts, src/auth/types.ts), all
auth tests pass, and the remaining work is the payments module. Correct?"
Claude: "Yes, that's correct. [confirms details]"
You: "Good. Now investigate the payments pipeline. Start with
src/payments/index.ts."检查点有两个作用:
- 在对话历史中创建一个摘要化的锚点,压缩时更可能被保留
- 在上下文被新信息搅乱之前验证共识
策略 3:每个会话一个任务
最强大的上下文工程策略也是最简单的:每个会话只做一件事。
# Session 1: Fix the auth bug
claude "Fix the 401 error in token rotation. The issue is in src/auth/."
# Session 2: Add rate limiting
claude "Add rate limiting to /api/upload. Max 10/minute per user. Use Redis."
# Session 3: Update tests
claude "The auth middleware changed. Update tests in tests/auth/ to match."每个会话从干净的上下文窗口开始。Claude 有最大的工作上下文空间。不存在早期任务污染后续任务的上下文腐化风险。
当单任务会话不可行时(例如多步功能),使用"压缩前主动总结"模式在较长对话中创建自然的会话边界。
策略 4:前置关键上下文
Transformer 模型的注意力机制(Attention Mechanism)在上下文窗口中不是均匀分布的。研究(包括 Anthropic 自己关于长上下文检索的研究)表明,模型对以下位置的注意力最强:
- 上下文的最开头(系统提示、CLAUDE.md)
- 上下文的最末尾(你最近的提示)
- 独特或结构化的内容(标题、代码块、格式化列表)
中间的内容受到的注意力较少。这有时被称为"中间遗失"效应(Lost in the Middle,Liu et al., 2023)。
实际影响:
- 将最重要的规则放在 CLAUDE.md 中(始终在上下文开头)
- 在当前提示中重复关键约束(始终在上下文末尾)
- 使用结构化格式(标题、要点列表)用于不可遗漏的指令
- 不要依赖 30 轮对话中第 5 轮的一个随意提及会被记住
# Effective: critical constraint repeated at point of action
You (turn 15): "Now implement the payments endpoint. IMPORTANT: do NOT
modify the existing /api/users endpoint -- it's in production and any
change will break the mobile app."策略 5:控制工具结果大小
工具结果通常是上下文中最大且最浪费的条目。你可以影响它们的大小:
# Bad: reads entire file when you only need the function signature
You: "Read the auth module and find the validateToken function."
# Claude reads the entire 500-line file (+3,000 tokens)
# Good: directs Claude to search, not read
You: "Find the validateToken function signature in src/auth/.
Use grep, don't read the whole file."
# Claude greps for the function (+200 tokens)# Bad: runs verbose test output
You: "Run the tests."
# Claude runs pytest with full output (+5,000 tokens)
# Good: controls output verbosity
You: "Run pytest with -q flag. Only show failures."
# Claude runs pytest -q (+500 tokens)这看起来像微优化,但在 20 轮会话中,控制工具结果大小可能是保持在黄金区域和触发上下文腐化之间的差别。
上下文腐化:质量为何下降
上下文腐化的表现
上下文腐化(Context Rot)是随着上下文窗口填满,Claude 输出质量逐步下降的现象。它表现为:
- 忘记需求:Claude 忽略你之前声明的约束
- 矛盾输出:Claude 做出与之前轮次中同意的相反的事情
- 幻觉上下文:Claude 引用不存在的文件或函数(将早期文件内容与当前状态混淆)
- 具体性下降:Claude 的回复变得更通用,不再针对你的代码库定制
- 重复增加:Claude 重新读取已读文件,或重新解释已解释过的内容
发生的原因
上下文腐化有多个交互原因:
注意力稀释:随着上下文增长,每个 Token 与更多 Token 竞争注意力权重。第 2 轮的重要指令在被 100K Token 的文件内容和测试输出包围时获得的注意力更少。
信噪比降低:会话早期,上下文主要是信号(你的指令、相关代码)。会话晚期,主要是噪声(中间结果、放弃的方案、冗长的输出)。模型难以提取信号。
位置偏差:"中间遗失"效应意味着长对话中间轮次的信息最可能被忽视。你在第 5 轮的关键约束到第 20 轮可能已经不可见。
压缩伪影:当自动压缩触发时,它摘要早期轮次。这些摘要不可避免地丢失细微差别。经过多次压缩循环后,模型对原始需求的理解可能已经严重偏移。
如何对抗上下文腐化
| 策略 | 何时使用 | 如何帮助 |
|---|---|---|
| 每个会话一个任务 | 始终,在可行时 | 完全防止腐化 |
任务间使用 /clear | 做多个任务时 | 在不相关工作间重置上下文 |
| 手动总结 | 在上下文使用约 50% 时 | 你控制什么被保留 |
| 重复关键约束 | 在每次重大操作前 | 刷新对重要规则的注意力 |
| 短会话 | 用于高精度工作 | 减少腐化累积的时间 |
管道模式 (| claude -p) | 用于独立子任务 | 每次调用获得干净上下文 |
进阶:团队的上下文工程
共享 CLAUDE.md 作为团队上下文
当团队在代码仓库中共享 CLAUDE.md 文件时,他们创建了一个共享的上下文工程层。每个团队成员的 Claude Code 会话都以相同的指令、模式和约束开始。
这很强大,但需要纪律:
# Team CLAUDE.md best practices
## DO include:
- Code style rules that apply to ALL files
- Architecture patterns that ALL code must follow
- Testing requirements that ALL changes must meet
- File naming and organization conventions
## DO NOT include:
- Individual preferences (use personal ~/.claude/CLAUDE.md)
- Temporary instructions ("until the migration is done")
- Documentation that belongs in docs/ or README
- Rules for code that no longer existsCI/CD 中的上下文工程
当 Claude Code 在 CI/CD 中运行(管道模式)时,上下文工程采取不同形式。没有交互式对话 -- 你只有一次提示机会。每个 Token 都必须有价值:
# CI/CD context engineering: front-load everything
git diff main..HEAD | claude -p "$(cat <<'EOF'
You are reviewing a pull request. The diff is piped as input.
Project rules:
- TypeScript strict mode, no any types
- All public functions need JSDoc
- No console.log in production code (use the logger)
- SQL queries must use parameterized statements
Review for:
1. Rule violations from the list above
2. Logic bugs
3. Security issues (OWASP Top 10)
Output format: one finding per line, with severity (HIGH/MEDIUM/LOW),
file path, line number, and description.
EOF
)"在这个模式中,系统提示和 CLAUDE.md 被管道提示中的显式上下文块替代。diff 提供工作上下文。没有对话历史需要管理。
上下文工程检查清单
在开始新项目或优化现有工作流时使用此清单:
- [ ] CLAUDE.md 在 2,000 Token 以内,仅包含活跃规则
- [ ] CLAUDE.md 中没有粘贴文档或代码示例
- [ ] 复杂任务被拆分为单会话单元
- [ ] 关键约束在行动点处的提示中被重复
- [ ] 工具结果受到控制(先 grep 再 read,测试使用 quiet 标志)
- [ ] 使用
/cost或/context监控上下文使用量 - [ ] 在 60% 阈值前进行手动摘要
- [ ] 团队成员在仓库中共享经过审查的 CLAUDE.md
- [ ] 个人偏好在
~/.claude/CLAUDE.md中,而非项目文件中
核心要点
- 上下文工程关乎所有输入,而不仅仅是提示。 你的提示不到 Claude 看到内容的 1%。优化那另外 99%。
- 黄金区域是 40-60% 的上下文使用量。 超过 60%,质量明显下降。超过 80%,严重下降。
- 每个会话一个任务是最强大的策略。 干净的上下文胜过任何巧妙的上下文管理。
- CLAUDE.md 是你的持久化指令层。 保持精简、聚焦并持续维护。像对待代码一样 -- 审查它、版本化它、精简它。
- 积极控制工具结果大小。 用 grep 代替 read,命令使用 quiet 标志,用有针对性的搜索代替广泛探索。
- 上下文腐化是真实的、可预测的。 通过设置检查点、摘要和在自然边界处重置来应对。
另见:A01 提示工程 了解如何专门优化提示层,以及 A07 Token 经济学 了解上下文预算决策的成本影响。