A14: RAG vs 长上下文
相关章节:第 7 章 MCP 服务器与外部数据
检索增强生成(Retrieval-Augmented Generation, RAG)和长上下文窗口是两种为 LLM 提供训练数据之外信息的竞争策略。大多数 AI 应用使用其中之一或两者兼用。Claude Code 采用了一种不寻常的方法:它既不使用传统的 RAG,也不使用暴力上下文加载。相反,它实现了一种基于工具的检索策略,结合了两者的优势。本附录解释原因、每种方法何时胜出,以及如何为大型代码库构建混合系统。
传统 RAG 的工作原理
RAG(检索增强生成)是一种在查询时从数据库中获取相关文档并注入到 LLM 提示中的模式:
┌──────────────────────────────────────────────────────────┐
│ Traditional RAG Pipeline │
│ │
│ User Query ──► Embedding Model ──► Vector Search ──┐ │
│ │ │
│ Document Store ◄───── Top-K Results ◄──────────────┘ │
│ │ │
│ ▼ │
│ Retrieved Chunks ──► [System Prompt + Chunks + Query] │
│ │ │
│ ▼ │
│ LLM Response │
└──────────────────────────────────────────────────────────┘RAG 流程:
- 索引:文档被分割成块并转换为向量嵌入(Vector Embeddings)
- 检索:用户查询被嵌入,通过向量相似性搜索找到最近的块
- 增强:检索到的块被添加到提示前面
- 生成:LLM 使用检索到的上下文生成回复
RAG 的优势:
- 可以搜索数百万个文档
- 只加载相关内容(高效的上下文使用)
- 索引可以增量更新,无需重新处理所有内容
- 适用于任何 LLM,无论上下文窗口大小
RAG 的弱点:
- 嵌入质量决定检索质量(语义不匹配很常见)
- 块边界可能拆分相关信息
- Top-K 检索可能遗漏重要文档
- 需要基础设施:向量数据库、嵌入流程、索引管理
- 无法跨未一起检索的文档进行推理
长上下文的工作原理
长上下文方法更简单:将模型可能需要的所有内容直接加载到上下文窗口中。
┌──────────────────────────────────────────────────────────┐
│ Long Context Approach │
│ │
│ [System Prompt] │
│ [Full Document 1] │
│ [Full Document 2] │
│ [Full Document 3] │
│ [...all potentially relevant documents...] │
│ [User Query] │
│ │ │
│ ▼ │
│ LLM Response │
│ (model attends to everything simultaneously) │
└──────────────────────────────────────────────────────────┘长上下文的优势:
- 没有检索错误:模型看到所有内容
- 可以跨远距离信息进行推理
- 不需要基础设施:无向量数据库、无嵌入流程
- 实现和调试更简单
长上下文的弱点:
- 上下文窗口有限制(截至 2026 年为 200k-1M Token)
- 成本随上下文大小线性增长
- "中间遗漏"效应(Lost in the Middle):模型可能低估很长上下文中间部分的信息权重
- 无法处理真正海量的语料库(数百万文档)
Claude Code 的方法:既非 RAG 也非长上下文
Claude Code 不使用传统 RAG。没有向量数据库,没有嵌入流程,没有块检索。Claude Code 也不会盲目地将整个文件加载到上下文中。相反,它使用基于工具的检索策略:
┌──────────────────────────────────────────────────────────┐
│ Claude Code's Retrieval Strategy │
│ │
│ User Query ──► Claude reasons about what info it needs │
│ │ │
│ ▼ │
│ Glob (find files by pattern) │
│ Grep (search content by regex) │
│ Read (load specific file sections) │
│ │ │
│ ▼ │
│ Claude reviews results │
│ Decides if more info is needed │
│ │ │
│ ┌── YES ─┘──── NO ──┐ │
│ ▼ ▼ │
│ More tool calls Generate response │
└──────────────────────────────────────────────────────────┘这与 RAG 和长上下文都有根本区别:
| 方面 | RAG | 长上下文 | Claude Code |
|---|---|---|---|
| 谁决定检索什么? | 嵌入相似度(自动化) | 开发者(手动) | Claude(Agent 式) |
| 检索机制 | 向量搜索 | 全部加载 | Glob/Grep/Read 工具 |
| 能否迭代? | 否(单次检索步骤) | 否(固定上下文) | 是(多回合工具循环) |
| 需要的基础设施 | 向量数据库 + 嵌入 | 无 | 文件系统访问 |
| 处理代码效果好? | 差 | 是 | 是 |
Glob/Grep/Read 模式详解
当 Claude Code 需要在代码库中查找信息时,它使用三步模式:
步骤 1:Glob -- 通过名称/路径模式查找候选文件
Claude thinks: "The user mentioned authentication. Let me find auth-related files."
Tool call: Glob("**/auth*", "**/*auth*")
Result: [src/auth/middleware.ts, src/auth/types.ts, src/auth/service.ts, tests/auth.test.ts]步骤 2:Grep -- 搜索文件内容以查找特定模式
Claude thinks: "I need to find where tokens are validated."
Tool call: Grep("validateToken|verifyToken|checkToken", "src/auth/")
Result: [src/auth/middleware.ts:42, src/auth/service.ts:118]步骤 3:Read -- 加载特定的相关部分
Claude thinks: "Let me read the token validation logic."
Tool call: Read("src/auth/middleware.ts", lines 35-60)
Result: [actual code content]这种模式是一种检索形式,但它是:
- 语义化的:Claude 基于理解来决定搜索什么,而不是基于嵌入相似度
- 迭代的:Claude 可以根据初始结果改进搜索
- 精确的:Claude 读取确切的行,而不是固定大小的块
- 代码感知的:模式匹配比向量嵌入更适合代码
为什么这对代码比 RAG 效果更好
代码具有使传统 RAG 表现不佳的特性:
代码是结构化的,不是语义化的:字符串 "authenticate" 可能出现在注释、函数名、import 语句和测试描述中——每个的相关性不同。向量嵌入抹平了这种结构。
代码有长距离依赖:
auth/middleware.ts中的函数依赖于auth/types.ts中的类型,并在tests/auth.test.ts中被测试。RAG 独立检索块,可能遗漏这些联系。Claude Code 的迭代工具调用跟踪依赖链。代码命名是有信息量的:文件路径(
src/auth/middleware.ts)和函数名(validateJWT)对代码导航比向量相似度携带更多信号。Glob 和 Grep 直接利用这一点。代码频繁变化:RAG 在代码变化时需要重新索引。Claude Code 的工具直接读取文件系统,所以它们总是看到当前状态。
何时 RAG 更好
Claude Code 的方法有局限性。在以下情况下 RAG 是更好的选择:
极大型代码库(1M+ 行代码)
如果你的代码库超出了通过合理数量的 Glob/Grep/Read 工具调用可以导航的范围,RAG 提供更快的初始检索:
Codebase size Claude Code approach RAG approach
────────────── ──────────────────── ─────────────
10K LOC Excellent Overkill
50K LOC Excellent Unnecessary
200K LOC Good Comparable
500K LOC Workable (many calls) Better initial retrieval
1M+ LOC Slow (too many calls) Significantly faster外部知识库
对于代码库之外的信息——文档、Wiki 页面、Confluence 文章、Slack 历史——Claude Code 没有原生访问能力。RAG(或基于 MCP 的检索)弥补了这个差距:
# MCP server providing RAG over company documentation
# .claude/settings.json
{
"mcpServers": {
"company-docs": {
"command": "node",
"args": ["./mcp-servers/docs-rag/server.js"],
"env": {
"VECTOR_DB_URL": "http://localhost:6333",
"COLLECTION": "company_docs"
}
}
}
}多仓库搜索
Claude Code 在单个仓库内运行。对于跨仓库搜索(例如,"找到所有依赖我们认证库的服务"),索引多个仓库的 RAG 系统更实用。
MCP 作为检索层
MCP(Model Context Protocol,模型上下文协议)连接了 Claude Code 和外部检索系统。你可以构建实现 RAG 并将其作为工具暴露给 Claude Code 调用的 MCP 服务器:
Claude Code ──► MCP Server ──► Vector Database ──► Retrieved Context
│
├── search_docs(query) → relevant documentation
├── search_code(query) → code from other repos
└── search_tickets(query) → related JIRA issues提供语义代码搜索的 MCP 服务器示例:
// mcp-servers/code-search/server.ts
import { McpServer } from "@modelcontextprotocol/sdk/server";
const server = new McpServer({
name: "code-search",
version: "1.0.0",
});
server.tool(
"semantic_search",
"Search the codebase using semantic similarity",
{
query: { type: "string", description: "Natural language query" },
top_k: { type: "number", description: "Number of results", default: 5 },
},
async ({ query, top_k }) => {
// 1. Embed the query
const embedding = await embedModel.embed(query);
// 2. Search the vector database
const results = await vectorDB.search({
collection: "codebase",
vector: embedding,
limit: top_k,
});
// 3. Return formatted results
return {
content: results.map((r) => ({
type: "text",
text: `File: ${r.metadata.file}:${r.metadata.line}\n${r.payload.code}`,
})),
};
}
);然后 Claude Code 可以将此工具与其原生的 Glob/Grep/Read 工具一起使用,结合语义搜索和结构化导航。
大型 Monorepo 的混合策略
对于大型 Monorepo(500K+ 行代码),最优方案是将 Claude Code 的原生工具与 RAG 结合:
策略 1:用 RAG 发现,用工具深入
"Find the rate limiter" ──► RAG (fast semantic search across 1M LOC)
│
▼
"src/middleware/rate-limit.ts"
│
▼
"Understand and modify it" ──► Read/Grep/Edit (precise tool-based work)使用 RAG 缩小在哪里查找的范围,然后切换到 Claude Code 的原生工具来理解和修改代码。这结合了 RAG 的广度与基于工具的检索的精度。
策略 2:预构建的代码地图
不使用 RAG,而是预先计算代码库的高级地图并将其包含在 CLAUDE.md 中:
<!-- CLAUDE.md -->
## Codebase Map
### Core Services (src/services/)
- auth/: JWT authentication, token management (3K LOC)
- payments/: Stripe integration, billing (5K LOC)
- notifications/: Email, SMS, push notifications (2K LOC)
### API Layer (src/api/)
- routes/: Express route handlers
- middleware/: Auth, rate limiting, logging
- validators/: Zod schemas for request validation
### Shared (src/shared/)
- types/: TypeScript interfaces shared across services
- utils/: Helper functions (date formatting, string manipulation)
- constants/: App-wide configuration constants这相当于一个目录。Claude Code 读取地图,识别相关区域,然后使用 Glob/Grep/Read 深入具体细节。不需要向量数据库。
策略 3:分层检索
Tier 1: CLAUDE.md codebase map (always in context)
│
▼ (Claude identifies area of interest)
Tier 2: Glob/Grep to find specific files
│
▼ (Claude identifies files to read)
Tier 3: Read specific file sections
│
▼ (if needed, Claude searches further)
Tier 4: MCP semantic search for cross-cutting concerns每一层逐步更昂贵但更精确。大多数任务在 Tier 2-3 就能解决。只有横切关注点的任务(例如,"找到所有处理错误的地方")才受益于 Tier 4。
比较表
| 维度 | 传统 RAG | 长上下文 | Claude Code (Glob/Grep/Read) | 混合 (Claude Code + MCP RAG) |
|---|---|---|---|---|
| 最大代码库 | 无限 | ~200K Token(~5K 行) | ~200K 行(有效) | 无限 |
| 检索质量 | 取决于嵌入 | 完美(看到所有) | 高(语义工具使用) | 最高 |
| 延迟 | 快(单次数据库查询) | 慢(大提示) | 中(多回合工具) | 中 |
| 每次查询成本 | 低(小上下文) | 高(大上下文) | 中(迭代读取) | 中高 |
| 设置成本 | 高(向量数据库、索引) | 无 | 无 | 中(MCP 服务器) |
| 处理代码效果好? | 差 | 是 | 是 | 是 |
| 迭代改进? | 否 | 否 | 是 | 是 |
| 跨仓库搜索? | 是 | 否 | 否 | 是 |
| 始终最新? | 否(需要重新索引) | 是 | 是 | 部分 |
上下文窗口何时让 RAG 过时
随着上下文窗口增长(2024 年 200K,2025 年 1M,2026 年以后可能更大),RAG 和长上下文之间的收支平衡点在移动:
RAG advantage zone
│
│ xxxxxxxx
│ xxxxxxxxx
│ xxxxxxxxxx
│ xxxxxxxxxxx Long context advantage zone
│ xxxxxxxxxxxx
│ xxxxxxxxxxxxx oooooo
│ xxxxxxxxxxxxxx oooooooo
│ xxxxxxxxxxxxxxx oooooooooo
│ xxxxxxxxxxxxxxxx oooooooooooo
│ xxxxxxxxxxxxxxxxx ooooooooooooo
└──────────────────────────────────────── Context window size
32K 128K 200K 500K 1M 4M 16M今天(1M 上下文):RAG 对于超过约 200K 行代码的代码库或外部知识库仍然必要。对于大多数单仓库项目,Claude Code 的原生工具就够了。
未来(4M+ 上下文):大多数单仓库项目将完全放入上下文。RAG 将仅在真正海量的代码库、多仓库搜索和外部知识集成中保持相关性。
趋势很明确:对于代码任务,长上下文正在蚕食 RAG 的地盘。但 RAG 永远不会完全消失,因为外部知识库和多源搜索本质上是检索问题。
实际建议
中小型项目(< 100K 行代码)
使用 Claude Code 的原生工具。不需要 RAG 基础设施:
# Just use Claude Code normally -- it will find what it needs
claude "Fix the authentication bug in the middleware layer"大型项目(100K - 500K 行代码)
在 CLAUDE.md 中添加代码库地图,并考虑使用 Opus 的 1M 上下文窗口:
# Use Opus for tasks spanning the full codebase
claude --model claude-opus-4-7-20250219 \
"Audit all API endpoints for consistent error handling"超大型项目(500K+ 行代码)或多仓库
构建提供语义搜索的 MCP 服务器:
# Set up MCP-based code search
# Then Claude Code can use both native tools AND semantic search
claude "Find all services that depend on the user authentication module"
# Claude will use both Grep and the MCP semantic_search tool外部知识(文档、Wiki、工单)
通过 MCP 的 RAG 是正确的方法——这些信息不在文件系统上:
# MCP server wrapping your documentation system
# Claude can search docs alongside code
claude "How does the billing webhook work? Check both the code and the design docs."核心要点
- Claude Code 不使用传统 RAG。它使用 Agent 式的 Glob/Grep/Read 模式,更适合代码导航。
- 大多数项目不需要额外的检索基础设施。Claude Code 的原生工具可以有效处理高达约 200K 行代码的代码库。
- RAG 在 Claude Code 工具不擅长的地方表现出色:超大代码库(1M+ 行代码)、多仓库搜索和外部知识库。
- MCP 弥合了差距,连接 Claude Code 和 RAG 系统。在需要时构建 MCP 服务器来暴露语义搜索。
- 长上下文正在胜出对抗代码任务中的 RAG。随着上下文窗口增长,需要 RAG 的门槛在提高。
- 混合策略(CLAUDE.md 中的代码库地图 + 原生工具 + MCP 搜索)是大型项目的实际最优方案。
参见:A06 MCP 协议深入 了解如何实现检索 MCP 服务器,以及 A02 上下文工程 了解如何管理进入 Claude Code 上下文窗口的内容。