Skip to content

A14: RAG vs 长上下文 ​

相关章节:第 7 章 MCP 服务器与外部数据

检索增强生成(Retrieval-Augmented Generation, RAG)和长上下文窗口是两种为 LLM 提供训练数据之外信息的竞争策略。大多数 AI 应用使用其中之一或两者兼用。Claude Code 采用了一种不寻常的方法:它既不使用传统的 RAG,也不使用暴力上下文加载。相反,它实现了一种基于工具的检索策略,结合了两者的优势。本附录解释原因、每种方法何时胜出,以及如何为大型代码库构建混合系统。

传统 RAG 的工作原理 ​

RAG(检索增强生成)是一种在查询时从数据库中获取相关文档并注入到 LLM 提示中的模式:

┌──────────────────────────────────────────────────────────┐
│                   Traditional RAG Pipeline                │
│                                                          │
│  User Query ──► Embedding Model ──► Vector Search ──┐    │
│                                                     │    │
│  Document Store ◄───── Top-K Results ◄──────────────┘    │
│       │                                                  │
│       ▼                                                  │
│  Retrieved Chunks ──► [System Prompt + Chunks + Query]   │
│                              │                           │
│                              ▼                           │
│                        LLM Response                      │
└──────────────────────────────────────────────────────────┘

RAG 流程:

  1. 索引:文档被分割成块并转换为向量嵌入(Vector Embeddings)
  2. 检索:用户查询被嵌入,通过向量相似性搜索找到最近的块
  3. 增强:检索到的块被添加到提示前面
  4. 生成:LLM 使用检索到的上下文生成回复

RAG 的优势:

  • 可以搜索数百万个文档
  • 只加载相关内容(高效的上下文使用)
  • 索引可以增量更新,无需重新处理所有内容
  • 适用于任何 LLM,无论上下文窗口大小

RAG 的弱点:

  • 嵌入质量决定检索质量(语义不匹配很常见)
  • 块边界可能拆分相关信息
  • Top-K 检索可能遗漏重要文档
  • 需要基础设施:向量数据库、嵌入流程、索引管理
  • 无法跨未一起检索的文档进行推理

长上下文的工作原理 ​

长上下文方法更简单:将模型可能需要的所有内容直接加载到上下文窗口中。

┌──────────────────────────────────────────────────────────┐
│                    Long Context Approach                   │
│                                                          │
│  [System Prompt]                                         │
│  [Full Document 1]                                       │
│  [Full Document 2]                                       │
│  [Full Document 3]                                       │
│  [...all potentially relevant documents...]              │
│  [User Query]                                            │
│           │                                              │
│           ▼                                              │
│     LLM Response                                         │
│  (model attends to everything simultaneously)            │
└──────────────────────────────────────────────────────────┘

长上下文的优势:

  • 没有检索错误:模型看到所有内容
  • 可以跨远距离信息进行推理
  • 不需要基础设施:无向量数据库、无嵌入流程
  • 实现和调试更简单

长上下文的弱点:

  • 上下文窗口有限制(截至 2026 年为 200k-1M Token)
  • 成本随上下文大小线性增长
  • "中间遗漏"效应(Lost in the Middle):模型可能低估很长上下文中间部分的信息权重
  • 无法处理真正海量的语料库(数百万文档)

Claude Code 的方法:既非 RAG 也非长上下文 ​

Claude Code 不使用传统 RAG。没有向量数据库,没有嵌入流程,没有块检索。Claude Code 也不会盲目地将整个文件加载到上下文中。相反,它使用基于工具的检索策略:

┌──────────────────────────────────────────────────────────┐
│              Claude Code's Retrieval Strategy              │
│                                                          │
│  User Query ──► Claude reasons about what info it needs  │
│                      │                                   │
│                      ▼                                   │
│              Glob (find files by pattern)                 │
│              Grep (search content by regex)               │
│              Read (load specific file sections)           │
│                      │                                   │
│                      ▼                                   │
│              Claude reviews results                      │
│              Decides if more info is needed               │
│                      │                                   │
│              ┌── YES ─┘──── NO ──┐                       │
│              ▼                   ▼                       │
│         More tool calls    Generate response             │
└──────────────────────────────────────────────────────────┘

这与 RAG 和长上下文都有根本区别:

方面RAG长上下文Claude Code
谁决定检索什么?嵌入相似度(自动化)开发者(手动)Claude(Agent 式)
检索机制向量搜索全部加载Glob/Grep/Read 工具
能否迭代?否(单次检索步骤)否(固定上下文)是(多回合工具循环)
需要的基础设施向量数据库 + 嵌入无文件系统访问
处理代码效果好?差是是

Glob/Grep/Read 模式详解 ​

当 Claude Code 需要在代码库中查找信息时,它使用三步模式:

步骤 1:Glob -- 通过名称/路径模式查找候选文件

Claude thinks: "The user mentioned authentication. Let me find auth-related files."
Tool call: Glob("**/auth*", "**/*auth*")
Result: [src/auth/middleware.ts, src/auth/types.ts, src/auth/service.ts, tests/auth.test.ts]

步骤 2:Grep -- 搜索文件内容以查找特定模式

Claude thinks: "I need to find where tokens are validated."
Tool call: Grep("validateToken|verifyToken|checkToken", "src/auth/")
Result: [src/auth/middleware.ts:42, src/auth/service.ts:118]

步骤 3:Read -- 加载特定的相关部分

Claude thinks: "Let me read the token validation logic."
Tool call: Read("src/auth/middleware.ts", lines 35-60)
Result: [actual code content]

这种模式是一种检索形式,但它是:

  • 语义化的:Claude 基于理解来决定搜索什么,而不是基于嵌入相似度
  • 迭代的:Claude 可以根据初始结果改进搜索
  • 精确的:Claude 读取确切的行,而不是固定大小的块
  • 代码感知的:模式匹配比向量嵌入更适合代码

为什么这对代码比 RAG 效果更好 ​

代码具有使传统 RAG 表现不佳的特性:

  1. 代码是结构化的,不是语义化的:字符串 "authenticate" 可能出现在注释、函数名、import 语句和测试描述中——每个的相关性不同。向量嵌入抹平了这种结构。

  2. 代码有长距离依赖:auth/middleware.ts 中的函数依赖于 auth/types.ts 中的类型,并在 tests/auth.test.ts 中被测试。RAG 独立检索块,可能遗漏这些联系。Claude Code 的迭代工具调用跟踪依赖链。

  3. 代码命名是有信息量的:文件路径(src/auth/middleware.ts)和函数名(validateJWT)对代码导航比向量相似度携带更多信号。Glob 和 Grep 直接利用这一点。

  4. 代码频繁变化:RAG 在代码变化时需要重新索引。Claude Code 的工具直接读取文件系统,所以它们总是看到当前状态。

何时 RAG 更好 ​

Claude Code 的方法有局限性。在以下情况下 RAG 是更好的选择:

极大型代码库(1M+ 行代码) ​

如果你的代码库超出了通过合理数量的 Glob/Grep/Read 工具调用可以导航的范围,RAG 提供更快的初始检索:

Codebase size    Claude Code approach    RAG approach
──────────────   ────────────────────    ─────────────
10K LOC          Excellent               Overkill
50K LOC          Excellent               Unnecessary
200K LOC         Good                    Comparable
500K LOC         Workable (many calls)   Better initial retrieval
1M+ LOC          Slow (too many calls)   Significantly faster

外部知识库 ​

对于代码库之外的信息——文档、Wiki 页面、Confluence 文章、Slack 历史——Claude Code 没有原生访问能力。RAG(或基于 MCP 的检索)弥补了这个差距:

bash
# MCP server providing RAG over company documentation
# .claude/settings.json
{
  "mcpServers": {
    "company-docs": {
      "command": "node",
      "args": ["./mcp-servers/docs-rag/server.js"],
      "env": {
        "VECTOR_DB_URL": "http://localhost:6333",
        "COLLECTION": "company_docs"
      }
    }
  }
}

多仓库搜索 ​

Claude Code 在单个仓库内运行。对于跨仓库搜索(例如,"找到所有依赖我们认证库的服务"),索引多个仓库的 RAG 系统更实用。

MCP 作为检索层 ​

MCP(Model Context Protocol,模型上下文协议)连接了 Claude Code 和外部检索系统。你可以构建实现 RAG 并将其作为工具暴露给 Claude Code 调用的 MCP 服务器:

Claude Code ──► MCP Server ──► Vector Database ──► Retrieved Context
                   │
                   ├── search_docs(query) → relevant documentation
                   ├── search_code(query) → code from other repos
                   └── search_tickets(query) → related JIRA issues

提供语义代码搜索的 MCP 服务器示例:

typescript
// mcp-servers/code-search/server.ts
import { McpServer } from "@modelcontextprotocol/sdk/server";

const server = new McpServer({
  name: "code-search",
  version: "1.0.0",
});

server.tool(
  "semantic_search",
  "Search the codebase using semantic similarity",
  {
    query: { type: "string", description: "Natural language query" },
    top_k: { type: "number", description: "Number of results", default: 5 },
  },
  async ({ query, top_k }) => {
    // 1. Embed the query
    const embedding = await embedModel.embed(query);

    // 2. Search the vector database
    const results = await vectorDB.search({
      collection: "codebase",
      vector: embedding,
      limit: top_k,
    });

    // 3. Return formatted results
    return {
      content: results.map((r) => ({
        type: "text",
        text: `File: ${r.metadata.file}:${r.metadata.line}\n${r.payload.code}`,
      })),
    };
  }
);

然后 Claude Code 可以将此工具与其原生的 Glob/Grep/Read 工具一起使用,结合语义搜索和结构化导航。

大型 Monorepo 的混合策略 ​

对于大型 Monorepo(500K+ 行代码),最优方案是将 Claude Code 的原生工具与 RAG 结合:

策略 1:用 RAG 发现,用工具深入 ​

"Find the rate limiter" ──► RAG (fast semantic search across 1M LOC)
                                    │
                                    ▼
                            "src/middleware/rate-limit.ts"
                                    │
                                    ▼
"Understand and modify it" ──► Read/Grep/Edit (precise tool-based work)

使用 RAG 缩小在哪里查找的范围,然后切换到 Claude Code 的原生工具来理解和修改代码。这结合了 RAG 的广度与基于工具的检索的精度。

策略 2:预构建的代码地图 ​

不使用 RAG,而是预先计算代码库的高级地图并将其包含在 CLAUDE.md 中:

markdown
<!-- CLAUDE.md -->
## Codebase Map

### Core Services (src/services/)
- auth/: JWT authentication, token management (3K LOC)
- payments/: Stripe integration, billing (5K LOC)
- notifications/: Email, SMS, push notifications (2K LOC)

### API Layer (src/api/)
- routes/: Express route handlers
- middleware/: Auth, rate limiting, logging
- validators/: Zod schemas for request validation

### Shared (src/shared/)
- types/: TypeScript interfaces shared across services
- utils/: Helper functions (date formatting, string manipulation)
- constants/: App-wide configuration constants

这相当于一个目录。Claude Code 读取地图,识别相关区域,然后使用 Glob/Grep/Read 深入具体细节。不需要向量数据库。

策略 3:分层检索 ​

Tier 1: CLAUDE.md codebase map (always in context)
    │
    ▼ (Claude identifies area of interest)
Tier 2: Glob/Grep to find specific files
    │
    ▼ (Claude identifies files to read)
Tier 3: Read specific file sections
    │
    ▼ (if needed, Claude searches further)
Tier 4: MCP semantic search for cross-cutting concerns

每一层逐步更昂贵但更精确。大多数任务在 Tier 2-3 就能解决。只有横切关注点的任务(例如,"找到所有处理错误的地方")才受益于 Tier 4。

比较表 ​

维度传统 RAG长上下文Claude Code (Glob/Grep/Read)混合 (Claude Code + MCP RAG)
最大代码库无限~200K Token(~5K 行)~200K 行(有效)无限
检索质量取决于嵌入完美(看到所有)高(语义工具使用)最高
延迟快(单次数据库查询)慢(大提示)中(多回合工具)中
每次查询成本低(小上下文)高(大上下文)中(迭代读取)中高
设置成本高(向量数据库、索引)无无中(MCP 服务器)
处理代码效果好?差是是是
迭代改进?否否是是
跨仓库搜索?是否否是
始终最新?否(需要重新索引)是是部分

上下文窗口何时让 RAG 过时 ​

随着上下文窗口增长(2024 年 200K,2025 年 1M,2026 年以后可能更大),RAG 和长上下文之间的收支平衡点在移动:

RAG advantage zone
│
│  xxxxxxxx
│  xxxxxxxxx
│  xxxxxxxxxx          
│  xxxxxxxxxxx         Long context advantage zone
│  xxxxxxxxxxxx        
│  xxxxxxxxxxxxx         oooooo
│  xxxxxxxxxxxxxx        oooooooo
│  xxxxxxxxxxxxxxx       oooooooooo
│  xxxxxxxxxxxxxxxx      oooooooooooo
│  xxxxxxxxxxxxxxxxx     ooooooooooooo
└──────────────────────────────────────── Context window size
   32K   128K  200K  500K   1M    4M    16M

今天(1M 上下文):RAG 对于超过约 200K 行代码的代码库或外部知识库仍然必要。对于大多数单仓库项目,Claude Code 的原生工具就够了。

未来(4M+ 上下文):大多数单仓库项目将完全放入上下文。RAG 将仅在真正海量的代码库、多仓库搜索和外部知识集成中保持相关性。

趋势很明确:对于代码任务,长上下文正在蚕食 RAG 的地盘。但 RAG 永远不会完全消失,因为外部知识库和多源搜索本质上是检索问题。

实际建议 ​

中小型项目(< 100K 行代码) ​

使用 Claude Code 的原生工具。不需要 RAG 基础设施:

bash
# Just use Claude Code normally -- it will find what it needs
claude "Fix the authentication bug in the middleware layer"

大型项目(100K - 500K 行代码) ​

在 CLAUDE.md 中添加代码库地图,并考虑使用 Opus 的 1M 上下文窗口:

bash
# Use Opus for tasks spanning the full codebase
claude --model claude-opus-4-7-20250219 \
  "Audit all API endpoints for consistent error handling"

超大型项目(500K+ 行代码)或多仓库 ​

构建提供语义搜索的 MCP 服务器:

bash
# Set up MCP-based code search
# Then Claude Code can use both native tools AND semantic search
claude "Find all services that depend on the user authentication module"
# Claude will use both Grep and the MCP semantic_search tool

外部知识(文档、Wiki、工单) ​

通过 MCP 的 RAG 是正确的方法——这些信息不在文件系统上:

bash
# MCP server wrapping your documentation system
# Claude can search docs alongside code
claude "How does the billing webhook work? Check both the code and the design docs."

核心要点 ​

  1. Claude Code 不使用传统 RAG。它使用 Agent 式的 Glob/Grep/Read 模式,更适合代码导航。
  2. 大多数项目不需要额外的检索基础设施。Claude Code 的原生工具可以有效处理高达约 200K 行代码的代码库。
  3. RAG 在 Claude Code 工具不擅长的地方表现出色:超大代码库(1M+ 行代码)、多仓库搜索和外部知识库。
  4. MCP 弥合了差距,连接 Claude Code 和 RAG 系统。在需要时构建 MCP 服务器来暴露语义搜索。
  5. 长上下文正在胜出对抗代码任务中的 RAG。随着上下文窗口增长,需要 RAG 的门槛在提高。
  6. 混合策略(CLAUDE.md 中的代码库地图 + 原生工具 + MCP 搜索)是大型项目的实际最优方案。

参见:A06 MCP 协议深入 了解如何实现检索 MCP 服务器,以及 A02 上下文工程 了解如何管理进入 Claude Code 上下文窗口的内容。

基于 MIT 许可发布