Skip to content

A14: RAG vs Long Context ​

Related chapters: Ch7 MCP Servers & External Data

Retrieval-Augmented Generation (RAG) and long context windows are two competing strategies for giving an LLM access to information beyond its training data. Most AI applications use one or both. Claude Code takes an unusual approach: it uses neither traditional RAG nor brute-force context loading. Instead, it implements a tool-based retrieval strategy that combines the strengths of both. This appendix explains why, when each approach wins, and how to build hybrid systems for massive codebases.

How Traditional RAG Works ​

RAG (Retrieval-Augmented Generation) is a pattern where relevant documents are fetched from a database and injected into the LLM's prompt at query time:

┌──────────────────────────────────────────────────────────┐
│                   Traditional RAG Pipeline                │
│                                                          │
│  User Query ──► Embedding Model ──► Vector Search ──┐    │
│                                                     │    │
│  Document Store ◄───── Top-K Results ◄──────────────┘    │
│       │                                                  │
│       ▼                                                  │
│  Retrieved Chunks ──► [System Prompt + Chunks + Query]   │
│                              │                           │
│                              ▼                           │
│                        LLM Response                      │
└──────────────────────────────────────────────────────────┘

The RAG pipeline:

  1. Indexing: Documents are split into chunks and converted to vector embeddings
  2. Retrieval: The user's query is embedded, and the nearest chunks are found via vector similarity search
  3. Augmentation: Retrieved chunks are prepended to the prompt
  4. Generation: The LLM generates a response using the retrieved context

Strengths of RAG:

  • Can search across millions of documents
  • Only loads relevant content (efficient context usage)
  • Index can be updated incrementally without reprocessing everything
  • Works with any LLM, regardless of context window size

Weaknesses of RAG:

  • Embedding quality determines retrieval quality (semantic mismatch is common)
  • Chunk boundaries can split relevant information
  • Top-K retrieval may miss important documents
  • Requires infrastructure: vector database, embedding pipeline, index management
  • Cannot reason across documents that were not retrieved together

How Long Context Works ​

The long context approach is simpler: load everything the model might need directly into the context window.

┌──────────────────────────────────────────────────────────┐
│                    Long Context Approach                   │
│                                                          │
│  [System Prompt]                                         │
│  [Full Document 1]                                       │
│  [Full Document 2]                                       │
│  [Full Document 3]                                       │
│  [...all potentially relevant documents...]              │
│  [User Query]                                            │
│           │                                              │
│           ▼                                              │
│     LLM Response                                         │
│  (model attends to everything simultaneously)            │
└──────────────────────────────────────────────────────────┘

Strengths of long context:

  • No retrieval errors: the model sees everything
  • Can reason across distant pieces of information
  • No infrastructure required: no vector database, no embedding pipeline
  • Simpler to implement and debug

Weaknesses of long context:

  • Context windows have limits (200k-1M tokens as of 2026)
  • Cost scales linearly with context size
  • "Lost in the middle" effect: models may underweight information in the middle of very long contexts
  • Cannot handle truly massive corpora (millions of documents)

Claude Code's Approach: Neither RAG Nor Long Context ​

Claude Code does not use traditional RAG. There is no vector database, no embedding pipeline, no chunk retrieval. Claude Code also does not blindly load entire files into context. Instead, it uses a tool-based retrieval strategy:

┌──────────────────────────────────────────────────────────┐
│              Claude Code's Retrieval Strategy              │
│                                                          │
│  User Query ──► Claude reasons about what info it needs  │
│                      │                                   │
│                      ▼                                   │
│              Glob (find files by pattern)                 │
│              Grep (search content by regex)               │
│              Read (load specific file sections)           │
│                      │                                   │
│                      ▼                                   │
│              Claude reviews results                      │
│              Decides if more info is needed               │
│                      │                                   │
│              ┌── YES ─┘──── NO ──┐                       │
│              ▼                   ▼                       │
│         More tool calls    Generate response             │
└──────────────────────────────────────────────────────────┘

This is fundamentally different from both RAG and long context:

AspectRAGLong ContextClaude Code
Who decides what to retrieve?Embedding similarity (automated)Developer (manual)Claude (agentic)
Retrieval mechanismVector searchLoad everythingGlob/Grep/Read tools
Can it iterate?No (single retrieval step)No (fixed context)Yes (multi-turn tool loop)
Infrastructure neededVector DB + embeddingsNoneFilesystem access
Handles code well?Poorly (code embeddings are unreliable)YesYes

The Glob/Grep/Read Pattern in Detail ​

When Claude Code needs to find information in your codebase, it uses a three-step pattern:

Step 1: Glob -- Find candidate files by name/path pattern

Claude thinks: "The user mentioned authentication. Let me find auth-related files."
Tool call: Glob("**/auth*", "**/*auth*")
Result: [src/auth/middleware.ts, src/auth/types.ts, src/auth/service.ts, tests/auth.test.ts]

Step 2: Grep -- Search file contents for specific patterns

Claude thinks: "I need to find where tokens are validated."
Tool call: Grep("validateToken|verifyToken|checkToken", "src/auth/")
Result: [src/auth/middleware.ts:42, src/auth/service.ts:118]

Step 3: Read -- Load the specific relevant sections

Claude thinks: "Let me read the token validation logic."
Tool call: Read("src/auth/middleware.ts", lines 35-60)
Result: [actual code content]

This pattern is a form of retrieval, but it is:

  • Semantic: Claude decides what to search for based on understanding, not embedding similarity
  • Iterative: Claude can refine its search based on initial results
  • Precise: Claude reads exact lines, not fixed-size chunks
  • Code-aware: Pattern matching works better for code than vector embeddings

Why This Works Better Than RAG for Code ​

Code has properties that make traditional RAG perform poorly:

  1. Code is structural, not semantic: The string "authenticate" might appear in comments, function names, import statements, and test descriptions -- all with different relevance. Vector embeddings flatten this structure.

  2. Code has long-range dependencies: A function in auth/middleware.ts depends on types in auth/types.ts and is tested in tests/auth.test.ts. RAG retrieves chunks independently and may miss these connections. Claude Code's iterative tool calls follow the dependency chain.

  3. Code naming is informative: File paths (src/auth/middleware.ts) and function names (validateJWT) carry more signal for code navigation than vector similarity. Glob and Grep exploit this directly.

  4. Code changes frequently: RAG requires re-indexing when code changes. Claude Code's tools read the filesystem directly, so they always see the current state.

When RAG Would Be Better ​

Claude Code's approach has limits. RAG becomes the better choice when:

Extremely Large Codebases (1M+ Lines of Code) ​

If your codebase exceeds what can be navigated through Glob/Grep/Read in a reasonable number of tool calls, RAG provides faster initial retrieval:

Codebase size    Claude Code approach    RAG approach
──────────────   ────────────────────    ─────────────
10K LOC          Excellent               Overkill
50K LOC          Excellent               Unnecessary
200K LOC         Good                    Comparable
500K LOC         Workable (many calls)   Better initial retrieval
1M+ LOC          Slow (too many calls)   Significantly faster

External Knowledge Bases ​

For information outside the codebase -- documentation, wiki pages, Confluence articles, Slack history -- Claude Code has no native access. RAG (or MCP-based retrieval) bridges this gap:

bash
# MCP server providing RAG over company documentation
# .claude/settings.json
{
  "mcpServers": {
    "company-docs": {
      "command": "node",
      "args": ["./mcp-servers/docs-rag/server.js"],
      "env": {
        "VECTOR_DB_URL": "http://localhost:6333",
        "COLLECTION": "company_docs"
      }
    }
  }
}

Claude Code operates within a single repository. For cross-repo search (e.g., "find all services that depend on our auth library"), a RAG system indexing multiple repositories is more practical.

MCP as a Retrieval Layer ​

MCP (Model Context Protocol) bridges Claude Code and external retrieval systems. You can build MCP servers that implement RAG and expose it as tools Claude Code can call:

Claude Code ──► MCP Server ──► Vector Database ──► Retrieved Context
                   │
                   ├── search_docs(query) → relevant documentation
                   ├── search_code(query) → code from other repos
                   └── search_tickets(query) → related JIRA issues

Example MCP server providing semantic code search:

typescript
// mcp-servers/code-search/server.ts
import { McpServer } from "@modelcontextprotocol/sdk/server";

const server = new McpServer({
  name: "code-search",
  version: "1.0.0",
});

server.tool(
  "semantic_search",
  "Search the codebase using semantic similarity",
  {
    query: { type: "string", description: "Natural language query" },
    top_k: { type: "number", description: "Number of results", default: 5 },
  },
  async ({ query, top_k }) => {
    // 1. Embed the query
    const embedding = await embedModel.embed(query);

    // 2. Search the vector database
    const results = await vectorDB.search({
      collection: "codebase",
      vector: embedding,
      limit: top_k,
    });

    // 3. Return formatted results
    return {
      content: results.map((r) => ({
        type: "text",
        text: `File: ${r.metadata.file}:${r.metadata.line}\n${r.payload.code}`,
      })),
    };
  }
);

Claude Code then uses this tool alongside its native Glob/Grep/Read tools, combining semantic search with structural navigation.

Hybrid Strategies for Massive Monorepos ​

For large monorepos (500K+ LOC), the optimal approach combines Claude Code's native tools with RAG:

Strategy 1: RAG for Discovery, Tools for Detail ​

"Find the rate limiter" ──► RAG (fast semantic search across 1M LOC)
                                    │
                                    ▼
                            "src/middleware/rate-limit.ts"
                                    │
                                    ▼
"Understand and modify it" ──► Read/Grep/Edit (precise tool-based work)

Use RAG to narrow down where to look, then switch to Claude Code's native tools to understand and modify the code. This combines RAG's breadth with tool-based retrieval's precision.

Strategy 2: Pre-Built Context Maps ​

Instead of RAG, pre-compute a high-level map of the codebase and include it in CLAUDE.md:

markdown
<!-- CLAUDE.md -->
## Codebase Map

### Core Services (src/services/)
- auth/: JWT authentication, token management (3K LOC)
- payments/: Stripe integration, billing (5K LOC)
- notifications/: Email, SMS, push notifications (2K LOC)

### API Layer (src/api/)
- routes/: Express route handlers
- middleware/: Auth, rate limiting, logging
- validators/: Zod schemas for request validation

### Shared (src/shared/)
- types/: TypeScript interfaces shared across services
- utils/: Helper functions (date formatting, string manipulation)
- constants/: App-wide configuration constants

This acts like a table of contents. Claude Code reads the map, identifies relevant areas, then uses Glob/Grep/Read to dive into the specifics. No vector database required.

Strategy 3: Tiered Retrieval ​

Tier 1: CLAUDE.md codebase map (always in context)
    │
    ▼ (Claude identifies area of interest)
Tier 2: Glob/Grep to find specific files
    │
    ▼ (Claude identifies files to read)
Tier 3: Read specific file sections
    │
    ▼ (if needed, Claude searches further)
Tier 4: MCP semantic search for cross-cutting concerns

Each tier is progressively more expensive but more precise. Most tasks are resolved at Tier 2-3. Only cross-cutting tasks (e.g., "find all places we handle errors") benefit from Tier 4.

Comparison Table ​

DimensionTraditional RAGLong ContextClaude Code (Glob/Grep/Read)Hybrid (Claude Code + MCP RAG)
Max codebaseUnlimited~200K tokens (~5K LOC)~200K LOC effectivelyUnlimited
Retrieval qualityDepends on embeddingsPerfect (sees everything)High (semantic tool use)Highest
LatencyFast (single DB query)Slow (large prompt)Medium (multi-turn tools)Medium
Cost per queryLow (small context)High (large context)Medium (iterative reads)Medium-High
Setup costHigh (vector DB, indexing)NoneNoneMedium (MCP server)
Handles code well?PoorlyYesYesYes
Iterative refinement?NoNoYesYes
Cross-repo search?YesNoNoYes
Always up-to-date?No (requires re-indexing)YesYesPartially

When Context Windows Make RAG Obsolete ​

As context windows grow (200K in 2024, 1M in 2025, potentially larger in 2026+), the breakeven point between RAG and long context shifts:

RAG advantage zone
│
│  xxxxxxxx
│  xxxxxxxxx
│  xxxxxxxxxx          
│  xxxxxxxxxxx         Long context advantage zone
│  xxxxxxxxxxxx        
│  xxxxxxxxxxxxx         oooooo
│  xxxxxxxxxxxxxx        oooooooo
│  xxxxxxxxxxxxxxx       oooooooooo
│  xxxxxxxxxxxxxxxx      oooooooooooo
│  xxxxxxxxxxxxxxxxx     ooooooooooooo
└──────────────────────────────────────── Context window size
   32K   128K  200K  500K   1M    4M    16M

Today (1M context): RAG is necessary for codebases over ~200K LOC or for external knowledge bases. For most single-repo projects, Claude Code's native tools are sufficient.

Future (4M+ context): Most single-repo projects will fit entirely in context. RAG will remain relevant only for truly massive codebases, multi-repo search, and external knowledge integration.

The trend is clear: long context is eating RAG's lunch for code tasks. But RAG will never fully disappear because external knowledge bases and multi-source search are inherently retrieval problems.

Practical Recommendations ​

For Small to Medium Projects (< 100K LOC) ​

Use Claude Code's native tools. No RAG infrastructure needed:

bash
# Just use Claude Code normally -- it will find what it needs
claude "Fix the authentication bug in the middleware layer"

For Large Projects (100K - 500K LOC) ​

Add a codebase map to CLAUDE.md and consider using Opus for its 1M context window:

bash
# Use Opus for tasks spanning the full codebase
claude --model claude-opus-4-7-20250219 \
  "Audit all API endpoints for consistent error handling"

For Massive Projects (500K+ LOC) or Multi-Repo ​

Build an MCP server providing semantic search:

bash
# Set up MCP-based code search
# Then Claude Code can use both native tools AND semantic search
claude "Find all services that depend on the user authentication module"
# Claude will use both Grep and the MCP semantic_search tool

For External Knowledge (Docs, Wikis, Tickets) ​

RAG via MCP is the right approach -- this information is not on the filesystem:

bash
# MCP server wrapping your documentation system
# Claude can search docs alongside code
claude "How does the billing webhook work? Check both the code and the design docs."

Key Takeaways ​

  1. Claude Code does not use traditional RAG. It uses an agentic Glob/Grep/Read pattern that is better suited for code navigation.
  2. For most projects, no additional retrieval infrastructure is needed. Claude Code's native tools handle codebases up to ~200K LOC effectively.
  3. RAG excels where Claude Code's tools do not: massive codebases (1M+ LOC), multi-repo search, and external knowledge bases.
  4. MCP bridges the gap between Claude Code and RAG systems. Build MCP servers to expose semantic search when needed.
  5. Long context is winning against RAG for code tasks. As context windows grow, the threshold for needing RAG increases.
  6. Hybrid strategies (codebase map in CLAUDE.md + native tools + MCP search) are the practical optimum for large projects.

See also: A06 MCP Protocol Deep Dive for implementing retrieval MCP servers, and A02 Context Engineering for managing what goes into Claude Code's context window.

Released under MIT License