A14: RAG vs Long Context
Related chapters: Ch7 MCP Servers & External Data
Retrieval-Augmented Generation (RAG) and long context windows are two competing strategies for giving an LLM access to information beyond its training data. Most AI applications use one or both. Claude Code takes an unusual approach: it uses neither traditional RAG nor brute-force context loading. Instead, it implements a tool-based retrieval strategy that combines the strengths of both. This appendix explains why, when each approach wins, and how to build hybrid systems for massive codebases.
How Traditional RAG Works
RAG (Retrieval-Augmented Generation) is a pattern where relevant documents are fetched from a database and injected into the LLM's prompt at query time:
┌──────────────────────────────────────────────────────────┐
│ Traditional RAG Pipeline │
│ │
│ User Query ──► Embedding Model ──► Vector Search ──┐ │
│ │ │
│ Document Store ◄───── Top-K Results ◄──────────────┘ │
│ │ │
│ ▼ │
│ Retrieved Chunks ──► [System Prompt + Chunks + Query] │
│ │ │
│ ▼ │
│ LLM Response │
└──────────────────────────────────────────────────────────┘The RAG pipeline:
- Indexing: Documents are split into chunks and converted to vector embeddings
- Retrieval: The user's query is embedded, and the nearest chunks are found via vector similarity search
- Augmentation: Retrieved chunks are prepended to the prompt
- Generation: The LLM generates a response using the retrieved context
Strengths of RAG:
- Can search across millions of documents
- Only loads relevant content (efficient context usage)
- Index can be updated incrementally without reprocessing everything
- Works with any LLM, regardless of context window size
Weaknesses of RAG:
- Embedding quality determines retrieval quality (semantic mismatch is common)
- Chunk boundaries can split relevant information
- Top-K retrieval may miss important documents
- Requires infrastructure: vector database, embedding pipeline, index management
- Cannot reason across documents that were not retrieved together
How Long Context Works
The long context approach is simpler: load everything the model might need directly into the context window.
┌──────────────────────────────────────────────────────────┐
│ Long Context Approach │
│ │
│ [System Prompt] │
│ [Full Document 1] │
│ [Full Document 2] │
│ [Full Document 3] │
│ [...all potentially relevant documents...] │
│ [User Query] │
│ │ │
│ ▼ │
│ LLM Response │
│ (model attends to everything simultaneously) │
└──────────────────────────────────────────────────────────┘Strengths of long context:
- No retrieval errors: the model sees everything
- Can reason across distant pieces of information
- No infrastructure required: no vector database, no embedding pipeline
- Simpler to implement and debug
Weaknesses of long context:
- Context windows have limits (200k-1M tokens as of 2026)
- Cost scales linearly with context size
- "Lost in the middle" effect: models may underweight information in the middle of very long contexts
- Cannot handle truly massive corpora (millions of documents)
Claude Code's Approach: Neither RAG Nor Long Context
Claude Code does not use traditional RAG. There is no vector database, no embedding pipeline, no chunk retrieval. Claude Code also does not blindly load entire files into context. Instead, it uses a tool-based retrieval strategy:
┌──────────────────────────────────────────────────────────┐
│ Claude Code's Retrieval Strategy │
│ │
│ User Query ──► Claude reasons about what info it needs │
│ │ │
│ ▼ │
│ Glob (find files by pattern) │
│ Grep (search content by regex) │
│ Read (load specific file sections) │
│ │ │
│ ▼ │
│ Claude reviews results │
│ Decides if more info is needed │
│ │ │
│ ┌── YES ─┘──── NO ──┐ │
│ ▼ ▼ │
│ More tool calls Generate response │
└──────────────────────────────────────────────────────────┘This is fundamentally different from both RAG and long context:
| Aspect | RAG | Long Context | Claude Code |
|---|---|---|---|
| Who decides what to retrieve? | Embedding similarity (automated) | Developer (manual) | Claude (agentic) |
| Retrieval mechanism | Vector search | Load everything | Glob/Grep/Read tools |
| Can it iterate? | No (single retrieval step) | No (fixed context) | Yes (multi-turn tool loop) |
| Infrastructure needed | Vector DB + embeddings | None | Filesystem access |
| Handles code well? | Poorly (code embeddings are unreliable) | Yes | Yes |
The Glob/Grep/Read Pattern in Detail
When Claude Code needs to find information in your codebase, it uses a three-step pattern:
Step 1: Glob -- Find candidate files by name/path pattern
Claude thinks: "The user mentioned authentication. Let me find auth-related files."
Tool call: Glob("**/auth*", "**/*auth*")
Result: [src/auth/middleware.ts, src/auth/types.ts, src/auth/service.ts, tests/auth.test.ts]Step 2: Grep -- Search file contents for specific patterns
Claude thinks: "I need to find where tokens are validated."
Tool call: Grep("validateToken|verifyToken|checkToken", "src/auth/")
Result: [src/auth/middleware.ts:42, src/auth/service.ts:118]Step 3: Read -- Load the specific relevant sections
Claude thinks: "Let me read the token validation logic."
Tool call: Read("src/auth/middleware.ts", lines 35-60)
Result: [actual code content]This pattern is a form of retrieval, but it is:
- Semantic: Claude decides what to search for based on understanding, not embedding similarity
- Iterative: Claude can refine its search based on initial results
- Precise: Claude reads exact lines, not fixed-size chunks
- Code-aware: Pattern matching works better for code than vector embeddings
Why This Works Better Than RAG for Code
Code has properties that make traditional RAG perform poorly:
Code is structural, not semantic: The string "authenticate" might appear in comments, function names, import statements, and test descriptions -- all with different relevance. Vector embeddings flatten this structure.
Code has long-range dependencies: A function in
auth/middleware.tsdepends on types inauth/types.tsand is tested intests/auth.test.ts. RAG retrieves chunks independently and may miss these connections. Claude Code's iterative tool calls follow the dependency chain.Code naming is informative: File paths (
src/auth/middleware.ts) and function names (validateJWT) carry more signal for code navigation than vector similarity. Glob and Grep exploit this directly.Code changes frequently: RAG requires re-indexing when code changes. Claude Code's tools read the filesystem directly, so they always see the current state.
When RAG Would Be Better
Claude Code's approach has limits. RAG becomes the better choice when:
Extremely Large Codebases (1M+ Lines of Code)
If your codebase exceeds what can be navigated through Glob/Grep/Read in a reasonable number of tool calls, RAG provides faster initial retrieval:
Codebase size Claude Code approach RAG approach
────────────── ──────────────────── ─────────────
10K LOC Excellent Overkill
50K LOC Excellent Unnecessary
200K LOC Good Comparable
500K LOC Workable (many calls) Better initial retrieval
1M+ LOC Slow (too many calls) Significantly fasterExternal Knowledge Bases
For information outside the codebase -- documentation, wiki pages, Confluence articles, Slack history -- Claude Code has no native access. RAG (or MCP-based retrieval) bridges this gap:
# MCP server providing RAG over company documentation
# .claude/settings.json
{
"mcpServers": {
"company-docs": {
"command": "node",
"args": ["./mcp-servers/docs-rag/server.js"],
"env": {
"VECTOR_DB_URL": "http://localhost:6333",
"COLLECTION": "company_docs"
}
}
}
}Multi-Repository Search
Claude Code operates within a single repository. For cross-repo search (e.g., "find all services that depend on our auth library"), a RAG system indexing multiple repositories is more practical.
MCP as a Retrieval Layer
MCP (Model Context Protocol) bridges Claude Code and external retrieval systems. You can build MCP servers that implement RAG and expose it as tools Claude Code can call:
Claude Code ──► MCP Server ──► Vector Database ──► Retrieved Context
│
├── search_docs(query) → relevant documentation
├── search_code(query) → code from other repos
└── search_tickets(query) → related JIRA issuesExample MCP server providing semantic code search:
// mcp-servers/code-search/server.ts
import { McpServer } from "@modelcontextprotocol/sdk/server";
const server = new McpServer({
name: "code-search",
version: "1.0.0",
});
server.tool(
"semantic_search",
"Search the codebase using semantic similarity",
{
query: { type: "string", description: "Natural language query" },
top_k: { type: "number", description: "Number of results", default: 5 },
},
async ({ query, top_k }) => {
// 1. Embed the query
const embedding = await embedModel.embed(query);
// 2. Search the vector database
const results = await vectorDB.search({
collection: "codebase",
vector: embedding,
limit: top_k,
});
// 3. Return formatted results
return {
content: results.map((r) => ({
type: "text",
text: `File: ${r.metadata.file}:${r.metadata.line}\n${r.payload.code}`,
})),
};
}
);Claude Code then uses this tool alongside its native Glob/Grep/Read tools, combining semantic search with structural navigation.
Hybrid Strategies for Massive Monorepos
For large monorepos (500K+ LOC), the optimal approach combines Claude Code's native tools with RAG:
Strategy 1: RAG for Discovery, Tools for Detail
"Find the rate limiter" ──► RAG (fast semantic search across 1M LOC)
│
▼
"src/middleware/rate-limit.ts"
│
▼
"Understand and modify it" ──► Read/Grep/Edit (precise tool-based work)Use RAG to narrow down where to look, then switch to Claude Code's native tools to understand and modify the code. This combines RAG's breadth with tool-based retrieval's precision.
Strategy 2: Pre-Built Context Maps
Instead of RAG, pre-compute a high-level map of the codebase and include it in CLAUDE.md:
<!-- CLAUDE.md -->
## Codebase Map
### Core Services (src/services/)
- auth/: JWT authentication, token management (3K LOC)
- payments/: Stripe integration, billing (5K LOC)
- notifications/: Email, SMS, push notifications (2K LOC)
### API Layer (src/api/)
- routes/: Express route handlers
- middleware/: Auth, rate limiting, logging
- validators/: Zod schemas for request validation
### Shared (src/shared/)
- types/: TypeScript interfaces shared across services
- utils/: Helper functions (date formatting, string manipulation)
- constants/: App-wide configuration constantsThis acts like a table of contents. Claude Code reads the map, identifies relevant areas, then uses Glob/Grep/Read to dive into the specifics. No vector database required.
Strategy 3: Tiered Retrieval
Tier 1: CLAUDE.md codebase map (always in context)
│
▼ (Claude identifies area of interest)
Tier 2: Glob/Grep to find specific files
│
▼ (Claude identifies files to read)
Tier 3: Read specific file sections
│
▼ (if needed, Claude searches further)
Tier 4: MCP semantic search for cross-cutting concernsEach tier is progressively more expensive but more precise. Most tasks are resolved at Tier 2-3. Only cross-cutting tasks (e.g., "find all places we handle errors") benefit from Tier 4.
Comparison Table
| Dimension | Traditional RAG | Long Context | Claude Code (Glob/Grep/Read) | Hybrid (Claude Code + MCP RAG) |
|---|---|---|---|---|
| Max codebase | Unlimited | ~200K tokens (~5K LOC) | ~200K LOC effectively | Unlimited |
| Retrieval quality | Depends on embeddings | Perfect (sees everything) | High (semantic tool use) | Highest |
| Latency | Fast (single DB query) | Slow (large prompt) | Medium (multi-turn tools) | Medium |
| Cost per query | Low (small context) | High (large context) | Medium (iterative reads) | Medium-High |
| Setup cost | High (vector DB, indexing) | None | None | Medium (MCP server) |
| Handles code well? | Poorly | Yes | Yes | Yes |
| Iterative refinement? | No | No | Yes | Yes |
| Cross-repo search? | Yes | No | No | Yes |
| Always up-to-date? | No (requires re-indexing) | Yes | Yes | Partially |
When Context Windows Make RAG Obsolete
As context windows grow (200K in 2024, 1M in 2025, potentially larger in 2026+), the breakeven point between RAG and long context shifts:
RAG advantage zone
│
│ xxxxxxxx
│ xxxxxxxxx
│ xxxxxxxxxx
│ xxxxxxxxxxx Long context advantage zone
│ xxxxxxxxxxxx
│ xxxxxxxxxxxxx oooooo
│ xxxxxxxxxxxxxx oooooooo
│ xxxxxxxxxxxxxxx oooooooooo
│ xxxxxxxxxxxxxxxx oooooooooooo
│ xxxxxxxxxxxxxxxxx ooooooooooooo
└──────────────────────────────────────── Context window size
32K 128K 200K 500K 1M 4M 16MToday (1M context): RAG is necessary for codebases over ~200K LOC or for external knowledge bases. For most single-repo projects, Claude Code's native tools are sufficient.
Future (4M+ context): Most single-repo projects will fit entirely in context. RAG will remain relevant only for truly massive codebases, multi-repo search, and external knowledge integration.
The trend is clear: long context is eating RAG's lunch for code tasks. But RAG will never fully disappear because external knowledge bases and multi-source search are inherently retrieval problems.
Practical Recommendations
For Small to Medium Projects (< 100K LOC)
Use Claude Code's native tools. No RAG infrastructure needed:
# Just use Claude Code normally -- it will find what it needs
claude "Fix the authentication bug in the middleware layer"For Large Projects (100K - 500K LOC)
Add a codebase map to CLAUDE.md and consider using Opus for its 1M context window:
# Use Opus for tasks spanning the full codebase
claude --model claude-opus-4-7-20250219 \
"Audit all API endpoints for consistent error handling"For Massive Projects (500K+ LOC) or Multi-Repo
Build an MCP server providing semantic search:
# Set up MCP-based code search
# Then Claude Code can use both native tools AND semantic search
claude "Find all services that depend on the user authentication module"
# Claude will use both Grep and the MCP semantic_search toolFor External Knowledge (Docs, Wikis, Tickets)
RAG via MCP is the right approach -- this information is not on the filesystem:
# MCP server wrapping your documentation system
# Claude can search docs alongside code
claude "How does the billing webhook work? Check both the code and the design docs."Key Takeaways
- Claude Code does not use traditional RAG. It uses an agentic Glob/Grep/Read pattern that is better suited for code navigation.
- For most projects, no additional retrieval infrastructure is needed. Claude Code's native tools handle codebases up to ~200K LOC effectively.
- RAG excels where Claude Code's tools do not: massive codebases (1M+ LOC), multi-repo search, and external knowledge bases.
- MCP bridges the gap between Claude Code and RAG systems. Build MCP servers to expose semantic search when needed.
- Long context is winning against RAG for code tasks. As context windows grow, the threshold for needing RAG increases.
- Hybrid strategies (codebase map in CLAUDE.md + native tools + MCP search) are the practical optimum for large projects.
See also: A06 MCP Protocol Deep Dive for implementing retrieval MCP servers, and A02 Context Engineering for managing what goes into Claude Code's context window.