A03: LLM Fundamentals for Claude Code Users
Related chapters: Ch1 Installation & First Session, Ch3 Context Window Management
You do not need a PhD in machine learning to use Claude Code effectively. But understanding the key concepts behind large language models transforms you from someone who follows recipes into someone who can predict how Claude will behave, diagnose when something goes wrong, and optimize your workflow based on how the model actually works. This appendix covers the fundamentals at a practical level -- mental models, not math.
Transformer Architecture: The Engine Under the Hood
The Big Picture
Claude is a Transformer-based large language model (LLM). The Transformer architecture, introduced in the 2017 paper "Attention Is All You Need" (Vaswani et al.), is the foundation of virtually all modern LLMs.
At the highest level, a Transformer does one thing: given a sequence of tokens, it predicts the next token. It does this by running the input through a series of layers, each of which transforms the representation of every token based on its relationship to every other token in the sequence.
How Claude generates a response (simplified):
┌─────────────────────────────────────────────────┐
│ Input: "Fix the bug in src/auth/" │
│ │
│ Step 1: Tokenize │
│ ["Fix", " the", " bug", " in", " src", "/", │
│ "auth", "/"] │
│ │
│ Step 2: Embed (convert tokens to vectors) │
│ [0.23, -0.15, ...] [0.81, 0.42, ...] ... │
│ │
│ Step 3: Transformer layers (x96 for Opus) │
│ Each layer: Attention → Feed-forward → Repeat │
│ The representation of each token gets richer │
│ with each layer │
│ │
│ Step 4: Predict next token │
│ "I" (probability 0.35) │
│ "'ll" (probability 0.28) │
│ "Let" (probability 0.15) │
│ → Sample "I" → append → repeat from Step 3 │
│ │
│ Step 5: Keep generating until done │
│ "I'll start by reading the auth module..." │
└─────────────────────────────────────────────────┘Key Insight: One Token at a Time
Claude generates its entire response one token at a time, left to right. When you see Claude produce a multi-paragraph explanation with code examples, it did not "plan" the whole response and then write it. It predicted each token based on everything before it.
This has practical implications:
- Claude cannot "go back": once a token is generated, it influences all subsequent tokens. If Claude starts down a wrong path, it tends to continue rather than backtrack (unless it uses extended thinking to plan ahead).
- Longer responses are more expensive: each generated token requires a full forward pass through the model. A 1,000-token response costs 1,000 forward passes.
- Earlier tokens matter more: the first few tokens of Claude's response set the trajectory. This is why prompt engineering techniques like "think step by step" work -- they influence the initial trajectory of generation.
The Attention Mechanism: Why Context Windows Have Limits
What Attention Does
The attention mechanism is the core innovation of the Transformer. It allows each token in the sequence to "look at" every other token and determine how much to weight each one.
Think of it like a conversation in a room. In a traditional neural network (RNNs), information passes like a game of telephone -- each person only hears from the person before them, and the message degrades over distance. In a Transformer, everyone can hear everyone else directly. Token 5,000 can directly attend to token 3, without the information passing through tokens 4 through 4,999.
Attention as "who looks at whom":
Token: "Fix" "the" "bug" "in" "auth"
│ │ │ │ │
"Fix" ←── ●───────●──────●──────●──────●
"the" ←── ●───────●──────●──────●──────●
"bug" ←── ●───────●──────●──────●──────●
"in" ←── ●───────●──────●──────●──────●
"auth" ←── ●───────●──────●──────●──────●
Each token attends to ALL other tokens.
Connection strength varies (thicker = more attention).Quadratic Scaling: The Fundamental Constraint
Here is why context windows have limits. For N tokens, the attention mechanism computes N x N attention scores. This is quadratic scaling:
| Tokens | Attention computations | Relative cost |
|---|---|---|
| 1,000 | 1,000,000 | 1x |
| 10,000 | 100,000,000 | 100x |
| 100,000 | 10,000,000,000 | 10,000x |
| 200,000 | 40,000,000,000 | 40,000x |
Doubling the context window quadruples the computation. This is not a software limitation that will be patched away -- it is a mathematical property of the attention mechanism. Techniques like FlashAttention, sliding window attention, and sparse attention reduce the constant factor, but the fundamental scaling remains.
Practical implication: This is why context window management matters (Ch3). Even though Claude has a 200K token window, using 200K tokens is dramatically more expensive (in compute, latency, and money) than using 50K tokens. And the model's ability to attend to all information equally degrades as the window fills.
Multi-Head Attention: Different Perspectives
Claude does not have a single attention mechanism -- it has many "heads" of attention running in parallel. Each head learns to attend to different aspects of the input:
- One head might track syntactic structure (matching opening and closing braces in code)
- Another might track semantic relationships (connecting "bug" to "fix" to "auth")
- Another might track positional patterns (attending to nearby tokens for local context)
This is why Claude can simultaneously understand code syntax, follow your conversational intent, and remember project constraints. Different attention heads handle different aspects.
The "Lost in the Middle" Effect
Research by Liu et al. (2023) demonstrated that long-context models do not attend equally to all positions. Performance on retrieval tasks follows a U-shaped curve:
Retrieval accuracy vs. position in context:
High |● ●●
|●● ●●
| ●● ●●
| ●● ●●
| ●●● ●●
| ●●●● ●●●
| ●●●●●● ●●●●●●
Low | ●●●●●●●●●●●●●
|____________________________________________
Start Middle of context End
Models attend most to the start and end, least to the middle.This is not a bug -- it is a natural consequence of how positional encoding and attention interact. For Claude Code users, this means:
- CLAUDE.md (at the start of context) is well-attended
- Your latest prompt (at the end of context) is well-attended
- Information from turn 10 of a 30-turn conversation (in the middle) may be overlooked
This directly informs the context engineering strategies in A02 Context Engineering.
Why Claude "Hallucinates"
What Hallucination Is
Hallucination is when Claude generates text that is fluent and confident but factually incorrect. In Claude Code, this manifests as:
- Referencing files or functions that do not exist
- Claiming a test passed when it did not run
- Generating code that uses APIs that do not exist in the library
- Stating "facts" about your codebase that are wrong
Why It Happens
Hallucination is not a bug -- it is a natural consequence of how LLMs work. Claude is a probability distribution over next tokens. It generates text that is statistically likely given its training data and the current context. When the most likely continuation of a sentence is a plausible-sounding but incorrect fact, Claude will generate it.
Several factors increase hallucination risk:
Insufficient context: When Claude does not have enough information to answer accurately, it fills in gaps with plausible completions from its training data. This is why Claude might reference a popular API pattern that does not match your specific library version.
Over-confidence from training: Claude was trained on vast amounts of text where confident-sounding statements are usually correct. It learned that pattern and applies it even when it should express uncertainty.
Long context degradation: As described above, attention weakens over long contexts. Claude may "forget" a file it read earlier and hallucinate its contents.
Ambiguous instructions: When your prompt is vague, Claude has more degrees of freedom in generation, which means more room for hallucination.
How Tool Calling Reduces Hallucination
This is one of the most important insights for Claude Code users: tool calling is Claude's antidote to hallucination.
When Claude can read a file before describing it, it does not need to hallucinate the file's contents. When Claude can run a command before reporting its output, it does not need to guess. When Claude can search a codebase before claiming a function exists, it can verify rather than assume.
Without tools (high hallucination risk):
You: "What does the auth middleware do?"
Claude: "Based on common patterns, it probably validates JWT tokens
and checks for expired sessions..." (may be completely wrong)
With tools (low hallucination risk):
You: "What does the auth middleware do?"
Claude: [Read src/auth/middleware.ts]
Claude: "The auth middleware at src/auth/middleware.ts does three things:
1. Extracts the Bearer token from the Authorization header (line 12)
2. Validates the token signature using the JWT_SECRET (line 18)
3. Checks token expiration (line 24)
It does NOT check for session validity -- that's handled
separately in src/session/check.ts."
(verified by reading the actual file)Practical tip: When you need factual accuracy about your codebase, phrase your prompt to encourage tool use rather than recall. "Read the auth middleware and explain what it does" will produce more accurate results than "Explain how our auth works."
Temperature and Code Generation
What Temperature Is
Temperature is a parameter that controls the randomness of token selection. After the Transformer computes probabilities for all possible next tokens, temperature adjusts how "peaked" or "flat" that distribution is:
Next token probabilities for "import { useState } from '"
Temperature 0 (deterministic):
"react" ██████████████████████ 99%
"preact" █ 0.5%
"solid" ░ 0.1%
Other ░ 0.4%
→ Always picks "react"
Temperature 0.5 (low randomness):
"react" ████████████████ 85%
"preact" ██ 8%
"solid" █ 4%
Other █ 3%
→ Usually picks "react", occasionally "preact"
Temperature 1.0 (default):
"react" ████████████ 60%
"preact" ████ 18%
"solid" ███ 12%
Other ██ 10%
→ More variety in outputsTemperature in Claude Code
Claude Code uses a default temperature tuned for agentic tasks. You generally do not need to adjust it, but understanding its effects is useful:
- Lower temperature produces more predictable, conventional code. Good for boilerplate, standard patterns, and tasks with one correct answer.
- Higher temperature produces more creative, varied code. Good for brainstorming, generating alternatives, and tasks where novelty matters.
When Claude Code generates tool calls (deciding which tool to use, what parameters to pass), it uses lower effective temperature because tool calls need to be precise and well-structured. When generating explanations or creative solutions, it may use slightly higher temperature.
Why identical prompts give different results: Claude is a stochastic (probabilistic) system. Running the same prompt twice may produce different outputs because of temperature-based sampling. This is normal and expected. If you need deterministic results, it must be handled at the application level (e.g., by caching or by running in a mode with temperature set to 0).
BPE Tokenization: How Text Becomes Tokens
What BPE Is
Byte Pair Encoding (BPE) is the algorithm that converts text into the tokens that Claude processes. Understanding BPE explains why:
- Different text types have different token costs
- Code is more expensive than prose
- Some languages are more token-efficient than others
- Certain identifiers are "cheaper" for Claude to generate
How BPE Works
BPE starts with individual characters and iteratively merges the most frequent pairs:
Training phase (done once, on a large corpus):
Step 0: Start with individual characters
"the" → [t, h, e]
Step 1: Most frequent pair is "t, h" → merge to "th"
"the" → [th, e]
Step 2: Most frequent pair is "th, e" → merge to "the"
"the" → [the]
After thousands of merge steps, the vocabulary contains:
- Common words as single tokens: "the", "function", "return"
- Common subwords: "tion", "ing", "pre"
- Less common words split into subwords: "XMLHttpRequest" → ["XML", "Http", "Request"]
- Rare characters as individual tokens: emoji, unusual UnicodePractical Token Counting
Here is how different content types tokenize (approximate, based on Claude's tokenizer):
English prose:
"The function validates the authentication token."
→ 7 tokens (~6.5 chars/token)
Python code:
"def validate_token(token: str) -> bool:"
→ 11 tokens (~3.5 chars/token)
JSON:
{"name": "validate", "type": "function"}
→ 13 tokens (~3 chars/token)
Chinese text:
"这个函数验证认证令牌。"
→ 8 tokens (~1.4 chars/token)
Variable names:
"getUserById" → 4 tokens (camelCase splits)
"get_user_by_id" → 7 tokens (snake_case: more tokens)
"x" → 1 tokenKey takeaway: Code is 2-3x more token-expensive than English prose. A 200-line Python file might consume 2,000-3,000 tokens. A verbose JSON configuration file can consume tokens at an alarming rate. This is why controlling tool result size (see A02 Context Engineering) matters so much.
Why Tokenization Matters for Claude Code
Cost estimation: Knowing token counts helps you estimate session costs before starting. A codebase exploration that reads 20 files of ~200 lines each consumes roughly 40,000-60,000 tokens just in file reads.
Context budget planning: If your CLAUDE.md is 500 words of English prose, that is about 375 tokens. If it is 500 lines of code examples, that might be 2,000+ tokens. Choose wisely what goes in CLAUDE.md.
Output efficiency: Claude generating a 50-line code file costs more output tokens than generating a 5-line explanation. When you ask Claude to "explain your plan before coding," the explanation is cheap. The code generation is expensive.
Model Size: Haiku vs Sonnet vs Opus
What Model Size Means
Anthropic offers multiple Claude models of different sizes. "Size" refers to the number of parameters -- the learned weights that define the model's behavior.
| Model | Relative Size | Strengths | Trade-offs |
|---|---|---|---|
| Haiku | Smallest | Fast, cheap, good for simple tasks | Less capable on complex reasoning |
| Sonnet | Medium | Balanced speed/capability, good for most tasks | Slower than Haiku, cheaper than Opus |
| Opus | Largest | Best reasoning, most capable on hard tasks | Slowest, most expensive |
Mental Models for Each Size
Haiku -- Think "competent junior developer." Fast at straightforward tasks: formatting code, writing boilerplate, simple edits, running commands. Struggles with complex multi-step reasoning, subtle bugs, and architectural decisions. Use for high-volume, low-complexity tasks (e.g., CI/CD reviews, simple formatting).
Sonnet -- Think "solid mid-level developer." Handles most day-to-day tasks well: feature implementation, bug investigation, test writing, code review. Good balance of speed and capability. This is the model most Claude Code users should default to.
Opus -- Think "senior architect." Excels at complex reasoning: understanding multi-file architectures, finding subtle bugs, making architectural decisions, writing complex algorithms. Use for tasks where quality matters more than speed: difficult refactors, security reviews, algorithm design.
How Model Size Affects Tool Calling
Larger models are better at:
- Choosing the right tool: Opus more reliably picks the optimal tool for a situation. Haiku sometimes reads entire files when a grep would suffice.
- Constructing correct parameters: Opus produces more accurate file paths, search patterns, and command flags.
- Recovering from errors: When a tool call fails, Opus is better at diagnosing why and trying an alternative approach.
- Multi-step planning: Opus can maintain a coherent plan across many tool calls. Haiku may lose track of the plan after several steps.
Scaling Laws
Research on neural scaling laws (Kaplan et al., 2020; Hoffmann et al., 2022) shows that model capability increases predictably with model size, training data, and compute. The relationship is roughly logarithmic: to get a noticeable improvement in capability, you need to roughly 10x the model size.
This explains why the jump from Haiku to Sonnet feels significant, but the jump from Sonnet to Opus may feel smaller for many tasks. The capability difference is real but concentrated in the hardest tasks. For simple tasks, all three models perform similarly.
The "Needle in a Haystack" Research
What the Research Shows
"Needle in a haystack" (NIAH) tests evaluate how well a model can retrieve specific information embedded in a large context. A "needle" (a distinctive fact) is placed at various positions within a "haystack" (a large amount of filler text), and the model is asked to retrieve it.
Key findings from NIAH research across multiple models:
Models are not perfect retrievers: Even with 100% context available, models miss information, especially in the middle of the context.
Position matters: Retrieval accuracy is highest at the start and end of the context, lowest in the middle (the U-shaped curve described earlier).
Context length matters: As the haystack grows, retrieval accuracy for any given position drops. A model that perfectly retrieves a needle in 10K tokens of context may miss it in 100K tokens.
Multiple needles interfere: When there are multiple pieces of relevant information scattered throughout the context, the model may retrieve one but miss others.
Implications for Claude Code Users
The NIAH research directly explains several Claude Code behaviors:
Why Claude "forgets" earlier instructions: Your constraint from turn 3 is a needle in the haystack of 50K tokens of file reads, test output, and conversation history. It is not that Claude chose to ignore it -- the attention mechanism may simply attend less strongly to it.
Why CLAUDE.md is effective: CLAUDE.md sits at the very beginning of the context, in the high-attention zone. It is one of the most reliably attended positions.
Why repeating constraints works: Repeating a critical constraint in your current prompt places it at the end of the context, in the other high-attention zone. This "bookend" strategy exploits the U-shaped attention curve.
Why /compact can help: Compaction removes the haystack, making remaining needles easier to find. After compaction, your key instructions are no longer buried in tens of thousands of tokens of file contents.
Putting It All Together: A Mental Model for Claude Code
Here is a unified mental model that connects all these fundamentals:
Your prompt Claude's brain Claude's output
│ │ │
▼ ▼ ▼
┌────────┐ ┌──────────────────┐ ┌──────────────────┐
│Tokenize│───▶│ 96 Transformer │───▶│ Sample next │
│ your │ │ layers, each │ │ token from │
│ text │ │ with multi-head │ │ probability │──┐
└────────┘ │ attention over │ │ distribution │ │
│ ALL context │ └──────────────────┘ │
│ │ │ │
│ Context: │ ▼ │
│ - System prompt │ ┌──────────────────┐ │
│ - CLAUDE.md │ │ Is it a tool │ │
│ - History │ │ call or text? │ │
│ - Tool results │ └────────┬─────────┘ │
│ - Your prompt │ │ │
└──────────────────┘ Tool call │ Text │
│ │ │
▼ ▼ │
┌────────┐ Output │
│Execute │ to you │
│tool, │ │
│add │ │
│result │◀──────────────┘
│to │ (loop until
│context │ response
└────────┘ complete)Every concept from this appendix maps to a node in this diagram:
- Tokenization determines how much each piece of input costs
- Attention determines what information Claude actually "sees"
- Context window is the maximum capacity of the "ALL context" buffer
- Temperature controls how the next token is sampled
- Model size determines the richness of each Transformer layer
- Hallucination occurs when the probability distribution favors plausible but incorrect tokens
- Tool calling grounds generation in real data, reducing hallucination
Key Takeaways
- Claude generates one token at a time. It does not plan ahead (except with extended thinking). Understanding this explains many behaviors.
- Attention is quadratic. Doubling context length quadruples cost. This is physics, not a product limitation.
- The "lost in the middle" effect is real. Put critical information at the start (CLAUDE.md) or end (your prompt) of the context, not buried in the middle.
- Tool calling is the antidote to hallucination. Encourage Claude to read/verify rather than recall/guess.
- Code is token-expensive. Budget for 2-3x more tokens than you would expect from line counts alone.
- Model size matters most for hard tasks. For simple tasks, Sonnet and Opus produce similar results. For complex reasoning, the difference is significant.
- Temperature means identical prompts can give different results. This is by design, not a bug.
See also: A02 Context Engineering for strategies built on these fundamentals, and A07 Token Economics for detailed cost analysis based on tokenization.