Chapter 6: Skills — Build Your Own Slash Commands
Appendix links: A01 Prompt Engineering covers how to write effective skill instructions that produce consistent, high-quality output. Apply those techniques when writing your skill Markdown files.
What You'll Build
Three production skills adapted from the most-starred Claude Code repos:
/deep-research— Multi-source research with citations and structured reports (from ECC)/review— Confidence-based code review with severity ratings and file:line references (from gstack)/investigate— Root-cause debugging that reads logs, traces code paths, and identifies the bug
These aren't simplified versions. They use the same patterns that Everything Claude Code and gstack deploy in production teams.
How Skills Work
A Skill is a Markdown file in .claude/skills/ that becomes a slash command. When you type /review in Claude Code, it loads the corresponding file, substitutes your arguments, and Claude follows the instructions.
You type: /review src/auth/login.ts
|
v
Claude loads .claude/skills/review.md
|
v
Reads YAML frontmatter (name, tools, context mode)
|
v
Replaces $ARGUMENTS with "src/auth/login.ts"
|
v
Executes !`backtick` commands for dynamic context
|
v
Claude follows the skill instructionsSkill vs CLAUDE.md
| CLAUDE.md | Skill | |
|---|---|---|
| Loaded | Every session, automatically | On demand, when you type /command |
| Purpose | Project rules, coding standards | Specific workflow for a specific task |
| Scope | Always active | Active only during execution |
| Analogy | Team handbook | A power tool you pull out when needed |
Rule of thumb: If it applies to every interaction, put it in CLAUDE.md. If it's a specific workflow you trigger manually, make it a skill.
SKILL.md Format
---
name: review
description: "Confidence-based code review with severity ratings"
user-invocable: true
context: fork
allowed-tools:
- Read
- Glob
- Grep
- Bash
---
Your skill instructions here.
Arguments from the user: $ARGUMENTSFrontmatter Fields
| Field | Type | What It Does |
|---|---|---|
name | string | The slash command name (review becomes /review) |
description | string | Shown in the / autocomplete menu |
user-invocable | boolean | true = appears in the menu, false = only callable by agents |
allowed-tools | list | Whitelist of tools this skill can use |
context | string | fork runs in a subagent (keeps main context clean) |
disable-model-invocation | boolean | true = pure template, no model call |
Dynamic Context with Backtick Injection
Use !`command` to run a command at skill load time and inject the output:
Current branch: !`git branch --show-current`
Last 5 commits: !`git log --oneline -5`Claude sees the actual output, not the command. This saves tool calls.
context: fork — Keep Your Context Clean
Heavy skills (lots of file reads, long output) should use context: fork. The skill runs in a subagent with its own context window. Only the final result comes back to your main session.
Main session (stays clean)
|
+--> Forked subagent
| reads 50 files
| analyzes patterns
| generates report
|
<-- only the report comes back (~500 tokens)Without fork, all 50 file reads would eat into your main context, causing faster compaction and degraded quality.
Skill File Locations
.claude/skills/ <-- Project skills (commit to Git, shared with team)
~/.claude/skills/ <-- Personal skills (not committed, yours only)Project skills override personal skills if they share the same name.
Demo 15: /deep-research Skill
Inspired by ECC's
/deep-researchskill — multi-source research with structured output and citations.
The Problem
You need to understand a technology, evaluate a library, or research an architecture pattern. You could spend 30 minutes reading blog posts and docs, or you could have Claude do systematic research with proper citations in 2 minutes.
Build It
1. Set Up
mkdir -p ~/claude-demos/demo-15 && cd ~/claude-demos/demo-15
git init && git branch -M main
mkdir -p .claude/skills src
cat > CLAUDE.md << 'EOF'
# Research Demo Project
- When researching, cite specific sources
- Present findings in structured format
- Be honest about confidence levels
EOF
cat > package.json << 'EOF'
{
"name": "research-demo",
"version": "1.0.0",
"type": "module"
}
EOF
git add -A && git commit -m "feat: initial setup"2. Create the Skill
cat > .claude/skills/deep-research.md << 'SKILLEOF'
---
name: deep-research
description: "Multi-source research with citations and structured analysis"
user-invocable: true
context: fork
allowed-tools:
- Bash
- Read
- Glob
- Grep
---
# Deep Research
You are a senior engineer conducting thorough research. Your job is to produce a well-sourced, structured analysis that someone could use to make a real decision.
## Research Topic
$ARGUMENTS
## Process
### Phase 1: Gather Sources
Search for information using multiple approaches:
1. Check if there are relevant files in the current project (grep for mentions, check package.json dependencies)
2. Use `bash` to query package registries: `npm info <package>`, `pip show <package>`, etc.
3. Check GitHub stats if relevant: `gh api repos/<owner>/<repo> --jq '.stargazers_count, .open_issues_count, .pushed_at'`
4. Look at the project's own code for existing usage patterns
### Phase 2: Analyze
For each source of information, evaluate:
- **Recency**: When was this information last updated?
- **Reliability**: Official docs vs blog post vs Stack Overflow answer
- **Relevance**: Does this directly address the research question?
### Phase 3: Report
Output a structured report in this exact format:
```markdown
# Research: [Topic]
## Summary
[2-3 sentence executive summary]
## Key Findings
### Finding 1: [Title]
- **Source**: [Where this came from]
- **Confidence**: High/Medium/Low
- **Details**: [What you found]
### Finding 2: [Title]
...
## Comparison Table (if applicable)
| Criterion | Option A | Option B |
|-----------|----------|----------|
| ... | ... | ... |
## Recommendation
[What you would actually recommend and why]
## Caveats
- [What you're not sure about]
- [What needs more investigation]
## Sources
1. [Source with specific location/URL]
2. ...Rules
- Never make up information. If you can't find something, say so.
- Confidence levels matter: High = verified from official sources, Medium = from reliable community sources, Low = inferred or from a single source.
- Include version numbers and dates wherever possible.
- If you find conflicting information, present both sides. SKILLEOF
git add .claude/skills/deep-research.md git commit -m "feat: add /deep-research skill"
#### 3. Test It
```bash
claude/deep-research Compare Vitest vs Jest for a TypeScript monorepo — performance, DX, migration path> /deep-research Compare Vitest vs Jest for a TypeScript monorepo — performance, DX, migration path
⏳ Forking subagent for deep-research...
● Bash: npm info vitest version
→ 3.2.1
● Bash: npm info jest version
→ 30.0.2
● Bash: gh api repos/vitest-dev/vitest --jq '.stargazers_count'
→ 14832
● Read: package.json
→ Checking current test dependencies...
● Bash: npm info vitest time --json | jq 'to_entries | last'
→ {"key":"3.2.1","value":"2025-05-12T..."}
✔ Subagent complete. Result:
# Research: Vitest vs Jest for TypeScript Monorepo
## Summary
Vitest offers significantly faster execution for TypeScript projects
due to native ESM and Vite-based transforms. Jest 30 closed the gap
with its new ESM support, but migration cost is low if you are
already using Vite.
## Key Findings
...Try another one:
/deep-research What are the security implications of using JWT tokens stored in localStorage vs httpOnly cookies?What Just Happened?
Key takeaways:
- The skill ran in a forked subagent, keeping your main context clean
- Claude gathered data from multiple sources (npm registry, GitHub API, local files) before forming conclusions
- Every finding in the report includes a confidence level and source attribution
- The
context: forksetting means only the final report came back to your session
Why This Pattern Works
ECC's version of this skill is one of its most-used. The key design decisions:
context: fork— Research generates massive context (file reads, command outputs). Fork keeps your main session clean.- Structured output — The report format is rigid on purpose. It forces Claude to organize findings instead of writing a wall of text.
- Confidence levels — This is the difference between useful research and hallucinated research. Claude must qualify each finding.
- Multiple sources — The skill instructs Claude to cross-reference, not just grab the first answer.
Demo 16: /review Skill
Inspired by gstack's
/reviewskill — confidence-based code review with specific file:line references.
The Problem
Code review is the most common AI-assisted task, but most people just ask "review this code" and get vague feedback. gstack's approach is different: every finding gets a confidence score, a severity rating, and a specific file:line reference. Findings below 80% confidence are filtered out.
Build It
1. Create Reviewable Code
cd ~/claude-demos/demo-15
# Create code with real issues at specific lines
mkdir -p src/api
cat > src/api/users.js << 'EOF'
import { db } from '../db.js';
import { hash } from '../utils/crypto.js';
export async function createUser(req, res) {
const { username, email, password } = req.body;
// Line 7: No input validation at all
const hashedPassword = await hash(password);
// Line 10: SQL injection via string interpolation
const existing = await db.query(
`SELECT id FROM users WHERE email = '${email}'`
);
if (existing.rows.length > 0) {
return res.status(409).json({ error: 'Email taken' });
}
// Line 18: Another SQL injection
const result = await db.query(
`INSERT INTO users (username, email, password)
VALUES ('${username}', '${email}', '${hashedPassword}')`
);
// Line 23: Returning the hashed password in the response
return res.status(201).json({
id: result.rows[0].id,
username,
email,
password: hashedPassword
});
}
export async function deleteUser(req, res) {
// Line 31: No auth check — any user can delete any other user
const { id } = req.params;
await db.query(`DELETE FROM users WHERE id = ${id}`);
return res.json({ success: true });
}
export async function listUsers(req, res) {
// Line 38: No pagination — will OOM on large tables
const result = await db.query('SELECT * FROM users');
// Line 40: Returning all fields including password hashes
return res.json(result.rows);
}
EOF
git add -A && git commit -m "feat: add user API endpoints"2. Create the Skill
cat > .claude/skills/review.md << 'SKILLEOF'
---
name: review
description: "Confidence-based code review with severity ratings and file:line references"
user-invocable: true
context: fork
allowed-tools:
- Read
- Glob
- Grep
- Bash
---
# Code Review
You are a senior engineer performing a thorough code review. This isn't a surface-level scan — you're looking for real issues that would block a PR or cause incidents in production.
## Target
Review: $ARGUMENTS
If no specific files are provided, review the changes on the current branch:
!`git diff main --name-only 2>/dev/null | head -20`
## Review Process
### Step 1: Read the Code
Read every file in scope. Don't skim — read line by line.
### Step 2: Analyze Each Finding
For every issue you find, assign:
- **Severity**: P0 (security/data loss), P1 (bug), P2 (code quality), P3 (style/nit)
- **Confidence**: 0-100%. Only report findings with confidence >= 80%.
- **File:Line**: Exact location, e.g., `src/api/users.js:10`
### Step 3: Generate Report
Use this exact format for each finding:[P0] SQL Injection in user query File: src/api/users.js:10 Confidence: 95% Problem: User input is interpolated directly into SQL string. The email parameter from req.body is inserted without parameterization, allowing arbitrary SQL execution. Fix: Use parameterized queries. const existing = await db.query( 'SELECT id FROM users WHERE email = $1', [email] );
### Output FormatCode Review: [files reviewed]
Summary
- Files reviewed: N
- Findings: N (X critical, Y bugs, Z quality)
- Overall: APPROVE / REQUEST CHANGES / BLOCK
Critical (P0)
[findings...]
Bugs (P1)
[findings...]
Code Quality (P2)
[findings...]
Style (P3)
[findings...]
What's Good
[1-2 things the code does well — always include this]
### Rules
- **Be specific**: "auth.ts:47" not "the authentication module"
- **Confidence threshold**: Drop anything below 80%. False positives waste everyone's time.
- **Actionable fixes**: Every finding must include a concrete fix, not just "consider improving this"
- **No AI vocabulary**: Don't say "I'd suggest considering" — say "Change X to Y"
- P0 findings automatically mean REQUEST CHANGES or BLOCK
SKILLEOF
git add .claude/skills/review.md
git commit -m "feat: add /review skill"3. Test It
claude/review src/api/users.js> /review src/api/users.js
⏳ Forking subagent for review...
● Read: src/api/users.js
→ 42 lines read
● Grep: 'db.query' in src/
→ 5 matches across 1 file
● Grep: 'parameterized\|prepared\|\$[0-9]' in src/
→ 0 matches — no parameterized queries found
✔ Subagent complete. Result:
# Code Review: src/api/users.js
## Summary
- Files reviewed: 1
- Findings: 6 (3 critical, 1 bug, 2 quality)
- Overall: BLOCK
## Critical (P0)
[P0] SQL Injection in email lookup
File: src/api/users.js:11
Confidence: 98%
Problem: Email is interpolated directly into SQL via template literal.
Fix: Use parameterized query: db.query('SELECT id FROM users WHERE email = $1', [email])
[P0] SQL Injection in INSERT
File: src/api/users.js:19
Confidence: 98%
...
[P0] Missing authorization on deleteUser
File: src/api/users.js:31
Confidence: 95%
...What Just Happened?
Key takeaways:
- The 80% confidence threshold filtered out uncertain findings, keeping the report focused
- Every finding has a specific file:line reference, not vague descriptions
- The skill used Grep to cross-reference patterns, not just Read to scan the file
- P0 findings automatically triggered a BLOCK verdict
The "auth.ts:47" not "the authentication module" rule is straight from gstack's voice guidelines. It forces Claude to be precise instead of vague.
Demo 17: /investigate Skill
A root-cause debugging skill that traces through code like a senior engineer would.
The Problem
Someone reports a bug. You could ask Claude "why is this broken?" and get a guess, or you could use a structured investigation skill that reads logs, traces code paths, identifies the root cause, and proposes a fix with confidence levels.
Build It
1. Create a Codebase with a Real Bug
cd ~/claude-demos/demo-15
mkdir -p src/services src/utils
cat > src/services/order.js << 'EOF'
import { db } from '../db.js';
import { calculateTotal } from '../utils/pricing.js';
import { sendNotification } from '../utils/notify.js';
export async function placeOrder(userId, items) {
// Fetch product details
const products = [];
for (const item of items) {
const product = await db.query(
'SELECT * FROM products WHERE id = $1',
[item.productId]
);
products.push({ ...product.rows[0], quantity: item.quantity });
}
// Calculate total
const total = calculateTotal(products);
// Create order
const order = await db.query(
'INSERT INTO orders (user_id, total, status) VALUES ($1, $2, $3) RETURNING id',
[userId, total, 'confirmed']
);
// BUG: sendNotification is called with order.id but
// order is a query result object, not the order itself
await sendNotification(userId, order.id, total);
return { orderId: order.rows[0].id, total };
}
EOF
cat > src/utils/pricing.js << 'EOF'
export function calculateTotal(products) {
let total = 0;
for (const p of products) {
// BUG: floating point arithmetic without rounding
// 19.99 * 3 = 59.970000000000006
total += p.price * p.quantity;
}
// Missing: Math.round(total * 100) / 100
return total;
}
export function applyDiscount(total, discountPercent) {
// BUG: no validation — negative discount increases price
return total * (1 - discountPercent / 100);
}
EOF
cat > src/utils/notify.js << 'EOF'
export async function sendNotification(userId, orderId, total) {
// This will fail silently if orderId is undefined
// because order.id doesn't exist on the query result
console.log(`Order ${orderId}: $${total} for user ${userId}`);
// In production this would call an email/SMS service
// The silent failure means users never get confirmation
}
EOF
git add -A && git commit -m "feat: add order system with pricing and notifications"2. Create the Skill
cat > .claude/skills/investigate.md << 'SKILLEOF'
---
name: investigate
description: "Root-cause debugging — trace code paths, identify the bug, propose fix"
user-invocable: true
context: fork
allowed-tools:
- Read
- Glob
- Grep
- Bash
---
# Bug Investigation
You are a senior engineer doing root-cause analysis. Don't guess — trace through the code systematically.
## Bug Report
$ARGUMENTS
## Investigation Process
### Phase 1: Reproduce
1. Read the reported symptoms carefully
2. Find the entry point in the code
3. Trace the execution path step by step
### Phase 2: Evidence Gathering
1. Read every file in the execution path
2. Check for error handling (or lack thereof)
3. Look at data transformations — where do values change?
4. Search for related issues: `grep -rn` for the function names, variable names
5. Check git history for recent changes: `git log --oneline -10 -- <file>`
### Phase 3: Root Cause Analysis
For each potential root cause:
- **Evidence**: What specific code points to this being the issue?
- **Confidence**: High/Medium/Low
- **Impact**: What's the blast radius?
### Phase 4: Report
```markdown
# Investigation: [Bug Title]
## Symptom
[What the user reported]
## Root Cause
[What's actually broken and why]
## Evidence Trail
1. [Entry point] -> [Function A at file:line]
2. [Function A] -> [Function B at file:line]
3. [Function B]: HERE IS THE BUG
- Expected: [what should happen]
- Actual: [what happens instead]
- Code: [the specific line]
## Contributing Factors
- [Other issues that make this worse]
## Fix
### Immediate Fix (do this now)
[Specific code change with before/after]
### Proper Fix (do this soon)
[More thorough solution]
## Prevention
- [How to prevent this class of bug in the future]Rules
- Follow the data. Read the actual code, don't guess based on function names.
- Every claim needs a file:line reference.
- If you find multiple bugs during investigation, report all of them, but clearly mark which one causes the reported symptom.
- "I don't know" is better than a wrong diagnosis. SKILLEOF
git add .claude/skills/investigate.md git commit -m "feat: add /investigate skill"
#### 3. Test It
```bash
claude/investigate Users report that order confirmation notifications are never received, and some order totals show values like $59.970000000000006 instead of $59.97> /investigate Users report that order confirmation notifications are never received, and some order totals show values like $59.970000000000006 instead of $59.97
⏳ Forking subagent for investigate...
● Grep: 'sendNotification' in src/
→ 2 matches: order.js:27, notify.js:1
● Read: src/services/order.js
→ 32 lines read
● Read: src/utils/notify.js
→ 8 lines read
● Read: src/utils/pricing.js
→ 14 lines read
● Bash: git log --oneline -10 -- src/services/order.js
→ 1 commit: "feat: add order system..."
✔ Subagent complete. Result:
# Investigation: Missing Order Notifications + Broken Totals
## Root Cause
Two separate bugs:
1. **Notifications (Primary)**: order.js:27 calls
sendNotification(userId, order.id, total) but `order` is a
pg query result object — it has `rows`, `rowCount`, `command`,
but no `.id` property. The correct access is
`order.rows[0].id`. The notification receives `undefined`
as the orderId.
2. **Totals**: pricing.js:5 accumulates floating point products
without rounding. 19.99 * 3 = 59.970000000000006.
## Evidence Trail
1. placeOrder (order.js:5) -> calculateTotal (pricing.js:1)
2. pricing.js:5: total += p.price * p.quantity ← no rounding
3. placeOrder (order.js:27): order.id ← WRONG, should be
order.rows[0].id
4. notify.js:4: silently logs "Order undefined: $59.97..."
...What Just Happened?
Key takeaways:
- Claude followed the data flow, not guesses -- it traced from the symptom back to the root cause
- Both bugs were identified and clearly distinguished: which one caused the reported symptom vs. a contributing factor
- Every claim has a file:line reference backed by an actual Read of the code
- The git history check confirmed the bug was original, not a regression
When Things Go Wrong
Skill Not Found
You type /deep-research but nothing happens:
> /deep-research Compare React vs Vue
⚠ Unknown command: /deep-research
Available commands: /help, /clear, /compact, /review ...Common causes:
Wrong directory — Skills are loaded from the project's
.claude/skills/directory. If you started Claude in a different folder, it will not find the skill.bash# Wrong: started Claude in home directory cd ~ && claude # /deep-research won't be found because ~/claude-demos/demo-15/.claude/skills/ is not here # Right: start Claude in the project directory cd ~/claude-demos/demo-15 && claudeTypo in the
namefield — The frontmatternamemust match what you type. If the file isdeep-research.mdbut the frontmatter saysname: deepresearch(no hyphen), the command/deep-researchwill not resolve.Missing
user-invocable: true— Without this field, the skill exists but does not appear in the autocomplete menu and cannot be triggered by the user.
Skill File Syntax Errors
If the YAML frontmatter is malformed, the skill will fail to load:
> /review src/api/users.js
⚠ Error loading skill: review
YAML parse error at line 3: bad indentation of a mapping entryCommon YAML mistakes:
- Tabs instead of spaces (YAML requires spaces)
- Missing quotes around descriptions that contain colons
- Incorrect indentation in the
allowed-toolslist
# WRONG — colon in unquoted string breaks YAML
description: Confidence-based review: with ratings
# RIGHT — quote the string
description: "Confidence-based review with ratings"Conflicting Skill Names
If you have both .claude/skills/review.md (project) and ~/.claude/skills/review.md (personal), the project skill wins. This can be confusing if your personal version is the one you expect.
# Check which skills are loaded
ls .claude/skills/ # Project skills (higher priority)
ls ~/.claude/skills/ # Personal skills (lower priority)If you want to use your personal version in a specific project, either remove the project skill or rename one of them.
Skill Design Patterns
Pattern: Gatekeeper (from gstack's /ship)
---
name: ship
allowed-tools:
- Bash
- Read
---
Before deploying, verify ALL of these pass:
1. `git status` shows clean working tree
2. `npm test` passes with no failures
3. `npm run lint` passes with no errors
4. `npm run build` completes successfully
5. No P0/P1 issues in the last review
If ANY check fails, stop and report what failed. Do not proceed.Pattern: Multi-Role Pipeline (from gstack's /autoplan)
---
name: autoplan
context: fork
---
Evaluate this feature request from three perspectives:
**Product (CEO review)**: Is this worth building? Impact vs effort?
**Design (Design review)**: How should the UX flow work?
**Engineering (Eng review)**: What's the technical approach? What are the risks?
For each perspective, rate 1-10 and explain.
Conclude with: BUILD / DEFER / REJECT and next steps.Pattern: Eval-Driven (from ECC's /eval-harness)
---
name: eval
context: fork
---
Run the evaluation harness for: $1
1. Run the test suite 3 times: `npm test -- --reporter=json`
2. Calculate pass@1 (first run) and pass@3 (any of 3 runs)
3. Compare against the baseline in `.claude/eval-baseline.json`
4. Report regressions and improvementsExercise: Build a /changelog Skill
Create a /changelog skill that reads your git history and generates release notes:
- Read the last N commits (default 20, configurable via
$1) - Group by Conventional Commit type (Features, Fixes, Other)
- Generate markdown suitable for a GitHub Release
- Include the date range and contributor list
Bonus: Have it automatically detect breaking changes from commits with ! in the type (e.g., feat!: remove legacy API).
Success Criteria
- [ ] Skill file is in
.claude/skills/changelog.md - [ ] Uses
context: fork(it reads a lot of git history) - [ ] Supports
$1parameter for commit count - [ ] Uses
!`git log --oneline -5`for dynamic context - [ ] Groups commits by type
- [ ] Output is ready to paste into GitHub Releases
Knowledge Check
Summary
Skills turn repetitive workflows into one-command operations. The three patterns in this chapter — research, review, and investigation — are the most common skill types in production Claude Code setups.
Key points:
context: forkis essential for skills that read many files — it protects your main context- Confidence thresholds (from gstack's 80% rule) prevent false positives in reviews
- Structured output formats force Claude to organize instead of ramble
allowed-toolslimits blast radius — a review skill shouldn't be able to Write files!`backtick`injection saves tool calls by pre-loading context at skill load time
Going deeper: See A01 Prompt Engineering for techniques on writing skill instructions that produce consistent, high-quality output. The same principles that make a good system prompt make a good skill file.
Next: Chapter 7: MCP Server Integration — Connect Claude to GitHub, documentation servers, and browser automation through the Model Context Protocol.