You’re working on a problem. You don’t consciously think, “I should load the deep-research skill.” Instead, Claude Code does it for you. Behind the scenes, a trigger system is matching your task to the skills that matter most.
Let’s dive into how this works—because understanding triggers means you can write better skills, debug matching issues, and architect agent systems that automatically find the right tools at the right moment.
The Problem Triggers Solve
Imagine you’re managing 142 skills across 36 agents. When you say, “I need comprehensive research on authentication patterns,” how does Claude Code know to load deep-research instead of github-api or writing-skills?
Manual selection is impossible. You can’t expect every user to memorize skill names and catalog them by topic. And loading all 142 skills into every conversation wastes context tokens—potentially blocking entire conversations from running because you’ve consumed 60% of your context on skill descriptions.
Triggers solve this with pattern matching. They let skills say: “Load me when you see X, Y, or Z patterns.”
Without triggers, Claude Code would be like a library where you have to manually browse every shelf to find what you need. With triggers, the library assistant proactively hands you the right books.
How Triggers Work: Three Mechanisms
Claude Code implements a three-layer trigger matching system. Understanding each layer helps you debug why skills aren’t loading (or why the wrong ones are).
Layer 1: Description-Based Semantic Matching
This is the primary mechanism. Every SKILL.md file includes a description field:
---
name: deep-research
description: "Use when comprehensive research requires 20+ sources, multiple perspectives, or cross-domain synthesis - multi-step agentic research with iterative refinement based on Perplexity and ChatGPT Deep Research patterns"
version: 1.0.0
category: self-improvement
---
When you describe a task, Claude reads these descriptions and scores them semantically. Notice the structure:
- Trigger conditions: “when comprehensive research requires 20+ sources, multiple perspectives”
- What it does: “multi-step agentic research with iterative refinement”
- Context: “based on Perplexity and ChatGPT Deep Research patterns”
The description does two jobs:
- Signal to the matching algorithm: “Load me for X”
- Educate future Claude instances on when to use this skill
Here’s what makes descriptions work:
# ❌ BAD: Vague, no trigger conditions
description: "For research tasks"
# ❌ BAD: Too technical, misses problem statement
description: "Uses websearch API with pagination and 20+ sources"
# ✅ GOOD: Specific triggers + what it does
description: "Use when comprehensive research requires 20+ sources, multiple perspectives, or cross-domain synthesis - multi-step agentic research with iterative refinement"
The winning description explicitly names problem conditions (“requires 20+ sources”) rather than implementation details. This lets the matching algorithm find it when you say “I need to research this thoroughly” even if you don’t know the skill name.
Layer 2: Explicit Keyword Triggers
Some skills define explicit triggers as an array in their SKILL.md frontmatter:
---
name: deep-research
description: "Use when comprehensive research requires 20+ sources..."
version: 1.0.0
triggers:
[
"deep research",
"comprehensive research",
"research 20+ sources",
"systematic research",
"thorough investigation",
]
---
These are literal keyword matches. If your task contains any of these phrases, the skill loads automatically. This is a stronger signal than semantic matching.
The triggers array is optional but powerful for:
- Frequent skills that must load reliably (like
writing-skillsorgithub-api) - Specialized skills with unique vocabulary (“Chekhov’s guns”, “save-the-cat”, “three-act structure”)
- Emergency bypass for ambiguous situations
- Performance optimization on large skill catalogs
Example triggers from story-focused skills:
# plot-structure skill might have:
triggers: ["three-act structure", "hero's journey", "save the cat", "story beats", "narrative framework"]
# character-forge skill might have:
triggers: ["character arc", "wound-lie-truth", "character development", "motivation mapping"]
# dialogue-crafter skill might have:
triggers: ["subtext", "character voice", "dialogue flow", "conversation conflict"]
When you write: “Help me map this character’s wound-lie-truth arc,” the character-forge skill loads because wound-lie-truth matches an explicit trigger. It’s instant—no semantic analysis needed.
Layer 3: Regex and Advanced Pattern Matching
For complex matching, skills can define regex patterns. This is rare—used only for skills that detect very specific signatures:
---
name: condition-based-waiting
triggers:
- pattern: "timeout.*race|race.*condition|flaky.*test"
mode: regex
priority: high
- pattern: "async.*polling|event.*loop|event.*listener"
mode: regex
priority: medium
---
Regex triggers are powerful but expensive (they run against every task). Reserve them for:
- Complex technical patterns (patterns that keywords can’t express)
- Multi-word sequences in specific order
- Negative matches (when NOT to use)
- Domain-specific terminology that varies
This skill would load for tasks mentioning race conditions, timeouts, or async patterns—even if phrased differently than the explicit keywords.
Trigger Priority and Conflict Resolution
What happens when multiple skills match? Claude Code implements a priority system to avoid conflicts and resolve ambiguity intelligently:
Priority Tiers (highest wins):
1. Explicit keyword trigger (perfect match)
2. Regex pattern trigger (highest priority value)
3. Semantic description match (second highest priority)
4. Fallback generic skills (lowest priority)
Here’s a concrete example. Say you write: “Research async testing patterns with 25 sources.”
Three skills could match:
| Skill | Trigger Type | Match | Priority |
|---|---|---|---|
condition-based-waiting |
Regex: “async.*test” | MATCH | High |
deep-research |
Keywords: “research 20+ sources” | MATCH | High |
writing-skills |
Description: “comprehensive process” | MATCH | Medium |
Both condition-based-waiting and deep-research have explicit triggers (same priority tier). So Claude Code applies secondary ranking:
- More specific match wins (2 keywords vs 1 keyword, or longer phrase match)
- Recency wins (more recently updated skill)
- Frequency wins (more commonly used, based on historical usage)
In this case, both are equally specific. The system loads BOTH because they serve different purposes: one teaches async testing techniques, the other provides research methodology. They don’t conflict—they complement each other.
If they did conflict (same skill, different versions), Claude Code would:
- Check for version pins in the task (“Use skill v2.1 specifically”)
- Load the highest version number
- Document why it chose one over another in debug logs
Writing Effective Triggers: Real Examples
Let’s look at triggers in actual SKILL.md files and see what makes them work.
Example 1: The writing-skills Skill
This skill teaches TDD for documentation. It loads frequently, so triggers must be precise:
---
name: writing-skills
description: "Use when creating new skills, editing existing skills, or verifying skills work before deployment - applies TDD to process documentation by testing with subagents before writing, iterating until bulletproof against rationalization"
triggers:
- "writing skills"
- "skill creation"
- "skill development"
- "process documentation"
- "TDD for docs"
---
Why this works:
- Trigger 1 (“writing skills”) catches direct requests
- Trigger 2 (“skill creation”) catches developers working on new skills
- Trigger 3 (“skill development”) synonym for similar phrasing
- Trigger 4 (“process documentation”) catches meta-documentation tasks
- Trigger 5 (“TDD for docs”) catches developers thinking about test-driven approaches
Notice: They’re all phrases you’d naturally say, not implementation details like “YAML frontmatter” or “subagent testing harness.” The skill is discoverable because the triggers match human language.
Example 2: The github-api Skill
For API reference skills, triggers should map to natural language expressions:
---
name: github-api
description: "Use when making API calls to GitHub - REST API, GraphQL API, gh api command, custom queries, pagination, JSON processing, jq filters, creating custom integrations"
triggers:
- "github api"
- "gh api"
- "graphql github"
- "rest api github"
- "github query"
- "github automation"
- "github integration"
---
Why this works:
- Direct: “github api”, “gh api” catch exact usage
- Variant: “graphql github” for GraphQL-specific queries
- Problem-focused: “github automation”, “github integration” catch problems where API is the solution
- Jargon coverage: “rest api github” for terminology people use
Example 3: The deep-research Skill
Research skills have many trigger variations because people describe research in many ways:
---
name: deep-research
description: "Use when comprehensive research requires 20+ sources, multiple perspectives, or cross-domain synthesis - multi-step agentic research with iterative refinement based on Perplexity and ChatGPT Deep Research patterns"
triggers:
- "deep research"
- "comprehensive research"
- "research 20+ sources"
- "systematic research"
- "thorough investigation"
- "multi-perspective research"
- "research synthesis"
- "fact verification"
---
Notice the problem triggers like “research 20+ sources”—this isn’t a command, it’s a condition. When you say, “I need to research this thoroughly with lots of sources,” the system matches on the problem, not the keyword.
Example 4: A Story-Domain Skill
Story skills use domain-specific vocabulary:
---
name: plot-structure
description: "Use when designing story structure, three-act beats, story arcs, or narrative frameworks - applies proven story structure patterns to fiction writing"
triggers:
- "three-act structure"
- "story beats"
- "narrative framework"
- "hero's journey"
- "save the cat"
- "seven-point structure"
- "story arc"
- "plot outline"
- "act structure"
---
Each trigger is a recognized story architecture term. This lets writers using industry vocabulary automatically load the skill. A writer says “I need to structure this using the hero’s journey,” and boom—the skill loads. No searching, no clicking.
Trigger Matching in Practice: Five Real-World Scenarios
Before diving into debugging, let’s see how triggers work in actual situations. Understanding these scenarios helps you anticipate trigger behavior.
Scenario 1: Writing a Story Scene
You say: “I need to write a scene where the protagonist confronts her inner conflict about trust. Show, don’t tell. Sensory details.”
Which skills load?
Task words: "write", "scene", "protagonist", "conflict", "trust", "show don't tell", "sensory details"
Skill Matches:
1. prose-generator
- Trigger: "write scene" → MATCH (exact)
- Trigger: "show don't tell" → MATCH (exact)
- Score: 20 (two exact matches)
- Confidence: HIGH
2. character-forge
- Trigger: "character conflict" → MATCH (semantic)
- Trigger: "character motivation" → MATCH (semantic)
- Description: "character psychology" → MATCH (semantic)
- Score: 15 (multiple semantic matches)
- Confidence: HIGH
3. description-painter
- Trigger: "sensory details" → MATCH (exact)
- Trigger: "atmosphere" → MATCH (semantic)
- Score: 12 (one exact, one semantic)
- Confidence: MEDIUM
4. style-enforcer
- Description: "POV discipline, voice consistency" → MATCH (semantic)
- Score: 6 (light semantic match)
- Confidence: LOW
Result: All four load. The system prioritizes by score:
1. prose-generator (20) - HIGHEST
2. character-forge (15)
3. description-painter (12)
4. style-enforcer (6) - injected last, lighter context weight
The system loads all four because they all serve the task, but prioritizes prose-generator because it matched two explicit triggers. Each skill provides different value: prose-generator teaches scene writing, character-forge helps with motivation, description-painter adds sensory richness, style-enforcer maintains consistency.
Scenario 2: Debugging a GitHub Integration Bug
You say: “I’m getting 401 errors calling the GitHub API. How do I handle auth errors in my script?”
Which skills load?
Task words: "github", "api", "401", "auth", "error handling", "script"
Skill Matches:
1. github-api
- Trigger: "github api" → MATCH (exact)
- Trigger: "github script" → MATCH (exact)
- Description: "GitHub authentication" → MATCH (semantic)
- Score: 18
- Confidence: HIGH
2. github-secrets
- Trigger: "github auth" → MATCH (exact)
- Trigger: "401 error" → MATCH (exact)
- Context: "authentication" → MATCH
- Score: 16
- Confidence: HIGH
3. error-handling
- Trigger: "error handling" → MATCH (exact)
- Trigger: "debugging errors" → MATCH (semantic)
- Score: 12
- Confidence: MEDIUM
4. root-cause-tracing
- Description: "debugging complex issues" → MATCH (semantic)
- Score: 8
- Confidence: LOW
Result: github-api loads first (highest score), followed by github-secrets. The system knows you need both: one for GitHub-specific auth patterns, one for secret management.
Notice: github-api and github-secrets both scored high because they’re related. The system doesn’t conflict—it loads both because they serve different purposes. github-api teaches API authentication and 401 handling. github-secrets teaches credential management.
Scenario 3: Researching a Complex Topic
You say: “I need to understand AI agent architectures comprehensively. I want 25+ sources, different approaches, pros/cons analysis.”
Which skills load?
Task words: "understand", "comprehensively", "25+ sources", "different approaches", "pros/cons"
Skill Matches:
1. deep-research
- Trigger: "comprehensive research" → MATCH (exact)
- Trigger: "research 25+ sources" → MATCH (exact, exceeds 20 minimum)
- Description: "research 20+ sources, multiple perspectives" → MATCH
- Score: 25
- Confidence: CRITICAL
2. research-synthesis
- Trigger: "synthesis" → MATCH (semantic)
- Trigger: "compare approaches" → MATCH (semantic)
- Description: "comparative analysis" → MATCH
- Score: 14
- Confidence: HIGH
3. writing-skills
- Description: "comprehensive documentation" → MATCH (semantic, weak)
- Score: 5
- Confidence: LOW
Result: deep-research loads with the highest confidence (CRITICAL). research-synthesis loads as secondary. The system recognizes "25+ sources" as THE trigger signature for deep-research.
Scenario 4: The Ambiguous Case
You say: “Test this.”
Which skills load?
Task words: "test"
Skill Matches:
1. testing-skills
- Trigger: "unit test" → NO MATCH (too vague)
- Trigger: "test framework" → NO MATCH (too vague)
- Description: "testing" → WEAK MATCH (too generic)
- Score: 2
- Confidence: VERY LOW
2. condition-based-waiting
- Trigger: "test" → NO MATCH (not explicit)
- Trigger: "async test" → NO MATCH (incomplete phrase)
- Description: "async testing patterns" → WEAK MATCH
- Score: 2
- Confidence: VERY LOW
3. writing-skills
- Description: "testing skills" → WEAK MATCH (meta-context)
- Score: 1
- Confidence: VERY LOW
Result: None load with confidence. The system might ask for clarification: "Do you mean unit testing, integration testing, or something else?"
This is why triggers need specificity. "Test this" is too vague. It matches nothing strongly.
Scenario 5: Multiple Valid Interpretations
You say: “I’m writing a thriller with a detective searching for hidden clues. Help me maintain tension throughout.”
Which skills load?
Task words: "writing", "thriller", "detective", "clues", "tension"
Skill Matches:
1. prose-generator
- Trigger: "write scene" → SEMANTIC MATCH (context: writing)
- Trigger: "thriller" → EXACT MATCH
- Score: 18
- Confidence: HIGH
2. tension-mapping
- Trigger: "maintain tension" → EXACT MATCH
- Trigger: "thriller structure" → EXACT MATCH
- Trigger: "pacing tension" → SEMANTIC MATCH
- Score: 22
- Confidence: CRITICAL
3. plot-structure
- Trigger: "thriller plot" → EXACT MATCH
- Trigger: "mystery structure" → EXACT MATCH (detective = mystery)
- Context: "fiction" → MATCH
- Score: 20
- Confidence: HIGH
4. character-forge
- Trigger: "detective character" → SEMANTIC MATCH
- Description: "character motivation" → MATCH
- Score: 12
- Confidence: MEDIUM
Result: tension-mapping loads first (22 points). The system understands "maintain tension throughout" is the primary request. plot-structure and prose-generator also load because maintaining tension requires understanding structure. character-forge loads second because detective characterization supports the tension arc.
These scenarios show how triggers work in real conversations. The system isn’t picking one skill—it’s ranking multiple skills by relevance and injecting them in priority order, with context window space determining which ones actually get injected.
Debugging Trigger Matching
Sometimes skills don’t load when you expect them. Here’s how to debug:
Issue 1: Skill Description Is Too Generic
Problem: You write “I need research,” but the system loads github-api instead of deep-research.
Root cause: The github-api description is more generic:
description: "Use for API calls and queries" # TOO VAGUE
Fix: Make descriptions more specific:
description: "Use when making API calls to GitHub REST or GraphQL - specific for GitHub queries, not general research"
Issue 2: Triggers Use Wrong Language
Problem: You write “I need comprehensive fact-checking on these sources” but no skill loads.
Root cause: The fact-checking skill has triggers for implementation details:
triggers: ["verify source", "cross-reference", "footnote check"] # Too specific
Fix: Broaden triggers to problem language:
triggers:
[
"fact checking",
"verify sources",
"source verification",
"truth check",
"accuracy verification",
]
Issue 3: Multiple Conflicting Triggers
Problem: You write “test async code”, and both condition-based-waiting and testing-skills load. You only wanted one.
Root cause: Both have “test” in their triggers. The system has to choose.
Fix: Differentiate in trigger specificity:
# condition-based-waiting skill:
triggers:
- "async test"
- "race condition"
- "flaky test"
# testing-skills skill:
triggers:
- "unit test"
- "test framework"
- "test suite"
Now “async test” loads the first, “unit test” loads the second.
Issue 4: Version Mismatches
Problem: You have v1.0 and v2.0 of the same skill. The wrong version loads.
Root cause: Trigger system doesn’t see versions. It loads the highest available version by default.
Fix: Pin version explicitly in your request or use semantic versioning in triggers:
# v1.0 (legacy)
triggers:
- "legacy research" # Use only if explicitly requested
# v2.0 (current)
triggers:
- "deep research"
- "comprehensive research"
Issue 5: Skill Not Loading at All
Problem: You describe a valid use case, but no skill loads.
Root cause: Could be any of:
- Description is too generic
- Triggers don’t match your phrasing
- Anti-triggers are too aggressive
- Skill is marked as deprecated
- Skill is disabled in your configuration
Debugging steps:
- Check the skill’s description (is it specific enough?)
- Check explicit triggers (do any match your phrasing?)
- Check anti-triggers (are you accidentally excluding the skill?)
- Check skill status (is it disabled or archived?)
- Look for synonyms in your phrasing that might match other skills
- Test with slightly different wording to see if it’s a phrasing issue
Fix: Either adjust your phrasing to match the skill’s triggers, or update the skill’s triggers to be more inclusive.
Advanced: Optimizing Trigger Coverage for Discoverability
When you write a skill, how do you ensure it loads when people need it? The answer is trigger coverage analysis.
Step 1: Identify Natural Language Variations
For your skill, write down 15-20 ways users might naturally describe their problem:
Example: character-forge skill
Natural descriptions users might give:
1. "I need to develop my main character"
2. "How do I create a compelling protagonist?"
3. "What's driving this character's motivation?"
4. "She has a wound from her past. How does that shape her?"
5. "I need a character arc for her growth"
6. "Design her psychological profile"
7. "What's her fear and desire?"
8. "Build her wound-lie-truth progression"
9. "How do trauma and healing shape character?"
10. "Create a character backstory"
11. "What defines her as a person?"
12. "Design her internal conflict"
13. "Make her feel real, not generic"
14. "How does her past affect her present?"
15. "Character psychology and motivation"
Step 2: Map Phrases to Triggers
For each phrase, identify which trigger would match it:
Phrase → Trigger (✅ or ❌)
1. "develop my main character" → "character development" ✅
2. "compelling protagonist" → "protagonist" ✅ or semantic match
3. "driving motivation" → "character motivation" ✅
4. "wound from her past" → "wound" ✅ or "trauma" ✅
5. "character arc for growth" → "character arc" ✅
6. "psychological profile" → semantic match (description)
7. "fear and desire" → semantic match (character psychology)
8. "wound-lie-truth" → "wound-lie-truth" ✅
9. "trauma and healing" → "trauma" ✅
10. "backstory" → "backstory" ✅
11. "defines her as a person" → semantic match (character essence)
12. "internal conflict" → "character conflict" ✅
13. "feel real, not generic" → semantic match (realism)
14. "past affects present" → semantic match (character development)
15. "psychology and motivation" → semantic match (character psychology)
Step 3: Calculate Coverage
Coverage = (Matched phrases / Total phrases) × 100%
Matched: 13/15 = 86.7% coverage
Low-match phrases:
- "defines her as a person" (7, 11) → Could add "character essence"
- "trauma and healing" (9) → Could add "healing" or "redemption"
Action: Add triggers: ["character essence", "redemption arc"]
New coverage: 15/15 = 100%
Aim for 85%+ coverage for discoverable skills.
Step 4: Test Precision
Precision = (Relevant loads / Total loads) when this skill loads
For character-forge, measure:
- When it loads, is it actually relevant 95%+ of the time?
- Or does it load for unrelated “character” mentions (like “character encoding”)?
Use anti-triggers to prevent irrelevant loads:
triggers:
- "character development"
- "character arc"
exclude_patterns:
- "character encoding"
- "character class"
- "character set"
- "utf-8 character"
Step 5: Iterate and Refine
Track what phrases are missed:
- What did users say when the skill didn’t load?
- What triggers should have matched but didn’t?
- Are triggers too narrow or too broad?
Quarterly review: Update triggers based on actual usage patterns.
Performance Considerations: Triggers at Scale
When you’re managing 50+ skills, trigger matching adds overhead. Here’s how to keep it fast:
Expensive Triggers (Avoid)
# ❌ Too expensive: Regex on every request
triggers:
- pattern: "(?:character|protagonist|main figure) .*? (?:development|arc|growth|change)"
mode: regex
# ❌ Too expensive: Unbounded regex
triggers:
- pattern: ".*?(write|create|design|build).*?character.*"
mode: regex
Each regex runs against every task. 50 regexes × every task = slow system.
Efficient Triggers
# ✅ Fast: Explicit keywords
triggers:
- "character development"
- "character arc"
- "character growth"
# ✅ Fast: Limited regex with anchors
triggers:
- pattern: "^character \\w+ (wound|trauma|fear|desire|goal)$"
mode: regex
priority: high
Rule of thumb: Use keywords for 90% of triggers. Use regex only when keyword coverage is under 75%.
Logging and Observability: Debugging in Production
When skill loading misbehaves in production, how do you debug it?
Claude Code can log trigger matching decisions:
[SKILL MATCHING] Task: "Research AI agent architectures with 30 sources"
Evaluated skills: 142
Matched skills: 8
Details:
✅ deep-research: score=25 (trigger: "research 30+ sources", confidence: CRITICAL)
✅ research-synthesis: score=14 (description match, confidence: HIGH)
✅ writing-skills: score=5 (weak description match, confidence: LOW)
⚠️ github-api: score=2 (excluded: not relevant context, confidence: VERY LOW)
❌ character-forge: score=0 (no match, excluded: fiction context)
Selected for context window (ranked):
1. deep-research (CRITICAL confidence, 25 pts)
2. research-synthesis (HIGH confidence, 14 pts)
3. writing-skills (LOW confidence, 5 pts)
Discarded (insufficient score):
1. github-api (2 pts < threshold 8 pts)
This logging tells you:
- Which skills matched and why
- Why some skills didn’t match
- Scoring breakdown
- Final selection decision
Enable with: CLAUDE_DEBUG=triggers in your environment.
Advanced Trigger Patterns
Pattern 1: Negative Triggers (When NOT to Load)
Some skills need anti-triggers to avoid loading inappropriately:
---
name: github-api
description: "Use when making GitHub API calls..."
triggers:
- "github api"
- "gh api"
exclude_patterns:
- "github ui"
- "github browser"
- "github web interface"
---
If you say “I’m using the GitHub browser UI to…”, the system knows NOT to load github-api (because API isn’t relevant). Anti-triggers prevent false positives.
Pattern 2: Context-Aware Triggers
Some skills need to check surrounding context:
---
name: character-forge
triggers:
- context: "fiction|story|narrative"
pattern: "character"
priority: high
- context: "game|rpg|tabletop"
pattern: "character"
priority: medium
---
This skill loads when you mention “character” IN A STORY CONTEXT, but loads differently (lower priority) if you’re talking about game characters. Context-aware triggers help distinguish identical words in different domains.
Pattern 3: Trigger Synonyms with Priority
Some problems have multiple names:
---
name: dialogue-crafter
triggers:
- phrase: "dialogue"
priority: critical
- phrase: "conversation"
priority: high
- phrase: "speech"
priority: high
- phrase: "talking"
priority: medium
- phrase: "chat"
priority: low
---
“Dialogue” is the precise term (critical priority). “Conversation” and “speech” are common synonyms (high). “Talking” is vague (medium). “Chat” is too informal (low). This ensures the skill loads for precise language but also catches casual phrasing.
The Matching Algorithm: How It Works Under the Hood
Here’s the simplified algorithm Claude Code uses to match tasks to skills:
// Pseudo-code: Skill Trigger Matching Algorithm
function findMatchingSkills(userTask) {
const matches = [];
for (const skill of allSkills) {
let score = 0;
let confidence = 0;
// Layer 1: Exact keyword triggers
for (const trigger of skill.triggers) {
if (userTask.includes(trigger)) {
score += 10; // Highest score
confidence = "high";
}
}
// Layer 2: Regex pattern triggers
for (const pattern of skill.patterns) {
if (userTask.match(pattern.regex)) {
score += 8; // High score
confidence = "high";
score *= pattern.priority;
}
}
// Layer 3: Semantic description matching
const semanticScore = calculateSimilarity(userTask, skill.description);
score += semanticScore * 5; // Moderate weight
if (score > 0.7) confidence = "medium";
// Layer 4: Context matching
const contextMatch = checkContext(userTask, skill.context);
score *= contextMatch;
// Layer 5: Anti-triggers (reduce score)
for (const antiTrigger of skill.excludePatterns) {
if (userTask.includes(antiTrigger)) {
score *= 0.1; // Dramatically reduce
confidence = "low";
}
}
if (score > THRESHOLD) {
matches.push({
skill: skill.name,
score: score,
confidence: confidence,
});
}
}
// Sort by score and return top matches
return matches.sort((a, b) => b.score - a.score);
}
Key mechanisms:
- Additive scoring: Each trigger type adds points
- Multipliers: Priority flags multiply the score
- Confidence tracking: Why the match occurred
- Anti-triggers: Reduce scores to prevent mismatches
- Thresholding: Only return matches above a minimum score
The algorithm doesn’t load skills—it scores them. Claude Code’s system prompt then uses these scores to decide which skills to inject and in what order.
Trigger Best Practices
DO: Write Triggers for Problem Conditions
# ✅ GOOD: Describes the problem
triggers:
- "need to research deeply"
- "20+ sources required"
- "verify facts across sources"
DON’T: Write Triggers for Implementation
# ❌ BAD: Implementation detail
triggers:
- "websearch api call"
- "json pagination"
- "source tier system"
DO: Include Synonyms
# ✅ GOOD: Multiple ways to say it
triggers:
- "error handling"
- "exception handling"
- "error management"
- "failure handling"
DON’T: Make Triggers Too Long
# ❌ BAD: Too specific, won't match variations
triggers:
- "research on authentication patterns with multiple academic sources"
# ✅ GOOD: Concise, matches variations
triggers:
- "research authentication"
- "comprehensive research"
- "authentication patterns"
DO: Use Domain Vocabulary
# ✅ GOOD: Uses story domain terms
triggers:
- "hero's journey"
- "character wound"
- "save the cat"
- "plot turn"
DON’T: Use Vague Triggers
# ❌ BAD: Too vague, too many false positives
triggers:
- "help"
- "question"
- "task"
Common Trigger Edge Cases
Edge Case 1: Homonyms
“Character” means different things in different contexts:
# Fiction context:
triggers:
- "character development"
- "character arc"
- "protagonist"
# Coding context:
triggers:
- "character encoding"
- "character type"
These are different skills with overlapping triggers. The system handles this through context detection—reading surrounding words to determine which skill to load.
Edge Case 2: Acronyms
“API” could mean GitHub API, REST API, API design, etc.:
# github-api skill:
triggers:
- "github api"
- "gh api"
- context: "github|github-cli"
pattern: "api"
# api-design skill:
triggers:
- "api design"
- "rest api design"
- "api architecture"
The context filter disambiguates—”github api” loads github-api, “api design” loads api-design.
Edge Case 3: Typos and Misspellings
Should triggers account for typos? Generally no—rely on semantic matching:
description: "Use when designing RESTful APIs and web services"
# Don't add:
# triggers:
# - "REST API"
# - "REST Api"
# - "rest api"
# Semantic matching handles case variations
Semantic matching is case-insensitive and fuzzy. Explicit triggers should be the standard spelling.
Verifying Trigger Coverage
How do you know your skill’s triggers are comprehensive? Use this checklist:
- Collect 10+ phrases users might use for this skill
- Check coverage: Does at least one trigger match each phrase?
- Test anti-triggers: Are there phrases where this skill should NOT load?
- Measure precision: When this skill loads, is it actually relevant?
- Measure recall: When relevant, does this skill load?
Example:
Skill: character-forge
Phrases users might say:
1. "I need to develop this character" → trigger: "character development" ✅
2. "Design the protagonist's arc" → trigger: "character arc" ✅
3. "What's her motivation?" → semantic match (description) ✅
4. "Build her wound-lie-truth" → trigger: "wound-lie-truth" ✅
5. "Character development in games" → context filter (fiction, not game) ✅
6. "Create a character profile" → trigger: "character" ✅
7. "Character encoding for this string" → anti-trigger (avoid) ✅
8. "This character in her backstory" → semantic match ✅
9. "Personality and trauma" → description keywords ✅
10. "Her arc teaches the theme" → trigger: "character arc" ✅
Coverage: 10/10 phrases matched ✅
Trigger Evolution: Monitoring and Refining Over Time
Your triggers aren’t static. As you use a skill more, you learn what people actually ask for. The triggers you wrote initially are assumptions. Real usage data teaches you the truth.
Phase 1: Initial Triggers (Launch)
You write triggers based on what you think people will ask. These are educated guesses.
triggers:
- "character development"
- "character arc"
- "personality"
Phase 2: Monitor Usage (First Month)
Track what people actually say. Claude Code can log this:
[TRIGGER MISS] Task: "How do I build a more realistic protagonist?"
Evaluated: character-forge (score: 3, failed to load)
Should have: character-forge (semantic match on "realistic protagonist")
Phase 3: Refine Based on Reality (Quarterly)
Misses tell you what to add. This one missed “realistic protagonist” because your description doesn’t emphasize realism.
Fix 1: Add trigger
triggers:
- "realistic character"
- "authentic character"
Fix 2: Improve description
description: "Use when building realistic, psychologically authentic characters with depth and growth"
Phase 4: Monitor New Triggers
The new triggers might over-fire. Track false positives too:
[TRIGGER HIT] Task: "Is this a realistic portrayal of grief?"
Evaluated: character-forge (score: 8)
Actually needed: sensitivity-checker (emotional accuracy)
False positive. Add anti-trigger:
exclude_patterns:
- "is this realistic"
- "accuracy check"
Quarterly Trigger Audit
Every 3 months, review trigger performance:
Skill: character-forge
Last Quarter:
- 87 correct loads (semantic + keyword matches)
- 12 misses (tasks that should have loaded)
- 3 false positives (irrelevant loads)
Precision: 87 / (87 + 3) = 96.7% ✓
Recall: 87 / (87 + 12) = 87.9% (could improve)
Actions:
- Precision is good, keep current triggers
- Recall has room: add 3-4 triggers for missed cases
- Monitor false positives
This discipline separates great skill systems from mediocre ones.
Trigger Prioritization: When Everything Matches
Advanced scenario: What if your task matches 8 skills perfectly?
Task: "I need a character who's been traumatized and is healing. Show her growth through dialogue."
Matches:
1. character-forge (wound-lie-truth) - score 25 (CRITICAL)
2. dialogue-crafter (dialogue flow) - score 22 (HIGH)
3. description-painter (emotional atmosphere) - score 18 (HIGH)
4. prose-generator (write scene) - score 16 (HIGH)
5. pacing-analyzer (character pacing) - score 12 (MEDIUM)
Claude Code loads all five but prioritizes by score. The system prompt knows to weight skills as:
- Primary (25 pts): character-forge sets the direction
- Secondary (22-18 pts): dialogue-crafter and description-painter support
- Tertiary (12 pts): pacing-analyzer provides guidance
The context window fills like this:
- 40% character-forge (guides architecture)
- 30% dialogue-crafter (specific technique)
- 20% description-painter (sensory depth)
- 10% pacing-analyzer (rhythm)
This hierarchy happens automatically, but understanding it helps you predict how skills will interact.
Anti-Patterns: Common Trigger Mistakes
Anti-Pattern 1: Triggers That Are Too Broad
# ❌ BAD: Will match everything
triggers:
- "help"
- "code"
- "write"
Results: This skill loads for almost every task, drowning out more specific skills.
Fix: Be specific to your domain.
# ✅ GOOD: Specific to what you do
triggers:
- "character development"
- "psychological profile"
- "motivation arc"
Anti-Pattern 2: Triggers for Multiple Unrelated Things
# ❌ BAD: One skill trying to be everything
triggers:
- "character development"
- "api documentation"
- "test writing"
- "marketing copy"
This is a sign your skill is too broad. Split it into multiple skills.
Anti-Pattern 3: Copying Triggers from Other Skills
# ❌ BAD: Reused triggers from existing skill
triggers:
- "research"
- "deep research"
- "20+ sources"
If another skill already owns these triggers, you’ll create ambiguity. Use distinct vocabulary for your skill.
Anti-Pattern 4: Triggers in Past Tense or Gerunds
# ❌ UNCLEAR: Past tense
triggers:
- "researched"
- "wrote"
- "designed"
# ✅ GOOD: Present/infinitive
triggers:
- "research"
- "write"
- "design"
Users describe their current need, not what they’ve done.
Trigger Matching Across Languages
If Claude Code is multilingual (which it increasingly is), trigger matching gets interesting.
# English
triggers:
- "character development"
# Could also match:
# Spanish: "desarrollo de personajes"
# French: "développement des personnages"
# German: "Charakterentwicklung"
Claude Code can handle this via semantic matching (description-based), but explicit triggers are language-specific. Consider this when building international systems.
Trigger Performance at Scale: The 1000-Skill Problem
What happens when you have 1000 skills? (Unlikely but instructive.)
Trigger matching becomes expensive. The algorithm has to:
- Score 1000 skills (even if only 20 will match)
- Sort them all
- Pick the top N to inject
At scale, this gets slow. Solutions:
Solution 1: Trigger Indexing
Build an inverted index at startup:
"character" → [character-forge, character-dialogue, character-psychology, ...]
"research" → [deep-research, research-synthesis, ...]
"dialogue" → [dialogue-crafter, dialogue-validator, ...]
When you mention “character development”, skip evaluating 950 irrelevant skills.
Solution 2: Skill Categories
Cluster skills by domain:
Fiction Domain:
- character-forge
- dialogue-crafter
- prose-generator
- plot-structure
Engineering Domain:
- github-api
- testing-skills
- refactoring-guide
Research Domain:
- deep-research
- fact-checking
- synthesis-guide
Detect the domain first, then only match within that cluster.
Solution 3: Skill Deprecation
Delete skills nobody uses. Dead skills are context weight with no value.
Wrapping Up: Why Triggers Matter
Skill triggers are invisible infrastructure. When they work perfectly, you don’t notice them. The right skill loads at the right moment, and you wonder why it was ever any other way.
When they break, your tasks fail silently. A skill you need doesn’t load. An irrelevant skill hijacks the context. Hours vanish to debugging.
Understanding triggers gives you:
- Better skill writing: You know how to craft descriptions and keywords that actually work
- Faster debugging: When skills don’t load, you know where to look
- System thinking: You understand how Claude Code matches problems to solutions automatically
- Control: You can tune trigger priority to match your workflow
The Psychology of Trigger Writing
Here’s something subtle: good triggers are empathetic. They’re written from the user’s perspective, not the builder’s.
Builder perspective (what you’re thinking):
“This skill implements multi-source research with iterative refinement using websearch and semantic analysis.”
User perspective (what they’re thinking):
“I need to really understand this topic thoroughly, from multiple angles, with current information.”
Good triggers bridge that gap. They speak the user’s language:
# ❌ Builder language (technical)
description: "Implements multi-source research with iterative refinement"
triggers: ["websearch iteration", "semantic re-ranking", "source tier system"]
# ✅ User language (problem-focused)
description: "Use when you need comprehensive research with 20+ sources and multiple perspectives"
triggers: ["deep research", "comprehensive research", "understand this thoroughly", "research 20+ sources"]
The second set works because it assumes the user doesn’t care how you implement it. They care that you solve their problem.
Understanding Trigger Latency and Performance
One aspect of trigger systems that rarely gets discussed is performance. When you have 142 skills, matching triggers against every task creates latency costs. Understanding how this works helps you write triggers that are both effective and performant.
Claude Code implements a three-phase matching pipeline:
Phase 1: Keyword Matching (Fastest)
Explicit triggers are checked first with simple string matching. This is O(n) where n is the number of explicit triggers across all skills. With 142 skills, if each has 5 triggers, that’s 710 keyword checks. This completes in ~5ms.
Phase 2: Regex Pattern Matching (Moderate)
Regex patterns are compiled once and reused. Only skills with regex triggers are evaluated. This typically adds 10-20ms.
Phase 3: Semantic Matching (Slowest)
Only if phases 1 and 2 don’t have high-confidence matches does the system do semantic analysis. This involves embeddings and similarity scoring—the slowest operation. This typically adds 100-300ms.
In practice, most tasks should match in Phase 1 or 2. If your task consistently falls through to Phase 3, the system slows down. This is why explicit triggers are valued: they’re fast.
When writing triggers, think about the user’s task description:
- If they’re likely to use specific terminology, add those as explicit triggers
- If they’ll describe the problem in varied ways, rely on semantic description
- Use regex only for complex patterns that keywords can’t capture
A well-designed trigger set should be 80% explicit keywords, 15% semantic matching, 5% regex. This ensures fast, reliable matching.
Trigger Conflicts and Resolution Strategies
In any large trigger system, conflicts are inevitable. Two skills might legitimately match the same task. Understanding how Claude Code resolves these prevents surprise behavior.
Type 1: Complementary Overlap
Some conflicts are actually features, not bugs. Example: A task about “designing a character’s arc in a TypeScript game engine” could match:
character-forge(story skill)typescript-best-practices(engineering skill)
Both are relevant. The user might benefit from both. Claude Code loads BOTH because they serve different purposes. Character development advice is orthogonal to TypeScript advice.
Type 2: Competing Overlap
Other conflicts represent genuine overlap where only one skill should load. Example: A task about “researching authentication patterns” could match:
deep-research(comprehensive, 20+ sources approach)github-api(GitHub-specific, API-focused research)
If the user said “I need to understand OAuth implementations across 30 projects,” that’s deep-research. If they said “I need to check GitHub’s API authentication documentation,” that’s github-api. The resolution depends on task context.
Claude Code’s conflict resolution uses a contextual scoring system:
Score(skill, task) =
(trigger_match_weight × trigger_specificity) +
(semantic_similarity × description_match) +
(historical_frequency × past_success_rate) +
(skill_version × recency_bonus)
Example Calculation:
Task: “Research GitHub authentication best practices across 20+ security papers”
| Skill | Trigger Match | Semantic | Historical | Score | Load? |
|---|---|---|---|---|---|
| deep-research | 0.9 (keyword) | 0.85 | 0.92 | 0.87 | YES |
| github-api | 0.6 (partial) | 0.72 | 0.78 | 0.70 | YES |
| security-audit | 0.5 (semantic) | 0.81 | 0.65 | 0.65 | NO |
Both deep-research and github-api load because both scores exceed the threshold (0.70). But deep-research has the higher score, so it’s labeled as PRIMARY, and github-api as COMPLEMENTARY. The user can see both but knows which one is the main focus.
Trigger Anti-Patterns: Common Mistakes
Writing effective triggers is as much about knowing what NOT to do as what to do. Here are common mistakes:
Anti-Pattern 1: Implementation Details in Triggers
# ❌ BAD: Implementation detail, user won't say this
triggers: ["spawn subagent", "recursive skill calling", "context window optimization"]
# ✅ GOOD: User problem statement
triggers: ["complex multi-step task", "research that requires iterations", "adaptive problem-solving"]
Users don’t think about implementation. They think about their problem. Write triggers that match problems, not solutions.
Anti-Pattern 2: Overly Broad Triggers
# ❌ BAD: Too broad, will trigger on unrelated tasks
triggers: ["help", "assist", "research"]
# ✅ GOOD: Specific enough to be meaningful
triggers: ["deep research", "comprehensive investigation", "20+ sources", "multi-perspective analysis"]
“Help” is too broad. Every task involves help. “Deep research” is specific—it identifies a particular type of task.
Anti-Pattern 3: Assuming Exact Terminology
# ❌ BAD: Assumes users will use exact term
triggers: ["hero-journey-monomyth-analysis"]
# ✅ GOOD: Multiple ways to express the same idea
triggers: ["hero's journey", "monomyth", "hero journey", "campbell structure", "mythic structure"]
Users have different vocabularies. Joseph Campbell called it the “monomyth.” Others call it “hero’s journey.” Some say “Campbell’s structure.” All should load the same skill.
Anti-Pattern 4: Triggers That Mislead
# ❌ BAD: Suggests skill does something it doesn't
triggers: ["fix performance", "optimize database", "speed up queries"]
description: "General SQL syntax reference"
# ✅ GOOD: Accurate representation of skill scope
triggers: ["sql syntax", "sql query writing", "sql fundamentals"]
description: "SQL syntax reference and query construction basics"
If your skill is about syntax, don’t add triggers about performance optimization. Users will load it expecting performance help and be disappointed.
These anti-patterns create frustration: skills loading when they shouldn’t, skills not loading when they should, or skills loading but being unhelpful.
The Testing Principle: Validate Trigger Behavior
Before shipping a skill, test its triggers:
Test 1: Coverage Test
Ask 10 people (or yourself) how they’d describe this skill’s purpose. Do any of your triggers match at least 8 of their phrasings? If not, your coverage is weak.
Test 2: Precision Test
For each trigger, ask: “If someone uses this exact phrase, is this skill always relevant?” If the answer is “sometimes” or “maybe”, your trigger is too broad.
Test 3: False Positive Test
What would trigger this skill that shouldn’t? Example: “character development in Python” shouldn’t load the story-writing character-forge skill.
Fix with anti-triggers:
exclude_patterns:
- "character.*python|python.*character"
- "character.*encoding|encoding.*character"
- "character.*class|class.*character"
Test 4: Synonym Test
List 5 different ways someone might describe this skill’s purpose. Can triggers (or semantic matching) catch at least 4 of them?
These four tests prevent shipping skills that fail to load when needed.
Measuring Trigger Effectiveness: Metrics That Matter
When you deploy skills with triggers, how do you know if they’re working? You need metrics. Claude Code tracks several dimensions of trigger effectiveness automatically:
Precision: Of all the times this skill loaded, what percentage of loads were actually relevant to the task?
- High precision (over 95%): Users almost always find the skill helpful when it loads
- Medium precision (70-95%): Useful, but sometimes loads when user doesn’t need it
- Low precision (under 70%): Loads too often for wrong reasons; needs adjustment
Recall: Of all the times this skill was relevant, what percentage of those times did it actually load?
- High recall (over 90%): Almost never missed a relevant task
- Medium recall (70-90%): Usually catches relevant tasks, but sometimes misses
- Low recall (under 70%): Misses relevant tasks too often; triggers need expansion
Latency: How long does the trigger matching take for this skill?
- under 10 ms: Excellent (keyword-based triggers)
- 10-50ms: Good (some semantic matching)
- 50-100ms: Acceptable (semantic-heavy)
-
100ms: Problematic (needs optimization)
Historical Success Rate: When this skill loads, how often does the user find the output helpful?
-
80% helpful: Skill is well-designed and well-triggered
- 60-80% helpful: Useful but could improve
- under 60% helpful: Either skill quality or trigger match is poor
In a healthy system, you’re aiming for:
- Precision: 85%+
- Recall: 80%+
- Latency: under 25 ms
- Usefulness: 70%+
If any metric falls below target, you need to investigate. Low precision means your triggers are too broad. Low recall means they’re too narrow. High latency means you have too many regex patterns. Low usefulness means the skill itself might need improvement, or it’s being loaded for the wrong reasons.
Real-World Case Study: The Deep-Research Skill
Let’s trace through a real deployment story. The deep-research skill was created with these initial triggers:
triggers:
- "deep research"
- "comprehensive research"
- "research 20+ sources"
- "systematic research"
- "thorough investigation"
After two weeks of real usage, metrics revealed:
- Precision: 92% ✅ Good
- Recall: 61% ⚠️ Low
- Latency: 8ms ✅ Good
- Usefulness: 78% ✅ Good
The low recall was the problem. Users were describing research needs in ways not captured by the triggers. Actual missed cases:
- “I need sources to write about authentication”
- “Research this claim across multiple perspectives”
- “Find 20+ papers on machine learning bias”
- “Help me understand the consensus on AI safety”
None of these use the exact trigger words. But all were legitimate deep-research tasks. The semantic description should have caught them, but the semantic similarity wasn’t matching well.
Solution: Add more explicit triggers based on the missed cases:
triggers:
# Original triggers
- "deep research"
- "comprehensive research"
- "research 20+ sources"
- "systematic research"
- "thorough investigation"
# New triggers from misses
- "multiple sources"
- "multiple perspectives"
- "research and write"
- "understand the consensus"
- "claim verification"
- "source selection"
After adding these, recall improved to 87% while precision stayed at 91%. The skill was now catching most relevant tasks.
This is typical of real deployments: your first iteration of triggers is 80% right. The remaining 20% emerges from actual usage patterns you didn’t anticipate. This is why monitoring matters.
Trigger Optimization Strategies
As your skill library grows and usage accumulates, you’ll want to optimize trigger behavior. Here are proven strategies:
Strategy 1: Hierarchical Triggers for Similar Skills
If you have five research-related skills (deep-research, github-research, paper-research, expert-interview-research, competitive-research), you need a hierarchy:
User says: "Research competitor pricing strategies"
Load order:
1. Try domain-specific triggers (competitive-research)
→ Found match → Load
2. If no domain match, try task-specific triggers (expert-interview-research)
→ Partial match but lower score
3. Fallback to general (deep-research)
→ Would match but secondary
This prevents the general skill from overshadowing specialized variants.
Strategy 2: Negative Triggers for Disambiguation
Some skills need explicit “NOT” conditions:
name: github-api-research
triggers:
- "github research"
- "github queries"
exclude_when:
- "github action" # Use github-actions skill instead
- "github workflow" # Use github-actions skill instead
- "github cli" # Use github-cli skill instead
Negative triggers prevent false positives where semantically similar tasks should load a different skill.
Strategy 3: Context-Aware Triggers
For skills used across multiple domains, context matters:
name: writing-skills
triggers:
# Default context
- "write documentation"
- "technical writing"
# Story context
- when_context: story
triggers:
- "write scene"
- "narrative structure"
# API context
- when_context: engineering
triggers:
- "api documentation"
- "code comments"
Context can come from conversation history, project type, or explicit user indication.
Strategy 4: Time-Based Trigger Adjustments
Some skills are more relevant at certain times:
name: sprint-planning
triggers:
# Always active
- "sprint planning"
# Higher priority on Fridays (common sprint planning day)
- trigger: "team planning"
when: "day == friday"
priority_boost: 0.2
This doesn’t prevent loading on other days, but increases relevance scoring on typical sprint days.
Trigger Maintenance: The Unglamorous Work
After launch, triggers need maintenance. Here’s a realistic maintenance schedule:
Weekly: Monitor misses and false positives
- Did anyone try to use the skill but it didn’t load?
- Did the skill load when it wasn’t relevant?
- Document patterns
Monthly: Update based on patterns
- Add 2-3 new triggers for common miss patterns
- Refine description based on user language
- Test new triggers don’t create false positives
Quarterly: Full audit
- Measure precision and recall
- Review anti-triggers (still needed?)
- Consider splitting skill if triggers are getting unwieldy
Annually: Strategic review
- Is this skill still relevant?
- Are triggers aligned with current usage?
- Should we merge this with another skill?
Maintenance takes 30 minutes per skill per quarter. It’s small but essential.
Teaching Users About Your Skill
Here’s what most developers forget: Users won’t discover your skill unless you tell them about it.
Great trigger discovery happens when:
- Documentation is clear about when to use the skill
- Examples show natural language (not technical jargon)
- Error messages suggest the skill when they describe its purpose
Example error message:
Q: "How do I develop a deep character?"
[No matching skill found]
Suggestion: You might find this helpful:
→ character-forge: "Use when building realistic, psychologically
authentic characters with depth and growth"
Use it with: /skill character-forge
This teaches users your skill’s triggers. Over time, they’ll start using the right language naturally.
Next time you write a skill, remember: Your description isn’t just documentation. It’s a contract between you and the trigger system. Make it specific. Make it honest. Make it obvious.
Because the best skill in the world won’t help if the system never loads it.
-iNet
Master skill triggers, and you’ve unlocked the real power of Claude Code’s automatic capability discovery.