All Articles Claude Code

Skill Composition: Combining Multiple Skills in One Workflow

You're staring at a feature request that spans multiple concerns. You need a skill that validates code, runs tests, generates documentation, and commits changes.

You’re staring at a feature request that spans multiple concerns. You need a skill that validates code, runs tests, generates documentation, and commits changes. Individually, you’ve got skills for each. But how do you wire them together without creating a tangled mess of instruction conflicts and ordering problems?

This is the hidden art of skill composition—and it’s where Claude Code separates “I can write skills” from “I architect resilient workflows.” We’re not talking theory here. You’re about to learn the exact patterns that power the most complex Claude Code workflows: how to invoke skills in sequence, how to manage dependencies, how to reference one skill from another, and critically, how to avoid the instruction conflicts that tank real-world automations.

By the end of this article, you’ll have concrete, copyable patterns for building workflow skills from atomic skills—and you’ll understand why order matters, why isolation matters, and why documenting composition dependencies is non-negotiable. You’ll also understand the deep principles behind orchestration that apply far beyond Claude Code: how systems scale through composition, why explicit dependencies prevent silent failures, and how to think about complexity at higher levels of abstraction.

The Problem: Why Skill Composition Matters (And Why It’s Hard)

Let’s ground this in reality. You’re implementing a feature that touches multiple domains:

  1. Code validation: Lint, format checks, static analysis
  2. Testing: Unit tests, integration tests, coverage verification
  3. Documentation: Auto-generate README sections, API docs
  4. Git operations: Stage changes, create commits, push to PR

Each of these is a distinct responsibility. You could write one giant skill that does all four. Your skill would be 500+ lines. It would be untestable. A change to testing logic would require re-reviewing all the git code. Debugging would be a nightmare because you wouldn’t know which of four independent systems failed.

Or, you could compose four atomic skills—one per concern—and orchestrate them in a workflow skill. Now each skill is testable. Changes are isolated. Logic is reusable across workflows. This is the difference between maintenance burden that grows quadratically with code size versus linearly.

The catch: Composition isn’t free. You need to carefully manage several concerns. You need to enforce the correct order—tests must run after validation, documentation only makes sense if tests passed. You need to prevent conflicts where one skill’s instructions contradict another’s. You need to handle dependencies—documentation needs code to exist first, and it needs to know the test results. You need a way to pass data from one skill to the next, because the output of validation informs what testing needs to do. And you need to handle failures gracefully—if tests fail, what happens to documentation? Do we skip it? Do we mark it as incomplete?

This is why understanding composition patterns is critical. It separates scripts from systems. A script does one thing. A system is composed of many things that work together predictably.

Why Monolithic Skills Fail at Scale

When you pack everything into a single skill, the problems multiply rapidly. Testing becomes impossible—you can’t verify one concern without running the entire workflow, which means a small change to the validation logic requires running all tests and regenerating all documentation just to verify it works. Debugging multiplies in complexity. An error could come from any of four domains, and tracing through 500 lines of mixed logic with branching paths requires exhaustive analysis. You can’t tell if the problem is in validation or in git operations.

Reuse becomes impossible. If another workflow needs just the “validation” piece, you’re copy-pasting code. When you find a bug in that validation logic, you have to fix it in two places and pray they stay in sync. This introduces subtle divergence over time. Eventually you have three copies of validation logic, and they’re subtly different in ways nobody remembers.

Maintenance becomes a nightmare. Changing validation logic requires re-reviewing and re-testing everything else, even though you didn’t touch it. Your deployment risk increases. What was a simple change now requires full regression testing because you altered the monolith. You start being afraid to touch the skill because you don’t understand all the interactions.

Scaling hits a wall. At 1000+ lines, the skill becomes unmaintainable. New team members can’t understand it. New features require careful surgical edits. Refactoring is impossible because you can’t change anything without breaking something else. The skill becomes a legacy system—feared, fragile, and avoided.

The alternative is composition: small, focused skills that do one thing well, orchestrated by a master skill that coordinates them. This approach has transformed how software teams work for decades, and it applies equally to Claude Code skill orchestration.

Understanding Composition Dependencies: The Mental Model

Before we write code, let’s establish the mental model. A composition is a directed graph of skills where edges represent dependencies. Skill A depends on Skill B if B must run before A, or if A needs B’s output.

Validation → Testing → Documentation → Git Commit
     ↓
   (any failure stops downstream)

This linear dependency chain is the simplest composition. Validation runs first. If it passes, testing runs. If testing passes, documentation runs. If documentation succeeds, we commit. If anything fails, everything downstream is skipped or cancelled.

But dependencies get more complex. Maybe documentation doesn’t need testing to succeed—it can generate even if tests fail, just with a note saying “tests failed.” Maybe validation and security scanning are independent and can run in parallel. Maybe documentation and testing are independent and can run in parallel, but both need validation to complete first.

Understanding these dependencies prevents instruction conflicts. If you tell the testing skill “assume validation passed” but validation actually failed, the testing skill has bad assumptions and produces garbage. If you tell both testing and documentation skills to update the same file, they conflict.

Explicit dependency documentation prevents these conflicts. Write down: “Testing requires validation to pass. Documentation requires validation to pass but not testing to pass. Git operations require both testing and documentation to succeed.” This clarity prevents misunderstandings.

Building Your First Composition: A Release Workflow

Let’s build a practical composition: a release workflow that validates code, runs tests, generates documentation, and commits changes.

# .claude/workflows/release-workflow.yaml
name: "Release Workflow"
description: "Validate, test, document, and commit changes"

stages:
  - name: "Validation"
    skill: "code-validator"
    required: true
    timeout: "5m"
    on_failure: "stop"

  - name: "Testing"
    skill: "test-runner"
    required: true
    depends_on: ["Validation"]
    timeout: "10m"
    on_failure: "stop"
    env:
      RUN_EXTENDED_TESTS: true

  - name: "Documentation"
    skill: "doc-generator"
    required: false
    depends_on: ["Validation"]
    timeout: "5m"
    on_failure: "warn"
    env:
      STYLE_GUIDE: "api-reference"

  - name: "Git Commit"
    skill: "git-release-commit"
    required: true
    depends_on: ["Validation", "Testing"]
    timeout: "2m"
    on_failure: "stop"
    env:
      AUTO_PUSH: false

This YAML defines the composition. It says:
– Run Validation first (required, stop on failure)
– Run Testing after Validation succeeds (required, stop on failure)
– Run Documentation after Validation succeeds, but only if it was requested (optional, warn on failure)
– Run Git Commit after both Validation and Testing succeed (required, stop on failure)

Each stage specifies its skill, dependencies, timeout, and failure handling. This clarity lets the orchestrator understand the workflow and execute it correctly.

Managing State Between Skills: Data Contracts

Skills in a composition must communicate. The output of one skill becomes input to the next. This requires explicit contracts.

// .claude/skills/code-validator.mjs
export const outputContract = {
  type: "validation-result",
  version: "1.0",
  schema: {
    success: boolean,
    errors: [{
      file: string,
      line: number,
      message: string,
      severity: "error" | "warning"
    }],
    warnings: number,
    summary: string,
    timestamp: ISO8601
  }
};

export const handler = async (context) => {
  // ... validation logic ...

  return {
    success: errors.length === 0,
    errors,
    warnings: warnings.length,
    summary: `Found ${errors.length} errors, ${warnings.length} warnings`,
    timestamp: new Date().toISOString()
  };
};

And then the testing skill knows exactly what to expect:

// .claude/skills/test-runner.mjs
export const inputContract = {
  type: "validation-result",
  version: "1.0"
};

export const handler = async (context) => {
  const validationResult = context.previousStagOutput;

  if (!validationResult.success) {
    throw new Error("Cannot run tests: validation failed");
  }

  // Run tests...
};

Data contracts prevent misunderstandings. They make explicit what each skill expects and what it produces. When you change one skill’s output, the contract tells you which other skills are affected.

Handling Failures: The Waterfall and the Safety Net

Different stages fail for different reasons. Some failures are catastrophic (validation failed, stop everything). Some are warnings (documentation failed, but we still proceed to commit). Some are retryable (network timeout, try again).

Design failure handling intentionally:

stages:
  - name: "Validation"
    on_failure: "stop"  # Stop immediately

  - name: "Testing"
    on_failure: "stop"  # Stop immediately

  - name: "Documentation"
    on_failure: "warn"  # Log warning, continue

  - name: "Security"
    on_failure: "block"  # Continue but mark as requiring manual review

  - name: "Git Commit"
    on_failure: "stop"  # Stop immediately

“stop” means the entire composition fails and downstream stages don’t run. “warn” means log the failure but continue to the next stage. “block” means continue but flag the composition as needing human review before the final action (e.g., deployment).

These states let you build workflows where some failures are fatal and others are acceptable. A documentation generation failure shouldn’t block a release. A validation failure should.

Preventing Instruction Conflicts: Isolation and Clarity

When multiple skills run in a composition, their instructions can conflict. For example, if validation says “fix all warnings” and testing says “ignore warnings in test code,” what happens? Which takes priority?

Prevent conflicts through isolation:

// Each skill gets isolated context
const handler = async (context) => {
  // context.skillName: "code-validator"
  // context.config: skill-specific configuration
  // context.previousStageOutput: output from previous stages
  // context.environment: workflow-level env vars

  // This skill doesn't see other skills' instructions
  // It only knows its own responsibility
};

And through clarity in documentation:

# Code Validator Skill

## Responsibility
Lint, format, and static analysis checks.

## What This Skill Does
- Run ESLint for JavaScript files
- Run Prettier format check
- Run TypeScript type checking

## What This Skill Does NOT Do
- Run tests (that's the test-runner skill)
- Generate documentation (that's the doc-generator skill)
- Commit changes (that's the git-release-commit skill)

## Configuration
- `--fix`: Automatically fix formatting issues (default: false)
- `--strict`: Treat warnings as errors (default: false)
- `--ignore-test-files`: Skip linting test files (default: false)

## Output
Returns validation result with error and warning counts.

Clear documentation about responsibility prevents instructions from being applied to the wrong skill.

Sequential vs. Parallel Orchestration: When to Use Each

Simple compositions run sequentially: A → B → C. Skill A finishes, then B starts, then C starts. This is easy to reason about.

Complex compositions run parts in parallel. Maybe A and B are independent, so run them together. C waits for both, then runs.

   → A →
  /       \
Start      C → End
  \       /
   → B →

Parallel execution saves time. If A takes 30 seconds and B takes 20 seconds, running them together takes 30 seconds (the max), not 50 seconds (the sum).

But parallel execution requires careful synchronization. You need to:
1. Identify independent stages (A and B don’t depend on each other)
2. Wait for all parallel stages to complete before moving to dependent stages
3. Collect all outputs from parallel stages for downstream skills

Use parallel orchestration when:
– The time savings matter (several seconds or more per composition)
– The stages are clearly independent
– You understand the synchronization mechanics

Use sequential orchestration when:
– You’re getting started with composition (keep it simple)
– Stages are mostly sequential anyway
– Clarity matters more than speed

Premature parallelization is a common mistake. Start sequential. When profiling shows parallelization helps, add it.

Real-World Scenarios: Where Composition Prevents Disaster

Understanding composition conceptually is one thing. Seeing it prevent actual failures is more convincing.

Scenario 1: The Validation That Should Have Blocked Testing

A team was using a monolithic skill that validated, tested, and committed. Someone modified the validation instructions while testing was in progress. The new validation logic had stricter rules. By the time validation finished, it produced a list of 47 errors. But the skill continued to testing, ran tests on code that had obvious errors, tests failed for the wrong reasons, nobody understood why tests were failing, and the whole release was confused.

With composition, this is impossible. Validation is a separate skill. Its failure status is passed to testing. The composition checks: “Did validation succeed?” No. “Run testing?” No, because testing depends on validation succeeding. The composition stops. Testing never runs. The error is clear: “Validation failed. Review and fix before testing.”

Composition enforces the contract. Validation must pass before testing runs. This clarity prevents cascading failures.

Scenario 2: The Documentation That Broke Releases

A team had a massive monolithic skill that validated, tested, documented, and committed. Documentation generation would occasionally fail due to missing code comments or API signature changes. When documentation failed, the whole skill failed, and nobody got a release because documentation couldn’t generate a README.

With composition, documentation is separate. If it fails, a warning is logged, but testing and validation succeeded, so the code is safe to release. The release proceeds. The documentation failure is addressed in a follow-up. The team ships on schedule instead of being blocked by a documentation tool failure.

Composition enables fault isolation. Failures in one area don’t propagate to unrelated areas.

Scenario 3: The Parallel Execution That Cut Release Time in Half

A team had sequential validation and security scanning taking 15 minutes total (8 minutes validation, 7 minutes security). They discovered these stages were independent—security didn’t need validation to pass first, and validation didn’t need security results. By running them in parallel, total time dropped to 8 minutes (the max of the two). The release time improved 46%.

This optimization was possible because composition made the dependencies explicit. By reading the dependency graph, they could see what could run in parallel. Without composition, dependencies would be implicit and scattered through the monolithic code, making this optimization impossible to discover safely.

Debugging Strategies: Fixing Broken Compositions

When a composition breaks in production, you need a systematic debugging approach. The challenge is that the failure might be in any stage, or in the data flow between stages, or in the orchestrator itself. Without structure, debugging can consume enormous amounts of time as you go down rabbit holes investigating the wrong problems.

Step 1: Identify the failing stage. Look at the composition report. Which stage failed? This narrows the scope significantly. Was it validation? Testing? Documentation? Git operations? The composition report should tell you exactly which stage failed and what the error message was.

Step 2: Run the stage in isolation. Can you reproduce the failure by running just that stage with the same input? If yes, the bug is in the stage itself. If no, the bug is upstream—the input to this stage is wrong. This distinction is crucial because it tells you where to focus your investigation.

Step 3: Check the previous stage’s output. If the stage is failing because of bad input, examine what the previous stage produced. Does it match the contract? Is there a version mismatch? Look at the actual output structure, not just the summary. Maybe the previous stage produced results in a different format than expected.

Step 4: Check the data contract. Read the previous stage’s documentation carefully. What’s the schema of its output? Does the actual output match? If it doesn’t, either the skill changed and documentation wasn’t updated, or there’s a skill bug. Data contract violations are the most common cause of composition failures.

Step 5: Create a minimal reproduction. With real input from the actual failure, create a minimal test case that reproduces the problem. Minimal means: smallest possible input that triggers the failure. A test case you can run repeatedly to verify fixes. This gives you something to verify fixes against and builds confidence that you’ve really solved the problem.

This systematic approach prevents the trap of debugging the wrong thing and wasting hours chasing ghosts in the wrong part of the system. By moving through the steps methodically, you either identify the problem quickly or eliminate possibilities systematically until the answer becomes obvious.

Teaching Others to Design Compositions: Building Institutional Knowledge

As you build more compositions, you need to teach others how to design them. This knowledge transfer is crucial for scaling beyond the original architects.

Key principles to teach:

Start with atomic skills: Before composing, have atomic skills that each do one thing well. These are tested independently. They’re reusable. Composition builds on a solid foundation.

Document dependencies explicitly: Write them down. Don’t rely on tribal knowledge or implicit assumptions. Explicit documentation survives people leaving the team.

Think in stages, not steps: A stage groups related operations. “Deploy to staging” is a stage. Within it, multiple steps might run. But from the composition perspective, it’s a single unit with clear success/failure criteria.

Plan for failure: When designing a composition, think about failure modes. What happens when this stage fails? Does everything stop? Do we skip downstream stages? Is it a warning or a blocker? Design failure handling intentionally.

Measure before optimizing: Parallelize stages that are actually slow, not theoretically slow. Profile real compositions with realistic data. Optimize bottlenecks, not guesses.

Version everything: Skills have versions. Compositions have versions. Outputs have schemas with versions. When things change, versions help manage compatibility.

Teaching these principles is how you scale composition expertise from one person to a team to an organization.

Advanced Composition Patterns: Moving Beyond the Basics

Once you’ve mastered sequential orchestration, there are more sophisticated patterns that unlock even greater power. These patterns are what separate teams that use skills as tools from teams that build skill ecosystems. Understanding when to use each pattern is the art of composition architecture.

The first advanced pattern is conditional orchestration. Not every workflow follows the same path. Maybe your release workflow is different for hotfixes versus regular releases. Maybe your code review process is different for documentation changes versus production code. Conditional orchestration branches the workflow based on context, running different skills or different configurations depending on the input. This adds complexity, but it enables much more powerful automation because you’re not forcing all variations through the same path.

Implementing conditionals requires decision points in your composition. After validation, you might ask: “Is this a hotfix?” If yes, run a faster, abbreviated testing process. If no, run the full test suite. “Does this PR modify database schema?” If yes, add extra database review steps. If no, skip them. Conditionals let you optimize workflows for different scenarios, reducing unnecessary work and improving cycle time.

Another advanced pattern is parallel execution with synchronization. When multiple skills are independent, they can run in parallel. Lint checks don’t depend on test results. Security scanning doesn’t depend on documentation generation. Running these in parallel cuts wall-clock time dramatically. This is where you get compounding benefits—not just from each skill being fast, but from them not blocking each other.

Implementing parallelization requires understanding which skills are truly independent. You need to read their dependency documentation, verify they don’t share mutable state, and confirm the outputs don’t contradict. Then you need synchronization mechanisms to wait for all parallel skills to complete before moving to dependent stages. Modern composition frameworks handle this, but you need to understand the semantics.

Retry and recovery patterns are critical for production systems. What happens when a skill fails intermittently? Implement sophisticated retry logic where network timeouts trigger immediate retry (the failure was transient), quota exceeded errors trigger delayed retry with backoff (the system needs time to recover), and permission denied errors fail immediately without retrying (the failure is permanent). This makes workflows resilient to transient failures while failing fast for permanent errors.

The most advanced pattern is composition of compositions. You have a workflow skill that orchestrates atomic skills. Another workflow skill orchestrates different atomic skills. A master orchestrator coordinates the workflow skills. This multi-level composition lets you build hierarchical systems that scale to arbitrary complexity. A microservice orchestrates internal services. A platform orchestrates microservices. A meta-platform orchestrates multiple platforms. Each level handles its own complexity, delegating details to lower levels.

The Cultural Impact of Good Composition

Good composition patterns affect team culture. When skills are cleanly separated and well-composed, the codebase feels approachable. A new developer can understand one skill deeply without understanding the entire system. They can contribute to one skill without touching everything else. This builds confidence and enables faster onboarding.

Good composition patterns also enable ownership. When skills are atomic and responsibilities are clear, you can assign ownership cleanly. Alice owns the code-validation skill. Bob owns the test-runner skill. Each owns a piece they understand deeply. Each is accountable for their piece. This clarity prevents diffused responsibility where nobody knows who owns what.

Finally, good composition enables experimentation. Want to try a new testing framework? Change the test-runner skill without touching anything else. Want to experiment with a different documentation generator? Swap in a new doc-generator skill. Because skills are isolated, experimentation is safe. Changes are local. You’re not risking the entire system.

This cultural shift—from monolithic code that people fear to compositional code that people confidently modify and extend—is one of the biggest benefits of composition patterns. It makes your codebase a living system that evolves rather than a legacy system that ossifies.


-iNet

Composable skills are the future. Start building them now.

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.