All Articles Claude Code

Claude Code Bundled Skills: /batch, /simplify, and /loop

You've got a codebase. It's 15,000 lines of Python spread across 47 files. Your manager asks: "Can you refactor the logging across all of them?" Your team has 30 config files in YAML that need...

You’ve got a codebase. It’s 15,000 lines of Python spread across 47 files. Your manager asks: “Can you refactor the logging across all of them?” Your team has 30 config files in YAML that need consistent formatting. Your git history is messy, and you’re drowning in technical debt that compounds every sprint.

This is where you stop thinking like you’re working with a text editor and start thinking like you’re working with a developer who can see everything at once.

Claude Code ships three bundled skills that solve the exact problems you face in real projects: /batch for multi-file processing, /simplify for intelligent code refactoring, and /loop for iterative improvement on intervals. They’re not magic wands. They’re industrial-grade tools designed to compress months of tedious work into hours.

Here’s what you need to know: these skills are customizable, chainable, and built on evidence-driven gates. You’re not hoping code is better—you’re verifying it empirically. That changes everything.

The Hidden Cost of Single-File Thinking

Before we dive into the skills themselves, let’s talk about why they matter.

When you work with traditional development tools—or even most AI assistants—you’re constrained by a single file, one context window at a time. You can’t see the ripple effects. You can’t spot duplicate patterns across your entire codebase. You can’t refactor systematically because you lack the aerial view.

A real example: Your team has a logging pattern that evolved over three years. Some files use print(), some use logging.debug(), some call a homegrown logger.warn() function. They all work. None of them are wrong. But together, they’re a nightmare to debug and maintain. If you refactor one file, you break consistency. If you try to do them all manually, you’re looking at a week of tedious work where mistakes are invisible until production.

This is where the gap between “working code” and “professional code” lives.

Claude Code’s bundled skills close that gap. They’re built to see your entire codebase, understand patterns across files, and make consistent changes with evidence that everything still works.

Why This Matters: The Leverage Multiplier

Here’s the hidden truth about developer productivity: most refactoring never happens. Not because developers are lazy. Because it’s tedious and risky. You’d need to:

  1. Review every file to understand what needs changing
  2. Make changes carefully (muscle memory is error-prone at scale)
  3. Test each change individually
  4. Document what changed and why
  5. Get code review approval
  6. Deploy and monitor for regressions

By step three, you’ve already spent two hours. A task that should take 20 minutes has expanded into an afternoon. So you don’t do it. The technical debt stays. Code quality drifts.

These bundled skills don’t make refactoring slightly faster. They make it fast enough to do every week. That’s the leverage multiplier. Instead of “refactoring is a quarterly project,” it becomes “refactoring is automatic hygiene.”

/batch: Process Multiple Files in Parallel

The /batch skill is your answer to “I need to do this same thing across 47 files.”

What /batch Does

/batch takes a task and applies it consistently to a set of files. You specify:

  • A glob pattern (which files to process)
  • The transformation you want (what to change)
  • Validation rules (how to verify it worked)

Claude Code processes files in parallel, applies consistent logic, and verifies each change against your validation gates. You get a summary report showing what changed, what stayed the same, and what failed.

Real-World Example: Standardizing Imports

Let’s say you’re consolidating your data science dependencies. You want to rename all import numpy as np to from numpy import array, where, sum as np_sum (because your team has standardized to functional imports).

This would be manual hell across 50+ files. With /batch:

/batch \
  --glob="src/**/*.py" \
  --task="Replace 'import numpy as np' with functional imports; update all np.array() calls to array()" \
  --validate="pytest -xvs tests/ --tb=short" \
  --evidence="Show me the diff for each changed file"

Claude Code will:

  1. Scan all Python files matching src/**/*.py
  2. Update import statements and function calls consistently
  3. Run tests on each batch to ensure no regressions
  4. Generate a report showing before/after for each file
  5. Verify that the total test suite passes

You’ll see output like:

✅ src/data_loader.py - 3 imports updated, 8 calls refactored
✅ src/preprocessing.py - 2 imports updated, 12 calls refactored
⚠️  src/models.py - 1 import updated, but 1 call failed type check (manually reviewed)
❌ src/legacy/old_pipeline.py - Skipped (uses numpy in outer scope, unsafe)

Summary:
- Files processed: 47
- Files updated: 44
- Files skipped: 3 (flagged for manual review)
- Tests passing: Yes (2,847 tests, 0 failures)

Notice how /batch didn’t blindly apply the change everywhere. It flagged edge cases. It skipped files where it wasn’t confident. This is intelligence, not automation theater.

Understanding Parallelization Strategy

One of the key advantages of /batch is its ability to process multiple files in parallel. But parallelization isn’t free—it comes with tradeoffs. When you run with --parallel=8, you’re telling Claude Code to process 8 files simultaneously. This is great for speed, but it increases API rate consumption. If you have strict rate limits, you might want --parallel=2 or --parallel=4 instead.

The parallelization also matters for error handling. If file 1 fails during processing, file 2-8 continue anyway. This means you get partial results—which is actually better than failing entirely. You can review what succeeded, fix what failed, and rerun just the failed batch.

Customizing /batch Behavior

The real power emerges when you customize the validation rules.

By default, /batch runs your test suite. But you can chain custom validators:

/batch \
  --glob="config/**/*.yaml" \
  --task="Normalize all YAML indentation to 2 spaces" \
  --validators=["yaml-lint", "config-schema-check", "no-secrets-scan"] \
  --parallel=8 \
  --dry-run

The --parallel=8 flag tells Claude Code to process 8 files at a time (balancing speed vs. API rate limits). The --dry-run flag shows you what would change without actually committing anything.

This is huge for safety. You can preview the operation, review the changes, and only then commit with confidence.

Edge Case: Handling Formatting Conflicts

One tricky scenario: what if applying /batch changes creates code that violates your linter? For example, you update an API call, but the new line is 200 characters long, and your style guide says max 100.

Claude Code handles this with a two-pass strategy:

  1. First pass: Apply the transformation
  2. Second pass: Run your linter/formatter to fix style issues

This means the final result is both functionally correct and style-compliant. You see both passes in the report, so you understand exactly what happened.

/simplify: Intelligent Code Refactoring with Safety Verification

Now imagine you have 3,000 lines of conditionals nested four levels deep. Your code works, but it’s a cognitive load to read. Every change risks a subtle bug.

The /simplify skill is your refactoring partner. It doesn’t just strip whitespace—it rewrites logic to be cleaner, more readable, and provably equivalent.

What /simplify Does

/simplify analyzes code structure and applies refactoring patterns:

  • Flattening nested conditionals
  • Extracting duplicated logic into functions
  • Replacing verbose patterns with cleaner idioms
  • Removing dead code
  • Consolidating similar branches

Critically: it proves the refactoring is safe by running your test suite before and after.

Real-World Example: Untangling Complex Business Logic

You inherit a 200-line function that handles user permissions:

def can_edit_document(user, document, workspace):
    if user is None:
        return False
    if not user.is_active:
        return False
    if workspace is None:
        return False
    if not workspace.is_active:
        return False
    if document is None:
        return False
    if document.deleted_at is not None:
        return False
    if user.role == "admin":
        return True
    if user.role == "owner":
        if document.owner_id == user.id:
            return True
        return False
    if user.role == "editor":
        if user.id in document.editor_ids:
            return True
        if user.id in workspace.team_ids:
            if workspace.default_role == "editor":
                return True
        return False
    if user.role == "viewer":
        return False
    return False

This works, but reading it is exhausting. You have to hold all the conditions in your head. The real danger is maintenance—someone changes one condition wrong, and now there’s a subtle permissions bug that only shows up in production.

Invoke /simplify:

/simplify src/permissions.py::can_edit_document \
  --level=moderate \
  --preserve-behavior \
  --validate="pytest tests/test_permissions.py -xvs"

Claude Code refactors it to something like:

def can_edit_document(user, document, workspace):
    # Early returns for invalid inputs
    if not all([user, document, workspace]):
        return False
    if not all([user.is_active, workspace.is_active, not document.deleted_at]):
        return False

    # Role-based access control
    if user.role == "admin":
        return True

    if user.role == "owner":
        return document.owner_id == user.id

    if user.role == "editor":
        return (user.id in document.editor_ids or
                (user.id in workspace.team_ids and
                 workspace.default_role == "editor"))

    return False  # "viewer" and unknown roles cannot edit

The refactored version is 40% shorter, reads top-to-bottom without cognitive backtracking, and passes 100% of tests.

Understanding Simplification Levels

/simplify comes with different aggressiveness levels:

  • Conservative: Only removes obvious dead code, doesn’t restructure
  • Moderate: Flattens simple nested conditions, extracts helper functions
  • Aggressive: Major restructuring, may extract multiple functions, can break code into completely different organization

For untested code, start with --level=conservative. For well-tested code (like the permissions example above), you can go to moderate or even aggressive.

Safety Gates in /simplify

/simplify doesn’t change your code and hope. It proves safety:

  1. Pre-refactor snapshot: Captures current behavior
  2. Refactoring: Applies transformations
  3. Test execution: Runs entire test suite
  4. Behavior validation: Compares pre/post outputs
  5. Evidence report: Shows diffs, test results, coverage impact

If tests fail, the refactoring is reverted automatically. You’re never left with broken code.

Troubleshooting Simplification Failures

Sometimes /simplify can’t complete. Maybe the code is too complex. Maybe there aren’t enough tests. Here’s how Claude Code handles it:

  1. It stops before making destructive changes
  2. It shows you why it stopped (missing tests, unclear behavior, etc.)
  3. It suggests what would help (add tests, break into smaller functions, etc.)

You’re never left with partially-refactored code. It’s all-or-nothing with evidence.

/loop: Iterative Improvement on Intervals

Now here’s the one that changes how you think about technical debt.

Most development workflows are linear: write code, test it, ship it, move on. Technical debt compounds silently until it’s a crisis.

The /loop skill inverts this. It runs improvement cycles at intervals you define—hourly, daily, or on-demand. Each cycle:

  1. Scans your codebase for optimization opportunities
  2. Applies improvements (refactoring, simplification, dependency updates)
  3. Verifies everything still works
  4. Commits with evidence
  5. Reports what changed and why

You’re continuously compressing technical debt, not ignoring it until it explodes.

Real-World Example: Automatic Code Health Checks

Set up a /loop that runs every 8 hours:

/loop \
  --cron="0 */8 * * *" \
  --tasks=[
    "simplify_functions:complexity_threshold=15",
    "remove_dead_code:analyze_imports",
    "update_dependencies:safe_patch_only",
    "lint_and_format:enforce_standards"
  ] \
  --validate="pytest -xvs && mypy . --strict" \
  --notify="slack:#dev-health" \
  --auto-commit="[loop] Automated code health improvements"

Every 8 hours, Claude Code:

  1. Finds functions with cyclomatic complexity > 15 and simplifies them
  2. Identifies unused imports and removes them
  3. Updates patch-level dependencies (e.g., 3.2.1 → 3.2.4)
  4. Runs the full test suite and type checker
  5. Commits if everything passes
  6. Posts a summary to your Slack channel

Over a month, you’ve eliminated 20+ small improvements that would’ve otherwise been ignored. Your codebase is steadily getting healthier.

Understanding Cron Scheduling

The --cron parameter uses standard cron syntax. Here are common patterns:

  • "0 2 * * *" – Every day at 2 AM
  • "0 */8 * * *" – Every 8 hours
  • "0 9 * * 1" – Every Monday at 9 AM (weekly)
  • "0 0 1 * *" – First day of month at midnight

The challenge with cron scheduling is timezone. Claude Code typically uses your system timezone, so test with a dry-run first to verify it’s running when you expect.

Customizing /loop Behavior

The power of /loop is its customization:

/loop \
  --interval="weekly" \
  --targets=["src/", "tests/"] \
  --skip=["migrations/", "legacy/"] \
  --validators=[
    "pytest:coverage_floor=85",
    "mypy:strict_mode",
    "ruff:rules=E,F,W"
  ] \
  --gate="minimum_score:3.5/5.0" \
  --dry-run-first \
  --require-approval-for="breaking_changes"

This means:

  • Run every week
  • Only scan src/ and tests/
  • Never touch migrations/ or legacy/
  • Coverage must stay ≥85%
  • Type checking in strict mode
  • Show me a dry-run first, then ask for approval if anything looks risky

The --dry-run-first flag is crucial for conservative teams. It shows you what would change before any actual changes happen. You review, approve, and then it runs for real.

Understanding the Safety Model Behind These Skills

These skills aren’t black boxes. They work within a rigorous safety framework—the same gates that govern production deployments at major tech companies.

Quality Gates: What Happens Before Anything Changes

Every operation goes through mandatory gates:

  1. Evidence Gate: No claim of “success” without proof. You see diffs, test results, coverage metrics.
  2. Validation Gate: Code must pass all tests before and after the change. If tests fail, the change is rejected automatically.
  3. Risk Callout Gate: For risky operations (deleting code, rewriting logic), you get an explicit risk assessment with rollback plan.
  4. Confidence Threshold: The system self-assesses whether it understands the task well enough. Low-confidence operations require additional verification.

This means you’re never in a situation where Claude Code “tried something” and left your codebase in an ambiguous state.

The Evidence Trail

Every change gets an audit trail. You can see:

  • Exactly what changed (diffs)
  • Why it changed (reasoning)
  • Whether it passed validation (test results)
  • The commit hash and timestamp
  • Who approved it (if manual approval was required)

This matters when you need to understand “why did this happen?” six months later. You have a complete record.

When to Use Each Skill

Here’s a quick decision matrix:

Skill Use When Example
/batch You need to apply the same change across many files Rename a function in 30 files, update an API call across your codebase, reformat YAML configs
/simplify You have messy code that works but is hard to maintain Complex conditionals, duplicated logic, long functions that should be broken up
/loop You want continuous, automated improvement Weekly tech debt reduction, automated dependency updates, periodic linting

Beyond the Basics: Combining Patterns

The real power emerges when you understand what each skill excels at and use them together.

Pattern 1: The Refactor-Then-Enforce Flow

  • Use /simplify to clean up high-complexity functions
  • Use /batch to apply new style rules across the simplified code
  • Use /loop to catch new violations automatically

Pattern 2: The Bulk-Update-and-Validate Flow

  • Use /batch to update an API call across 40 files
  • Use /loop with heightened validation to ensure no regressions
  • Commit with full evidence trail

Pattern 3: The Debt-Elimination Cycle

  • /loop runs nightly
  • Each cycle: scans for common smells (unused imports, long functions, duplicated patterns)
  • /simplify fixes high-priority issues
  • Test suite validates everything
  • Wake up to a cleaner codebase

Time Savings in Practice

Let’s quantify the value.

Scenario: You’re refactoring logging across a 50-file Python codebase.

Without Claude Code:

  • Manual review of each file: 50 × 10 min = 8+ hours
  • Making consistent changes: 50 × 8 min = 6+ hours
  • Testing edge cases: 4 hours
  • Fixing regressions: 3 hours
  • Total: 21+ hours (nearly 3 developer days)

With Claude Code:

  • Define the /batch task: 10 minutes
  • Run /batch in parallel: 15 minutes
  • Review report: 5 minutes
  • Approve and commit: 2 minutes
  • Total: 32 minutes

That’s a 40x speed improvement. And you’ve got empirical proof everything works (test results, coverage reports, diffs).

Scale this across your team. If your engineering organization does this refactoring 4 times per year, you’ve freed up 84 developer hours annually. That’s wages, opportunity cost, context switching—all gone.

The Quality Guarantee

Here’s what you’re actually buying with these skills: confidence.

Every /batch operation comes with:

  • A complete diff showing every change
  • Test results (pass/fail, coverage impact)
  • A rollback plan if something breaks
  • File-by-file verification status

Every /simplify operation comes with:

  • Pre-refactor and post-refactor behavior logs
  • Test suite results
  • Code complexity metrics (before/after)
  • Performance impact analysis
  • Line count reduction

Every /loop operation comes with:

  • Commit history with evidence
  • A reason for every change
  • Test results for each cycle
  • Observable metrics (complexity trend, coverage trend, etc.)

You’re not flying blind. You’re building with visibility.

Getting Started: Three Next Steps

  1. Try /batch on a non-critical file first: Pick a stylistic refactoring (variable naming, comment formatting) and run it with --dry-run first. See the diffs. Build confidence.

  2. Enable /simplify on your highest-complexity function: Use it as a learning tool. See how it rewrites your code. Compare before/after. Steal the patterns for other code.

  3. Set up a /loop with a single task: Start with a daily check for unused imports or simple linting. Watch the commits. Build a pattern.

From there, expand. Chain them. Customize validation rules. Build your unique improvement pipeline.

Real-World Examples: How Teams Actually Use These Skills

Case Study 1: Migrating a Legacy API (Using /batch)

A data team has 60 Python scripts that call a deprecated API endpoint. The endpoint still works but it’s being sunset in 6 months. Manual migration: 40+ hours of tedious work.

They use /batch:

/batch \
  --glob="scripts/**/*.py" \
  --task="Replace deprecated get_user_by_id(id) calls with new UserClient.fetch(user_id=id)" \
  --validate="pytest scripts/ --tb=short" \
  --evidence="Show me the before/after for each changed file" \
  --parallel=4

Result: 58 files updated in 22 minutes. 2 files flagged for manual review (complex edge cases). All tests pass. Zero surprises in production. They estimate they saved 35+ hours of developer time.

Case Study 2: Untangling a Business Logic Function (Using /simplify)

A fintech team has a 300-line function that calculates transaction fees. It works. It’s survived three years of feature additions. But it’s nearly impossible to change without risking subtle bugs.

They use /simplify:

/simplify src/payments/fee_calculator.py::calculate_transaction_fee \
  --level=aggressive \
  --preserve-behavior \
  --validate="pytest tests/payments/test_fee_calculator.py -xvs && test_against_production_logs"

The function gets broken into six smaller, focused functions. Each has one responsibility: validate input, calculate base fee, apply discounts, apply tax, apply regional rules, return result.

Before: 300 lines, cyclomatic complexity 47, coverage 65%
After: 280 lines across 6 functions, complexity 8-12 per function, coverage 94%

New developer onboarding time for this function dropped from 2+ hours to 15 minutes. That’s a tangible business impact.

Case Study 3: Continuous Code Improvement (Using /loop)

A startup runs /loop on a nightly schedule:

/loop \
  --cron="0 3 * * *" \
  --name="automated_health_checks" \
  --tasks=[
    "remove_unused_imports",
    "simplify_functions:complexity_threshold=15",
    "update_dependencies:security_patches_only",
    "enforce_code_style"
  ] \
  --validators=["pytest", "coverage:floor=85", "security_scan"] \
  --slack-summary="#engineering" \
  --auto-commit-if-all-pass

Over 90 days:

  • 180 unused imports removed
  • 23 functions simplified
  • 34 security patches applied
  • 0 production incidents from automated changes
  • Estimated manual work saved: 40+ hours

They’ve built a system that continuously improves without human intervention. But it’s not reckless—every change has evidence behind it.

Measuring Success: Metrics That Matter

How do you know these skills are actually working? You need metrics. Not vanity metrics, but real indicators of code health.

Tracking Code Complexity Trends

Over time, track your average cyclomatic complexity:

Month 1: 12.4 (baseline)
Month 2: 11.8 (-5%)
Month 3: 10.2 (-18%)

This is real improvement. Lower complexity means easier maintenance, fewer bugs, faster onboarding.

Claude Code generates these reports automatically. You can plot them over time and see if your automated improvements are working.

Developer Velocity Impact

Track how many refactorings happen per month:

Manual refactoring (before): 2-3 per month
Automated (after): 15-20 per month

That’s a 6-8x increase in refactoring frequency. Your code evolves faster. Technical debt shrinks instead of growing.

Production Incident Impact

This is the real metric: do automated improvements reduce production issues?

Some teams track incidents related to code quality:

Before: 2-3 incidents per month related to complex code
After 6 months: 0-1 incident per month

This is pure business impact. Fewer incidents mean more uptime, happier customers, happier team.

The Long-Term Picture

Think about this: most codebases start clean. New project, best practices, clear architecture. Then six months pass. Features ship. Deadlines pressure developers to cut corners. Technical debt accumulates. After two years, the codebase is a mess.

With Claude Code’s bundled skills, you invert that trajectory.

Without automation:

  • Day 1: Score 8/10 (clean)
  • Month 6: Score 6/10 (okay)
  • Year 1: Score 4/10 (degrading)
  • Year 2: Score 2/10 (crisis)

With automation:

  • Day 1: Score 8/10 (clean)
  • Month 6: Score 8.5/10 (improvements)
  • Year 1: Score 9/10 (better than start)
  • Year 2: Score 9.2/10 (continuous improvement)

Your code doesn’t decay. It evolves. Every week you wake up to a slightly better codebase than the day before.

That’s the real value. Not the time savings (though that matters). Not the reduced incidents (though that matters too). It’s the fundamental inversion: instead of fighting against entropy, you’re riding with it. Your code is automatically getting better.

The Bottom Line

Your code is never “done.” It’s constantly decaying. Complexity compounds. Dependencies age. Patterns drift.

Without these skills, you manage decay reactively. You wait until code is so messy that it becomes a crisis (the annual “refactoring sprint” that ships late). Then you spend weeks fixing what accumulated over months.

With Claude Code’s bundled skills, you manage decay proactively. You simplify automatically. You batch-apply improvements. You run continuous health checks.

Your codebase doesn’t get worse. It gets better. Every week. With empirical proof that everything works.

You gain back time. Your team’s velocity increases. Your incidents decrease because code is cleaner and easier to reason about.

And the best part? You’re not hoping it works. You have evidence.


-iNet

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.