All Articles Claude Code

Hook-Based Test Runner: Auto-Test After Code Changes

You just pushed code changes. Your tests are scattered across three directories. Do you run all 200 tests and wait 45 seconds?

You just pushed code changes. Your tests are scattered across three directories. Do you run all 200 tests and wait 45 seconds? Or do you run the three tests that matter, in 1.2 seconds, and ship with confidence? The answer lies in a Git hook that intercepts your commit workflow, maps changed files to their test suites, runs only relevant tests, and prevents commits when tests fail. Let’s build that system.

Why Smart Testing Matters for Productivity

The difference between a test hook that takes 2 seconds and one that takes 30 seconds is massive. With 2-second tests, developers commit multiple times per minute if they want. With 30-second tests, developers commit once per minute or less. Over a developer’s career, this compounds. A developer with 40 commits per week times 30 seconds is 20 minutes of waiting per week per person. Across 200 developers, that’s 67 hours per week of waiting.

This waiting has consequences. Developers start committing less frequently to reduce waiting. They batch changes together. This reduces feedback frequency and makes debugging harder when tests fail. Developers disable the hook to ship faster. They skip testing intentionally. Once that pattern starts, it’s hard to stop. “Everyone skips testing, so why shouldn’t I?”

Intelligent test selection keeps tests fast. Running only relevant tests means developers rarely wait. When they don’t wait, they don’t disable hooks. When they don’t disable hooks, bugs get caught locally. The entire dynamic shifts from adversarial (developers vs. tests) to collaborative (developers with tests).

The Problem: Testing Doesn’t Scale

When you modify lib/auth.js, do you need to test your database migration library? No. But without intelligent mapping, you run the whole suite anyway. Real projects have this problem consistently. Changing one component triggers unrelated tests. Developers disable pre-commit hooks to ship faster. CI catches bugs that local testing should have caught. Time-to-feedback grows with repository size.

We need automation that runs the right tests, every time, without manual intervention. That’s what makes the difference between a local workflow that blocks development and one that accelerates it. When developers see instant feedback on what they changed, they move faster and ship more confidently. When they wait 45 seconds for unrelated tests, they start skipping the hook.

The Architecture: Understanding Test Mapping

Test mapping is conceptually simple but practically complex. In a well-organized codebase with clear conventions, one test file per source file makes mapping trivial. But real codebases have complexity. One test file might validate multiple source files (integration tests). One source file might be validated by multiple test files (unit tests plus integration tests). Some source files don’t need tests (configurations, data files). Some test files aren’t tied to specific source files (end-to-end tests, performance tests).

The solution is layered. First, try direct mapping—if source file lib/auth.js exists, look for test file test/auth.test.js. If found, run it. Second, try pattern matching—if direct mapping fails, apply configured patterns like “src/ → test/“. Third, try dependency analysis—scan the test files to see which ones import the changed source file. Tests that import changed files should be run.

This layered approach handles most real-world complexity. A few test mappings are ambiguous, but most are clear. For ambiguous cases, err on the side of running tests rather than skipping them. Running extra tests is slower but safer. Skipping tests that should run is dangerous.

The Hook Flow: How Everything Works Together

The test runner lives in a pre-commit hook. When you git commit, the hook detects changes by scanning git diff --cached for modified files. It maps those files to associated test suites. It runs only relevant tests. It exits with code 0 to allow commit or 1 to block. It reports results clearly showing developers exactly what passed and failed.

The key insight is intelligent filtering. You don’t run all tests—you run the tests that could possibly be affected by your changes. This keeps feedback fast while maintaining safety. A developer who changes a utility function runs all tests that depend on that utility. A developer who adds a new feature runs only the feature’s test suite.

Let’s walk through what actually happens when you commit. The hook starts by detecting code changes in your staging area. The --cached flag captures staged changes. The --diff-filter=ACMR excludes deletions—you don’t test deleted code. The regex filters for JavaScript and TypeScript files only. Why not test all file types? Because .md, .yaml, and .json changes don’t execute code. This saves time and eliminates false positives.

This filtering is critical for adoption. When developers see that their documentation change doesn’t trigger a 45-second test run, they appreciate the intelligence. When they see that a comment-only change doesn’t run tests, they understand the system is thoughtful. This builds trust.

Next comes test file mapping. Now we need to know which test file corresponds to lib/auth.js. Some conventions make this easy; others require a mapping file. The system uses a .test-map.json file if it exists, otherwise falls back to directory conventions. By returning only existing test files, we avoid errors from tests that don’t exist yet.

This mapper handles real-world complexity. Different teams organize code differently. A monorepo might have tests in a separate tests/ directory. A polyrepo might keep tests alongside source. Stale test mappings happen when people rename files. By checking existence first, we avoid chasing ghosts.

The third stage is executing tests intelligently. We need a test runner that executes only mapped tests, captures pass/fail status, reports results clearly, and stops on first failure. The runner abstracts over different test frameworks. If you’re using Jest, it runs Jest. If you switch to Vitest, change one constant and the rest stays the same.

The framework abstraction matters because test frameworks change. You might upgrade from Jest to Vitest. You might standardize on Vitest across multiple teams. When the test framework abstraction is built into your hook, these changes are cheap. You update one line of code.

The full hook orchestrates everything: get changed files, load mappings, find tests, run tests, report results, allow or block commit. The hook uses colors to make output readable. Green checkmarks for passed tests. Red X marks for failures. Dimmed text for meta information. This visual clarity helps developers understand what happened at a glance.

The hook uses environment variables to detect critical branches. On main, you might enforce tests regardless of size. On feature branches, you might be more lenient. This flexibility lets you apply different policies to different branches.

Handling the Messy Reality

Real systems have edge cases. Some files shouldn’t trigger tests—markdown docs, YAML configs, .env files, build outputs. Skipping them saves time. Deduplicating test paths prevents running the same test twice if you modified two files that share coverage. Expanding related tests discovers dependencies—maybe a shared utility test validates behavior across multiple test files.

Allowing --no-verify to bypass the hook is important for emergencies. A critical production bug at 3 AM might need to skip testing. But for critical branches like main, you might want to enforce tests regardless. This gives you safety rails where they matter most.

Integrating with Claude Code

When the hook succeeds, developers move forward. But what happens when Claude Code makes changes? We need to feed test results back into the AI’s context. The system creates a memory file that Claude can read on next invocation. This creates a feedback loop: code change → test run → AI learns. If Claude modified code and the hook runs, Claude immediately sees the test results.

This feedback loop is transformative. Claude sees that its change failed tests. On the next iteration, Claude reads the test memory and understands the failure. Claude can self-correct without human intervention. This is where AI development becomes powerful—when the AI learns from its own mistakes.

Making Installation Easy

Developers need a simple way to install and configure this hook. A setup script installs the pre-commit hook and creates default configuration. Running npm run setup:hooks or node scripts/setup-hooks.mjs handles everything. New developers don’t need to manually install hooks.

Performance Optimization: Caching

For large test suites, caching prevents re-running the same tests unnecessarily. Hash file contents to detect changes. If source hasn’t changed, use cached test result. If source changed, re-run. This optimization drops test time from seconds to milliseconds on subsequent commits with unchanged source files.

Caching requires care—a cached pass on old source shouldn’t override a failure on new source. The system stores source file hash alongside test result. If hashes don’t match, it re-runs. This ensures correctness while maintaining speed.

Dealing with Flaky Tests

Flaky tests—that pass sometimes and fail randomly—are a special problem. The hook needs to distinguish between real failures and transient failures. The system tracks test flakiness by keeping records of recent runs. If a test fails 20-80% of the time, it’s flaky. The system retries flaky tests multiple times before failing them.

Identifying flaky tests is valuable. A test that fails 50% of the time is telling you something is wrong. The hook tracks this data so you can investigate. Over time, you identify and fix the root causes of flakiness.

Observability and Metrics

A production-grade hook needs visibility. Metrics record test execution, duration, file count, pass/fail rates. Over time, these metrics reveal patterns. Tests are getting slower. Failure rate is increasing. Developers skip hooks more on certain branches. This data informs decisions about test architecture and optimization.

The metrics use JSONL format for easy analysis. Each line is a complete test run record. You can query these logs to answer questions like: “What’s the average test time over the last week?” or “Which tests are slowest?” or “Has failure rate increased recently?”

Why This Matters

You’ve now automated the “should I test this?” decision. The system always runs the right tests. Developers always get instant feedback. No more skipping tests to ship fast. No more catching bugs in CI that should have been caught locally.

The hook is intelligent about scope—it tests what matters, ignores what doesn’t, reports clearly. Over a career of commits, that saves hours of waiting for irrelevant test suites to run. A developer who commits 100 times a week saves 45 minutes per week if average test time drops from 27 seconds to 0.3 seconds. That compounds.

When you integrate caching, the average local test run drops from seconds to milliseconds on subsequent commits. When you track metrics, you can prove that the hook is working and identify when test architecture needs attention. When you feed results back to Claude Code, the AI learns from test failures and self-corrects.

The best testing infrastructure is one you don’t think about—it just works, runs only what matters, and gets out of the way. That’s what a well-designed pre-commit hook delivers. Developers ship faster. Bugs are caught locally. Your CI pipeline has fewer false positives. The entire team is happier and more productive.

Evolution and Long-Term Maintenance

Your test infrastructure will outlive individual test frameworks. You might start with Jest, but frameworks change. By building framework abstraction into your hook, you make framework migrations cheap. When the team decides to switch to Vitest, you update configuration, not the core logic.

This is important for long-lived projects. A codebase that’s been around for 10 years will probably outlive multiple test framework generations. Building for that reality saves you from rewrites. Your hook adapts rather than resists framework evolution.

Cross-Team Configuration

When you have multiple teams, each might want different test behavior. The frontend team wants strict testing—all tests run. The infrastructure team is okay with running only direct tests. The data team wants coverage reports. Your hook needs to be configurable per directory or per team.

This can be achieved through multiple .test-map.json files or through a centralized configuration with team overrides. Different teams get different behavior without conflicts. The platform team sets defaults, but individual teams can customize. This balance gives autonomy while maintaining standards.

Integration with IDEs

Most developers use an IDE or editor that supports git hooks. When you run git commit, your IDE sees the hook success/failure. Advanced IDE integration can show test results inline while you’re still editing. You see that your change breaks a test before you commit. This immediate feedback is incredibly powerful.

Claude Code can enhance this further. When Claude makes code changes, it runs the hook and sees test results. If tests fail, Claude reads the failure messages and self-corrects. This closes the loop completely—code generation, testing, and correction all happen automatically.

Scaling Performance

As your test suite grows, hook performance becomes critical. A hook that takes 2 seconds per commit is fine. A hook that takes 30 seconds gets skipped. Your caching strategy is essential. Only re-run tests when source files change. Use parallel test execution when possible. Keep test files close to source files so filtering is fast.

Metrics matter here—measure hook performance over time. If average execution time is increasing, investigate why. Maybe your test suite is growing faster than your caching improves. Maybe your mappings got inefficient. These metrics guide optimization work.

Learning from Failures

Each test failure is data. When a test fails consistently, that tells you something. When a test fails on some developers’ machines but not others, that suggests environment issues. When tests fail after certain code changes, that’s signal about code quality. The hook becomes a data gathering system if you treat failures as learning opportunities.

Over time, you can analyze test failure patterns. Most failures are in unit tests or integration tests? That might suggest your testing strategy needs adjustment. Certain files are always in failing PRs? That code might need refactoring. Certain developers have more test failures? Maybe they need mentoring. The data is there if you look.

Building Developer Trust

Developers will only respect a test hook if they trust it. Trust is earned through consistency, clarity, and responsiveness. A hook that sometimes fails on valid code destroys trust. A hook that is slow and frustrating destroys trust. A hook with unclear error messages confuses developers.

Building trust requires attention to detail. Make sure your test selection is accurate. Developers should rarely see “but I didn’t change that file” when a test fails. Make sure error messages are clear. Developers should understand what went wrong and how to fix it. Make sure the hook is fast. Developers should barely notice it’s running.

When developers trust the hook, they embrace it. They stop trying to skip it. They rely on it to catch bugs. They use it as a learning tool. That’s when the real value emerges. The hook becomes part of the culture, not an obstacle to work around.

Developer Experience and Performance

The difference between a 2-second hook and a 10-second hook is significant. At 2 seconds, developers don’t mind the hook overhead. At 10 seconds, developers start looking for ways to skip it. At 30 seconds, developers actively resist the hook.

So performance optimization isn’t optional—it’s critical. Use parallel test execution. Use caching. Use dependency analysis to run only relevant tests. Profile your hook to find bottlenecks. If test mapping is slow, optimize the mapping. If test execution is slow, parallelize it.

The hook should be transparent to developers. It should run, show results, and get out of the way. Developers shouldn’t think about the hook. They should think about their code and tests. The hook is invisible infrastructure.

Pre-commit vs. Pre-push: Different Strategies

Some teams prefer pre-commit hooks (run before local commit). Others prefer pre-push hooks (run before pushing to remote). Pre-commit catches issues before commits—faster feedback. Pre-push prevents bad commits from reaching shared branches. The choice depends on your team’s style.

You can use both—lightweight pre-commit (quick smoke tests), comprehensive pre-push (full test suite). This balances fast feedback with safety. Local development moves fast because commits are cheap. But remotes stay clean because push requires passing tests.

Monorepo Complexity

Large monorepos present special challenges. When you have 500+ test files and changes touch multiple subsystems, filtering matters a lot. A change in the auth subsystem shouldn’t run the payments subsystem tests. But a change in shared utilities should run all tests that depend on it.

This requires sophisticated dependency analysis. You need to understand not just file-level dependencies, but subsystem-level dependencies. A utility shared across 50 tests needs to run all 50 when modified. A subsystem-specific utility needs to run only its own tests. The complexity is high, but the payoff is huge—avoiding running 1000 tests when only 50 are affected.

Test Evolution as Your Codebase Grows

As your codebase evolves, tests need to evolve too. Old tests become obsolete. New patterns require new tests. The hook needs to handle this gracefully.

When source code is deleted, mapped tests might become orphaned. The hook detects this and suggests cleanup. When new test files appear, the hook can detect them and update mappings. This automated maintenance prevents test cruft.

Sometimes tests become slower as the codebase grows. A test that ran in 100ms might take 5 seconds after you add complex setup. The hook tracks this and flags tests that have degraded significantly. Maybe that test needs optimization. Maybe it needs to be broken into faster unit tests. The data helps you make good decisions.

Test coverage metrics from the hook show gaps. If you consistently commit code without tests, the hook can flag it and suggest adding tests before committing. This creates positive incentive to write tests—you can’t commit without them.

Building Testing Culture

There’s always pressure to skip hooks. Just before a deadline, someone wants to commit without testing. Without proper culture, the hook becomes something to work around rather than something to embrace. Preventing this requires leadership visibility. Leaders should commit with hooks passing. They should reinforce that testing is non-negotiable.

It also requires making hooks fast enough to not be a burden. If you address performance issues and provide clear feedback, developers don’t want to skip. They trust the hook and accept its verdict. They want it to pass. That’s when hooks are truly effective.

Continuous Learning Through Tests

Beyond catching bugs, test hooks are learning tools. When a test fails, it’s teaching you something. A test that fails frequently might indicate the code is fragile. A test that is complex might indicate the code being tested is complex. A test that requires setup might indicate the code has too many dependencies.

Experienced teams use test failures as signals about code health, not just as blockers to fixing bugs. They ask “why is this test fragile?” and refactor the code. They ask “why is this test complex?” and simplify the code or the test. Over time, this leads to codebases that are easier to test and therefore easier to modify.

The hook becomes a teaching channel between your codebase and developers. The tests communicate “this is important to preserve.” The failures communicate “this assumption broke.” Developers who listen to these signals write better code.

Testing Infrastructure as Competitive Advantage

Your code is only as good as your ability to verify it works. Your testing infrastructure is the linchpin of code quality. Invest in it early. Make it fast. Make it intelligent. Make it trustworthy. Teams with excellent testing infrastructure ship better code faster. That’s not an accident—it’s the compounding effect of small investments in the right place.

A pre-commit hook is one of the simplest and most effective tools you can deploy. It’s automation that works. It’s instant feedback. It’s early problem detection. It’s developer education. It’s all of those at once. Build a good one, maintain it, and you’ve just multiplied your team’s productivity and code quality. This is worth your time and attention today. The most successful teams invest in this infrastructure first and maintain it continuously over years. It pays dividends in quality and speed exponentially. Your testing infrastructure is where technical excellence begins. Make it a strategic priority and watch your team’s velocity and confidence soar.


-iNet

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.