Every program needs a graceful exit. That’s where stop hooks come in. When Claude Code is wrapping up a session—whether the user asked it to stop, hit an error, ran out of context, or encountered a timeout—you often need to clean up. Close database connections. Flush logs to disk. Save session state. Upload analytics. Reset temporary files. Coordinate with external systems to say “we’re done here.” Stop hooks are the mechanism that lets you run code when Claude Code is shutting down. Not at an arbitrary moment during execution, but at the very end, after all the work is done, during the organized shutdown sequence. This is where graceful shutdown happens. This is where you ensure that everything in your system is left in a consistent state, ready for the next session or ready to restart without data loss.
Think of it like turning off a restaurant kitchen at the end of the day. You don’t just flip the switch. You finish the last order, clean the grills, put away the knives, secure the door, count the register. The shutdown sequence is as important as startup. Stop hooks orchestrate that sequence for your Claude Code sessions. They’re the difference between “the session just ended abruptly” and “the session ended cleanly and everything is ready for the next one.”
Why Stop Hooks Matter: The Difference Between Graceful and Chaotic
Most engineers don’t think about the exit scenario. They focus on the happy path—code runs, does useful work, produces results. But production systems fail, timeout, or need to stop abruptly. When that happens, what state are you in? Are there open connections? Temporary files? Buffered data waiting to be written? Unsaved session state that you’ll need to reconstruct?
A real scenario helps illustrate the problem: Claude Code is running a long data migration. It’s been processing records for an hour. Suddenly, the context window fills up. Claude is told to stop and wrap up. At that moment, you want to:
- Flush buffers – Write any in-flight data to the database
- Save checkpoint – Record which record was last processed, so resuming is fast
- Close connections – Gracefully disconnect from the database
- Log final state – Record metrics about what completed, what didn’t
- Clean temp files – Remove any temporary files created during the run
- Notify downstream systems – Tell any dependent services “the migration is done”
- Generate summary – Create a report of what happened
Without stop hooks, you can’t guarantee any of this happens. You’re hoping Claude gracefully cleans up. And hope is not a strategy for production systems. Stop hooks are your guarantee. They run automatically when Claude Code stops. You define them once, trust them to run, and know your systems will be in a clean state. The difference between hope and certainty is worth understanding deeply.
Understanding When Stop Hooks Fire: The Lifecycle Moments
Stop hooks fire at specific moments in the Claude Code lifecycle. Understanding the timing is important because the state of the system is different at each point:
Normal Exit: The user types “exit” or the session completes naturally. All work is done, Claude Code is shutting down cleanly. This is the easiest case—there’s no rush, no errors, just orderly shutdown. You have time to do thorough cleanup. Every database connection can be closed properly. Every log can be flushed. Every notification can be sent.
Context Window Full: Claude has used up all available tokens. The session must end. There’s urgency here because the last few tokens are precious. Cleanup code needs to be fast and not generate more tokens if possible. You’re operating in resource-constrained conditions. A cleanup that takes 50 tokens might be acceptable. A cleanup that takes 5000 tokens means you can’t do as much work.
User Interrupt: The user hits Ctrl+C or closes the terminal. This is abrupt. There might not be much time. Stop hooks should be designed to fail gracefully. You might get 2 seconds instead of 30. You can’t assume you’ll have time to do thorough cleanup.
Error or Timeout: Something went wrong. A tool crashed, a network request timed out, or Claude encountered an error it can’t recover from. The system is stopping to prevent further damage. Stop hooks should be idempotent—safe to run even if something is broken. A database connection might be gone already. A log file might be locked. Your cleanup code has to handle these degraded conditions.
The key insight: stop hooks run in degraded conditions. The system might be partially broken. Resources might be low. Timeouts might happen. Design them defensively. Assume the worst and make your cleanup code bulletproof. This defensive posture is what keeps your systems reliable even when things fall apart.
Stop Hook Architecture: The Contract and Implementation
Stop hooks follow a simple contract. When Claude Code is exiting, it calls your stop hook with context about why it’s stopping, and gives you a limited time window to clean up.
// .claude/hooks/stop/cleanup-hook.mjs
export const handler = async (context) => {
// context contains:
// - reason: "exit" | "context-full" | "error" | "interrupt" | "timeout"
// - timestamp: ISO timestamp of when stop was triggered
// - duration: how long this session lasted (in ms)
// - toolCalls: count of tool calls made
// - tokenCount: approximate tokens used
// - error: if reason is "error", the error message
// - timeRemaining: ms before this hook is forcefully terminated
// Do your cleanup here
// Keep it simple, fast, and idempotent
return {
success: true,
cleaned: ["database", "tempfiles", "logs"],
message: "Cleanup completed successfully"
};
};
This is synchronous in intent but asynchronous in execution. Claude Code calls your hook, waits for it to complete (with a timeout), and if it finishes in time, great. If it doesn’t, Claude Code stops anyway. You don’t have unlimited time. The architecture is simple: one function, one contract, one guarantee. When the function returns, cleanup is done. When the timeout expires, Claude Code exits whether the function finished or not. This simplicity makes stop hooks reliable and predictable.
Building Your First Stop Hook: The Pattern
Let’s build a practical stop hook that handles the common cleanup scenarios. This is a template you can adapt for your specific needs. Understanding the pattern will help you build hooks for your specific systems.
// .claude/hooks/stop/flush-and-save-hook.mjs
export const handler = async (context) => {
const startTime = Date.now();
const results = {
flushedLogs: false,
savedState: false,
closedConnections: false,
errors: []
};
try {
// 1. Flush any pending logs
if (global.pendingLogs && global.pendingLogs.length > 0) {
try {
const logPath = join(".claude", "logs", "session.log");
await fs.mkdir(join(".claude", "logs"), { recursive: true });
const logContent = global.pendingLogs.join("\n");
await fs.appendFile(logPath, logContent + "\n");
results.flushedLogs = true;
global.pendingLogs = [];
} catch (err) {
results.errors.push(`Failed to flush logs: ${err.message}`);
}
}
// 2. Save session state if we have one
if (global.sessionState) {
try {
const statePath = join(".claude", "session-state.json");
await fs.mkdir(".claude", { recursive: true });
const stateToSave = {
...global.sessionState,
endedAt: new Date().toISOString(),
stoppedBy: context.reason,
duration: context.duration,
toolCalls: context.toolCalls
};
await fs.writeFile(statePath, JSON.stringify(stateToSave, null, 2));
results.savedState = true;
} catch (err) {
results.errors.push(`Failed to save state: ${err.message}`);
}
}
// 3. Close database connections if we have them
if (global.dbConnection) {
try {
// Different database clients have different close methods
if (global.dbConnection.end) {
await global.dbConnection.end();
} else if (global.dbConnection.close) {
await global.dbConnection.close();
} else if (global.dbConnection.disconnect) {
await global.dbConnection.disconnect();
}
results.closedConnections = true;
} catch (err) {
results.errors.push(`Failed to close DB connection: ${err.message}`);
}
}
const elapsed = Date.now() - startTime;
return {
success: results.errors.length === 0,
results,
elapsed,
stopped: context.reason,
timeRemaining: context.timeRemaining - elapsed
};
} catch (err) {
return {
success: false,
error: err.message,
results
};
}
};
This hook demonstrates the pattern: try each cleanup in sequence, catch errors individually so one failure doesn’t stop other cleanups, report what you did, watch the clock. It’s idempotent—safe to run twice. One failure doesn’t stop other cleanups from running. This resilience is critical when operating in degraded conditions.
Advanced Pattern: Database Transaction Cleanup
A common scenario: Claude Code has open database transactions that need to be rolled back or committed. This requires care because incomplete transactions can lock tables and block other operations:
// .claude/hooks/stop/database-cleanup-hook.mjs
export const handler = async (context) => {
const results = {
transactionsRolledBack: 0,
connectionsClosedGracefully: 0,
connectionsForceKilled: 0,
errors: []
};
if (!global.dbPool) {
return { success: true, results, message: "No database pool found" };
}
try {
// For PostgreSQL clients
if (global.dbPool.query) {
try {
// Check for active transactions and roll them back
const activeQueries = await global.dbPool.query(
"SELECT pid, query_start, state FROM pg_stat_activity WHERE datname = current_database() AND state = 'active'"
);
if (activeQueries.rows && activeQueries.rows.length > 0) {
try {
await global.dbPool.query("ROLLBACK");
results.transactionsRolledBack++;
} catch (err) {
// Might already be rolled back, that's ok
}
}
// Gracefully close the pool
await global.dbPool.end();
results.connectionsClosedGracefully++;
} catch (err) {
results.errors.push(`PostgreSQL cleanup error: ${err.message}`);
}
}
// For MySQL/MariaDB clients
if (global.dbPool.getConnection) {
try {
await global.dbPool.end();
results.connectionsClosedGracefully++;
} catch (err) {
results.errors.push(`MySQL cleanup error: ${err.message}`);
try {
global.dbPool.destroy();
results.connectionsForceKilled++;
} catch (err2) {
results.errors.push(`MySQL force-kill failed: ${err2.message}`);
}
}
}
return {
success: results.errors.length === 0,
results,
reason: context.reason
};
} catch (err) {
return {
success: false,
error: err.message,
results
};
}
};
This hook handles the nuances of different database systems. Some need explicit rollback. Some need connection pool cleanup. Some need force-kill as a fallback. It tries them in order of gentleness, degrading gracefully if earlier approaches fail. The progression from gentle to forceful ensures that even if the database is unresponsive, you don’t hang waiting for timeouts.
Pattern: Checkpoint Saving for Resumable Work
If Claude Code is doing long-running work, save a checkpoint so resuming is fast. This is critical for batch jobs:
// .claude/hooks/stop/checkpoint-save-hook.mjs
export const handler = async (context) => {
if (!global.checkpoint) {
return { success: true, message: "No checkpoint tracking" };
}
try {
const checkpointPath = join(".claude", "checkpoints", `${global.checkpoint.jobId}.json`);
await fs.mkdir(join(".claude", "checkpoints"), { recursive: true });
const checkpoint = {
jobId: global.checkpoint.jobId,
lastProcessedId: global.checkpoint.lastProcessedId,
itemsProcessed: global.checkpoint.itemsProcessed,
itemsFailed: global.checkpoint.itemsFailed,
startedAt: global.checkpoint.startedAt,
lastCheckpoint: new Date().toISOString(),
stoppedBy: context.reason,
context: {
duration: context.duration,
toolCalls: context.toolCalls,
tokenCount: context.tokenCount
},
estimatedTimeRemaining: estimateRemainingTime(global.checkpoint),
resumeFrom: {
afterId: global.checkpoint.lastProcessedId,
skipCount: global.checkpoint.itemsProcessed,
retryFailed: global.checkpoint.itemsFailed > 0
}
};
// Write checkpoint
await fs.writeFile(checkpointPath, JSON.stringify(checkpoint, null, 2));
// Also write to a "latest checkpoint" file for quick access
const latestPath = join(".claude", "checkpoints", "latest.json");
await fs.writeFile(latestPath, JSON.stringify({
jobId: global.checkpoint.jobId,
path: checkpointPath,
savedAt: checkpoint.lastCheckpoint
}));
return {
success: true,
checkpoint,
path: checkpointPath,
resumable: true
};
} catch (err) {
return {
success: false,
error: err.message,
message: "Failed to save checkpoint"
};
}
};
function estimateRemainingTime(checkpoint) {
if (!checkpoint.startedAt || checkpoint.itemsProcessed === 0) {
return null;
}
const elapsedMs = Date.now() - new Date(checkpoint.startedAt).getTime();
const timePerItem = elapsedMs / checkpoint.itemsProcessed;
const remainingItems = (checkpoint.totalItems || 0) - checkpoint.itemsProcessed;
return Math.round(timePerItem * remainingItems / 1000); // seconds
}
This hook saves everything needed to resume work from where you left off. On restart, Claude can read this checkpoint and skip the already-processed items. For a 10-hour job that gets interrupted after 6 hours, this saves you from redoing 6 hours of work. Checkpointing transforms long-running jobs from “fragile and vulnerable to interruption” to “reliable and resumable.”
Pattern: Metrics and Analytics Flush
Long-running Claude Code sessions often accumulate metrics. On stop, upload them:
// .claude/hooks/stop/metrics-flush-hook.mjs
export const handler = async (context) => {
if (!global.metrics) {
return { success: true, message: "No metrics collected" };
}
try {
const finalMetrics = {
session: {
startedAt: global.metrics.sessionStart,
endedAt: new Date().toISOString(),
duration: context.duration,
stoppedBy: context.reason
},
operations: {
toolCalls: context.toolCalls,
fileWrites: global.metrics.fileWrites || 0,
fileReads: global.metrics.fileReads || 0,
dbQueries: global.metrics.dbQueries || 0,
errors: global.metrics.errors || 0
},
performance: {
avgToolCallDuration: global.metrics.toolDurations?.length
? Math.round(global.metrics.toolDurations.reduce((a, b) => a + b, 0) / global.metrics.toolDurations.length)
: 0,
slowestOperation: global.metrics.toolDurations?.length
? Math.max(...global.metrics.toolDurations)
: 0
},
tokens: {
estimated: context.tokenCount,
utilized: Math.round((context.tokenCount / 200000) * 100)
}
};
// Write metrics to file
const metricsPath = join(".claude", "metrics", `session-${Date.now()}.json`);
await fs.mkdir(join(".claude", "metrics"), { recursive: true });
await fs.writeFile(metricsPath, JSON.stringify(finalMetrics, null, 2));
// Also update a "latest metrics" file
const latestPath = join(".claude", "metrics", "latest.json");
await fs.writeFile(latestPath, JSON.stringify(finalMetrics, null, 2));
// Optionally upload to analytics service
if (process.env.ANALYTICS_ENDPOINT) {
try {
const response = await fetch(process.env.ANALYTICS_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(finalMetrics),
signal: AbortSignal.timeout(5000)
});
if (!response.ok) {
console.warn(`Analytics upload returned ${response.status}`);
}
} catch (err) {
console.warn(`Failed to upload analytics: ${err.message}`);
// Don't fail the whole hook for this
}
}
return {
success: true,
metricsPath,
metrics: finalMetrics
};
} catch (err) {
return {
success: false,
error: err.message
};
}
};
This hook shows how to capture rich metrics and optionally upload them. Over time, these metrics become valuable data about what Claude Code is doing, how long operations take, where bottlenecks are. This telemetry feeds back into system improvements.
Managing Stop Hook Timeouts: The Race Against Time
Stop hooks have limited time. By default, Claude Code gives you 5 seconds to clean up. If your hook takes longer, it gets terminated. Design for this aggressive deadline:
// .claude/hooks/stop/timeout-aware-hook.mjs
export const handler = async (context) => {
const deadline = Date.now() + context.timeRemaining - 500; // Leave 500ms buffer
const results = [];
// Quick operations first
try {
if (Date.now() < deadline - 1000) {
results.push(await quickCleanup());
}
} catch (err) {
results.push({ success: false, error: err.message });
}
// Only do slow operations if we have time
try {
if (Date.now() < deadline - 2000) {
results.push(await mediumCleanup());
}
} catch (err) {
results.push({ success: false, error: err.message });
}
// Very slow operations only get attempted if we have plenty of time
try {
if (Date.now() < deadline - 3000) {
results.push(await slowCleanup());
}
} catch (err) {
// Don't bother logging if we're out of time
}
return {
success: true,
completed: results.filter(r => r.success).length,
skipped: results.filter(r => !r.success).length
};
};
Know how much time you have and prioritize ruthlessly. Do critical work first (flush databases, save state). Skip nice-to-have work if time is tight (generate summaries, upload analytics). This triage approach means you accomplish the essentials even under extreme time pressure.
Registering Stop Hooks
Stop hooks live in .claude/hooks/stop/ directory. Claude Code discovers and loads them automatically:
.claude/
└── hooks/
└── stop/
├── flush-and-save-hook.mjs
├── database-cleanup-hook.mjs
├── checkpoint-save-hook.mjs
└── metrics-flush-hook.mjs
Each file exports a handler function. The filename doesn’t matter; Claude Code runs all stop hooks in the directory. Order is not guaranteed, so make each hook independent. This means your hooks can’t assume other hooks have already run, but it also means you can add new hooks without affecting existing ones.
Best Practices for Stop Hooks
Keep them fast: Stop hooks have a timeout. Each millisecond counts. Don’t do complex work. Offload heavy operations to startup next time. A 100ms operation is acceptable. A 500ms operation better be critical. A 2000ms operation better be essential.
Make them idempotent: It’s possible a stop hook runs twice. Design them so running twice produces the same result as running once. If you write a checkpoint file, writing it twice should create the same file, not corrupt it. If you close a connection, closing an already-closed connection should not error out.
Fail gracefully: If a stop hook fails, Claude Code still exits. Don’t throw errors for non-critical failures. Log them and return what you can. One cleanup failure shouldn’t cascade into other failures.
Don’t generate tokens: Stop hooks run in the context of Claude exiting. Avoid operations that might trigger new Claude reasoning. Call external APIs, write files, close connections. Don’t ask Claude questions. This discipline keeps cleanup fast.
Test them: Stop hooks are hard to test because they run at exit. Write unit tests for the actual functions. Test the error paths. Verify they work when resources are missing. Create staging tests where you simulate cleanup conditions.
Monitor them: Log what your stop hooks do. When a session ends, check the logs to see if cleanup worked. Over time, you’ll spot patterns of what needs improvement. If you see certain errors recurring, fix them. If you see certain operations always timing out, optimize them.
Common Pitfalls and How to Avoid Them
Pitfall 1: Forgetting to Handle the Context Window Full Scenario
When context is full, you have very little time. Some developers write stop hooks that assume plenty of time. Then they timeout in production when context gets tight.
Solution: Design stop hooks to work with minimal context. Check context.timeRemaining and adapt. Some operations are deferred if time is tight. Always have a “must save” list that you prioritize over “nice to have.”
Pitfall 2: Blocking on I/O Without Timeouts
A stop hook that calls a database without a timeout can hang indefinitely. Then Claude Code can’t exit cleanly.
// ❌ Dangerous: Can hang forever
await database.query("FLUSH TABLES");
// ✓ Safe: Times out after 2 seconds
await Promise.race([
database.query("FLUSH TABLES"),
new Promise((_, reject) =>
setTimeout(() => reject(new Error("Timeout")), 2000)
)
]);
Every I/O operation needs a timeout. Period. This is non-negotiable.
Pitfall 3: Not Handling Missing Resources Gracefully
If the database connection is gone, your shutdown hook shouldn’t crash trying to close it.
// ❌ Dangerous: Assumes connection exists
await global.dbConnection.end();
// ✓ Safe: Checks first
if (global.dbConnection && typeof global.dbConnection.end === 'function') {
try {
await global.dbConnection.end();
} catch (err) {
// Connection was probably already gone, that's ok
}
}
Defense in depth: check existence, check type, wrap in try-catch, log errors but don’t throw.
Pitfall 4: Generating Output During Shutdown
Stop hooks should be quiet. They shouldn’t send data to stdout or generate new work. Every byte of output might use tokens that Claude needs for other things.
// ❌ Problematic: Outputs to console
console.log("Flushing logs...");
await flushLogs();
console.log("Done!");
// ✓ Better: Silent execution
await writeToLogFile("Shutdown sequence started");
Write to log files. Don’t write to stdout. Save logging for later analysis.
Conclusion: Building Reliable Systems
Stop hooks are the difference between graceful shutdown and chaotic exit. They’re simple to write, run automatically, and protect your systems when Claude Code stops. Invest time in building them. Test them. Monitor what they do. Over time, they become an essential part of your reliable Claude Code infrastructure.
The moment Claude Code stops is the moment you most need to ensure everything is clean. Stop hooks make that possible. From flushing buffers to notifying external systems to generating audit trails, stop hooks handle all the housekeeping that happens when work is done. Think of them as your safety net—they catch you on the way down and ensure you land safely.
Build them with intention. Make them fast. Make them idempotent. Make them observable. And when Claude Code shuts down at 3 AM because something went wrong, you’ll have the confidence that everything was cleaned up properly. Your logs are flushed. Your databases are disconnected. Your state is saved. Your system is ready for the next session.
That’s the reliability that matters in production.
Real-World Scenarios: Where Stop Hooks Prevent Disasters
Understanding stop hooks conceptually is one thing. Seeing them prevent actual disasters is more convincing. Let’s walk through several realistic scenarios where good stop hooks made the difference between graceful shutdown and catastrophic failure.
Scenario 1: The Incomplete Migration Recovery
A data engineering team was using Claude Code to migrate 200 million customer records from an old database to a new one. The process was taking about 10 hours. The migration worked by: read batch of 10,000 records, transform, insert, checkpoint, repeat.
After 6 hours, the migration had successfully processed 120 million records. The system logged a checkpoint every million records. Then, Claude Code hit a context limit and had to stop. Without stop hooks, the team would have had to restart from zero. That’s 6 wasted hours.
With a stop hook that saves checkpoints, here’s what happened: Claude Code stopped, the stop hook ran, it saved a checkpoint file with: “last processed ID: 120000000, items_processed: 120000000, items_failed: 0, resume_from: 120000001”. The team restarted Claude Code with the checkpoint file, and it picked up from record 120000001. The remaining 80 million records were processed in the next 5 hours. Without the stop hook, this would have been a 15+ hour job. With the stop hook, it was 11 hours total.
The difference is the checkpoint stop hook. It transformed an interruption (context window full) from a disaster (start over) into a minor inconvenience (resume from checkpoint).
Scenario 2: The Database Connection Crisis
A backend team was using Claude Code to perform a complex database refactoring. They opened a connection to their PostgreSQL database. Claude Code ran successfully for 45 minutes, executing migrations, creating indexes, refactoring data. Then it hit a timeout and had to exit abruptly.
Without a database cleanup stop hook, the PostgreSQL connection would have been left open. It would have held a transaction lock. Any subsequent connection attempts by the application would hang waiting for the lock. The application would become unresponsive. Customers would experience downtime. The on-call engineer would wake up at 3 AM to an alert. They’d have to manually kill the PostgreSQL connection. By the time they found and fixed the problem, there would be an incident review, post-mortem, and a week of stress.
With a database cleanup stop hook, here’s what happened: Claude Code stopped. The stop hook ran automatically. It detected the open connection. It sent ROLLBACK to abort the in-flight transaction. It closed the connection gracefully. The application never knew there was an issue. No downtime. No incident. No 3 AM alerts.
The cost of building this stop hook: 15 minutes. The cost of the incident it prevented: 6+ hours of engineer time, customer impact, incident review. The ROI is staggering.
Scenario 3: The Metrics That Disappeared
A DevOps team used Claude Code to run a daily infrastructure audit. The script checked all their cloud resources, validated configurations, and collected metrics about what was running and how it was configured. The metrics were important—they fed into dashboards that the team used for decision-making.
One day, Claude Code exited unexpectedly (network timeout, or context window, or error). The metrics that had been collected during the run were buffered in memory. They were never written to disk. They were lost. The dashboard showed no data for that day. The team had a gap in their time-series data that they’d need to explain to stakeholders.
With a metrics flush stop hook, here’s what happened: Claude Code stopped. The stop hook ran. It detected that metrics were buffered. It wrote them to disk and uploaded them to their metrics database. The dashboard showed complete data for the day. No gaps. Continuity.
This isn’t a catastrophe, but it’s the kind of preventable data loss that compounds over time. If it happens regularly, your metrics become untrustworthy. Users stop relying on them. Stop hooks prevent this erosion of data integrity.
Advanced Patterns: Multi-Stage Cleanup
As your systems grow more sophisticated, stop hooks need to handle more complex scenarios. Multi-stage cleanup is a pattern for this:
A multi-stage cleanup prioritizes operations by importance. Stage 1 (critical): must happen even if time is tight. Stage 2 (important): should happen if time permits. Stage 3 (nice-to-have): only if plenty of time.
// Multi-stage cleanup example
export const handler = async (context) => {
const deadline = Date.now() + context.timeRemaining;
// Stage 1: CRITICAL (must always run, even with <1s remaining)
const stage1 = await executeCritical(deadline);
// Stage 2: IMPORTANT (run if we have >2s remaining)
const stage2 = Date.now() < deadline - 2000
? await executeImportant(deadline)
: null;
// Stage 3: NICE-TO-HAVE (run if we have >5s remaining)
const stage3 = Date.now() < deadline - 5000
? await executeNiceToHave(deadline)
: null;
return {
success: stage1.success && (stage2?.success ?? true),
stages: { stage1, stage2, stage3 }
};
};
Critical operations (flush critical state, close database connections, save essential checkpoints) always run. Important operations (update metrics, write logs) run if time permits. Nice-to-have operations (generate summary, upload analytics) only run if time is abundant. This triage approach ensures that even in the worst case (very little time before forceful shutdown), the essential cleanup happens.
Monitoring Stop Hook Behavior: Building Observability
Stop hooks run at exit, which makes them hard to observe. But observability is crucial—you need to know if your stop hooks are working correctly. Here are strategies for visibility:
Write detailed logs: Stop hooks should log everything they do. Create .claude/logs/stop-hook-execution.log and write to it:
2026-03-17T14:32:15Z STOP_HOOK_START reason=context_full time_remaining=3500ms
2026-03-17T14:32:15Z STAGE_1_START Critical cleanup
2026-03-17T14:32:15Z DATABASE Closing connection pool
2026-03-17T14:32:16Z DATABASE Pool closed successfully (took 450ms)
2026-03-17T14:32:16Z STATE Saving checkpoint file
2026-03-17T14:32:16Z STATE Checkpoint saved: /path/to/checkpoint.json (took 120ms)
2026-03-17T14:32:17Z STAGE_1_COMPLETE All critical operations succeeded
2026-03-17T14:32:17Z STAGE_2_SKIP Insufficient time remaining (have 1500ms, need 2000ms)
2026-03-17T14:32:17Z STOP_HOOK_COMPLETE success=true duration=1800ms stages_run=1
These logs are your debugging tool. When something goes wrong, you read the logs to see what happened during shutdown.
Metrics about stop hooks: Track execution time, what stages ran, what stages were skipped, what failed. After a month, you see patterns: “Database cleanup takes an average of 600ms, 95th percentile is 1200ms.” This tells you whether your timeout budget is adequate.
Regular testing: In your staging environment, occasionally trigger stop scenarios manually. Kill the Claude Code session. Check the logs. Verify that cleanup ran correctly. This gives you confidence that stop hooks will work when you really need them.
Common Pitfalls in Stop Hook Design
Beyond the pitfalls we covered earlier, there are more subtle ones that teams encounter:
Pitfall 5: Assuming Cleanup Can Use Claude Reasoning
Stop hooks shouldn’t call back into Claude to reason about what to do. Cleanup needs to be deterministic and fast. Don’t write:
// ❌ Dangerous: Uses Claude to decide what to do
const decision = await askClaude("Should I save this data or discard it?");
if (decision === "save") {
await saveData();
}
This defeats the purpose of stop hooks. You’re in a degraded state trying to exit. Calling Claude creates circular dependencies and can cause timeout loops. Instead, make cleanup decisions based on simple logic:
// ✓ Safe: Deterministic, no Claude involvement
if (global.unsavedData && global.unsavedData.length > 0) {
await saveData();
}
Pitfall 6: Not Testing Degraded Conditions
Stop hooks need to work when the system is broken. Test them with broken database connections, missing files, permission errors:
// Test with broken connection
global.dbConnection = { end: async () => { throw new Error("Connection refused"); } };
// Run hook
const result = await handler(context);
// Verify it handles the error gracefully
assert(result.success === false);
assert(result.results.errors.length > 0);
Testing stop hooks in ideal conditions doesn’t prepare you for reality. Degraded systems are the only time stop hooks really matter.
Pitfall 7: Memory Leaks in Stop Hooks
If you’re not careful, your stop hook cleanup can leave things in memory. Don’t accumulate state that isn’t cleaned:
// ❌ Dangerous: Accumulates state
global.cleanupLog = global.cleanupLog || [];
global.cleanupLog.push("Cleaned up database"); // Grows over time
// ✓ Safe: Discards old state
const cleanupLog = [];
cleanupLog.push("Cleaned up database");
// cleanupLog gets garbage collected when function ends
Over many sessions, accumulating state can consume memory. Stop hooks should leave the system as clean as possible, not add to the mess.
Building Team Practices Around Stop Hooks
Stop hooks are infrastructure, but they require team practices to be effective. Here are practices that mature teams follow:
Everyone understands why they exist: When you introduce stop hooks to a team, explain the scenarios they prevent. Show real examples of what happens without them. When people understand the “why,” they take them seriously.
Stop hooks are owned by the team that created the system: If the data migration system creates state during its run, the team that owns the migration system owns the stop hook that cleans it up. Ownership prevents orphaned stop hooks that nobody maintains.
Stop hooks are versioned with the system: When you update the system that generates state, update the stop hook that cleans it up. They’re paired. If you change how state is structured, update the stop hook to handle both old and new formats.
Stop hooks are tested regularly: Don’t assume they work. Test them. Staging tests, simulation tests, whatever you can do to verify. A stop hook that hasn’t been tested is a stop hook waiting to fail.
Stop hooks have fallbacks: If the primary cleanup can’t run, what’s the fallback? If you can’t reach the database, can you write state to a local file instead? Document the fallback behavior.
Alerts when stop hooks fail: If a stop hook fails, you want to know. Log it prominently. Alert if the failure is critical (database cleanup failed, state not saved). This creates visibility so you can fix problems.
Measuring Stop Hook Effectiveness
After implementing stop hooks, measure their impact:
- Cleanups succeeded: How many times did stop hooks run and complete successfully?
- Data loss prevented: How much data would have been lost without stop hooks?
- Checkpoints saved and resumed: How many long-running operations were resumed from checkpoints instead of restarted?
- Incident response improved: Are investigations faster because of audit logs and metrics from stop hooks?
Track these metrics for a few months. You’ll see that stop hooks are preventing problems that would otherwise be invisible. The metrics justify the investment in building them.
Implementing Stop Hooks at Enterprise Scale
When you’re managing Claude Code across large organizations with dozens of active sessions running simultaneously, stop hooks become critical infrastructure. At scale, stop hooks protect you from cascading failures. One developer’s session ending badly shouldn’t impact the other ninety-nine sessions running in parallel. This is where enterprise-grade stop hook design matters.
The core challenge: you need hooks that handle not just individual service failures, but coordinated shutdown across multiple systems. A typical large deployment might have database connections, cache servers, message queues, file storage, monitoring systems, and analytics platforms all running simultaneously. When Claude Code stops, you need to gracefully disconnect from all of them without creating orphaned resources, dangling transactions, or incomplete writes.
Enterprise stop hooks implement a multi-phase shutdown. Phase 1 (critical and immediate): close connections in reverse order of their dependencies. If cache depends on database, close cache first, then database. Phase 2 (important but non-critical): flush logs, write final state, trigger monitoring alerts. Phase 3 (nice-to-have): generate summaries, upload analytics, clean temporary files. If you run out of time, you’ll have completed phases 1 and 2, leaving your systems in a consistent state. That’s what enterprise-grade means—guaranteed consistency even when deadlines are tight.
A real enterprise hook might coordinate shutdown across five different services:
// Enterprise shutdown coordination
const serviceOrder = ["cache", "queue", "database", "storage", "analytics"];
async function shutdownServices(context) {
for (const service of serviceOrder) {
const handler = serviceHandlers[service];
try {
await withTimeout(handler.shutdown(), handler.timeout);
} catch (err) {
logCritical(`${service} shutdown failed: ${err.message}`);
// Continue shutdown of remaining services
}
}
}
At enterprise scale, you measure and monitor stop hook performance obsessively. Which services are slow? Which fail most often? Which timeout patterns appear? Track these metrics over weeks and months. You’ll see patterns emerge. Maybe database shutdown averages 1.2 seconds but has a 95th percentile of 3.5 seconds, suggesting occasional deadlocks. Maybe analytics uploads are timing out 2% of the time, suggesting rate limiting issues. This data drives improvements. You optimize the slow services. You add retry logic for flaky ones. You adjust timeouts based on actual behavior, not guesses.
At this scale, stop hooks also become audit infrastructure. Every shutdown is logged in detail. Later, if something goes wrong—a support ticket says “why is my data inconsistent?”—you can review the stop hook logs from the sessions that touched that data. You can see exactly what happened during shutdown. This traceability is essential for incident investigation and for regulatory compliance in some organizations.
-iNet
Clean exits build reliable systems.