Your workflow is humming along beautifully. Tests pass, PRs get reviewed in seconds, and Claude Code is automatically fixing bugs before they hit production. Then, at 3 AM on a Saturday, the API goes offline for 45 minutes—and suddenly your entire automation pipeline fails. Your workflow crashes. No retry happens. No fallback kicks in. You wake up to 200 failed actions and a backlog of work that could have been handled gracefully.
The difference between a resilient system and a brittle one often comes down to one thing: what happens when things break. In this guide, we’ll walk through the reality of running Claude Code in GitHub Actions—the ways it fails, why it fails, and most importantly, how to build workflows that don’t just break, but bounce back. -iNet
Why Error Handling Matters in GitHub Actions
Here’s the uncomfortable truth: any system that calls an external API will eventually encounter errors. The Claude Code API is incredibly reliable, but it’s still subject to rate limits, occasional outages, network timeouts, and the random hiccups that affect every cloud service. Your GitHub Actions workflow sits at the intersection of two different systems—GitHub’s runners and Anthropic’s API—which means twice as many things that can go wrong.
Without proper error handling, a single API error kills your entire workflow. Your code reviews don’t happen. Your documentation doesn’t update. Your automated fixes don’t run. And worse, you often don’t know why until you dig through logs hours later. You might miss a critical security issue because the automated review crashed silently. You might let a buggy commit merge because the linter never ran.
With proper error handling, you get graceful degradation. Transient errors retry automatically. Persistent failures fall back to safer alternatives. You get visibility into what happened. And your team keeps moving forward instead of being blocked.
The good news? Building resilient workflows isn’t complicated. It requires understanding what can fail, categorizing those failures, and implementing strategies that match each failure type. The operational benefit isn’t just technical—it’s about trust. When your CI/CD system is reliable, your developers rely on it. It becomes part of your development velocity instead of a source of friction and frustration.
Understanding Error Types in Claude Code Workflows
Not all errors are created equal. A timeout is different from a rate limit, which is different from an authentication failure. Your retry strategy needs to match the error type, or you’ll just waste time and money retrying things that won’t succeed.
Transient Errors are temporary problems that usually resolve themselves if you retry:
- 429 (Too Many Requests): You’ve hit a rate limit. The API tells you to wait, then try again. This is common during traffic spikes or when multiple workflows hit the API simultaneously.
- 502/503 (Bad Gateway / Service Unavailable): The API is temporarily down or unreachable. A load balancer is struggling. A deployment is rolling out. Wait a bit, it’ll come back.
- timeout errors: The request took too long. Network hiccups, temporary load spikes, or just the random latency you get with cloud services. Waiting and retrying usually succeeds.
- Connection resets: The connection dropped mid-request. TCP hiccup. DNS flake. Retry it.
- 500 (Internal Server Error): Something went wrong on the server side, but it might be transient. A worker crashed but others are fine. A cache invalidation is in progress.
These errors are not your fault. You didn’t send a bad request. Your credentials are fine. You just have to wait and try again.
Permanent Errors won’t go away no matter how many times you retry:
- 400 (Bad Request): Your request was malformed. You sent invalid JSON, missing required fields, or a body that doesn’t match the schema. Fix the input, not the retry logic. Retrying the same bad request 100 times won’t magically fix it.
- 401 (Unauthorized): Your API key is invalid, missing, or expired. Check your credentials. Rotate your token. Retrying won’t help if the key itself is wrong.
- 403 (Forbidden): Your API key lacks permissions for this operation. Your organization has rate limits or usage restrictions. Contact Anthropic support to elevate permissions.
- 404 (Not Found): The endpoint doesn’t exist. Check your URL. Maybe you’re targeting the wrong API version or the endpoint was deprecated.
These errors mean: stop trying. Fix the underlying problem, or acknowledge that this operation will never succeed.
The Rate Limit Error deserves special mention because it’s in a weird middle ground. It’s technically “permanent” in that moment (the API will reject your request right now), but it’s “transient” in that waiting a bit and retrying will succeed. This is where the Retry-After header becomes your best friend.
The Anthropic API includes a Retry-After header in 429 responses that tells you exactly how many seconds to wait. This is gold. When you see 429, look at that header, wait that long, then retry. No guessing. No exponential backoff arithmetic. Just follow the header. If you ignore it and implement your own backoff, you might retry too quickly and hit rate limits again.
The Cost of Retry Strategies
Before diving into code, let’s understand why retry strategy matters economically. Each request to Claude Code’s API costs API credits. Dumb retries that don’t respect rate limits waste those credits. If you’re retry-looping rapidly against a rate limit, you might burn through your monthly credits on retries that will never succeed.
Exponential backoff (waiting 1 second, then 2, then 4, then 8) is efficient because it:
- Respects the API: You’re not hammering it while it’s already struggling
- Saves credits: Failed requests that time out waste less than rapid retry spam
- Gives transient problems time to recover: A 503 that lasts 3 seconds recovers while you’re sleeping for 4 seconds
- Is humane to runners: GitHub Actions runners that spend time sleeping don’t burn compute
But there’s a catch: exponential backoff can be too patient. If a transient error lasts 30 seconds, waiting 1→2→4→8→16 seconds (31 seconds total) might timeout your workflow. That’s why the Retry-After header is a lifesaver—it tells you exactly how long the API needs, not a guess.
Setting Up Error Handling in Your Workflow
Let’s build a real workflow that handles errors gracefully. We’ll start with the skeleton, then add error handling layer by layer.
name: Claude Code with Error Handling
on:
pull_request:
types: [opened, synchronize]
jobs:
code-review:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Run Claude Code Review
uses: anthropics/claude-code-action@v1
with:
prompt: "Review this pull request for code quality and security issues."
api_key: ${{ secrets.ANTHROPIC_API_KEY }}
This is a starting point, but it has no error handling. If the Claude Code action fails, the workflow stops. The job is marked as failed. All downstream steps don’t run. Let’s add a timeout first—this prevents hanging workflows that might be waiting on a network issue:
- name: Run Claude Code Review
uses: anthropics/claude-code-action@v1
with:
prompt: "Review this pull request for code quality and security issues."
api_key: ${{ secrets.ANTHROPIC_API_KEY }}
timeout-minutes: 10
Now we’re saying: “This step should complete in 10 minutes or fail.” But what happens when it fails? Nothing yet. Let’s add a continue-on-error flag so the workflow doesn’t stop:
- name: Run Claude Code Review
id: claude-review
uses: anthropics/claude-code-action@v1
with:
prompt: "Review this pull request for code quality and security issues."
api_key: ${{ secrets.ANTHROPIC_API_KEY }}
timeout-minutes: 10
continue-on-error: true
Now the workflow keeps going even if Claude Code fails. We’ve added an id: claude-review so we can reference whether this step succeeded or failed in later steps. Here’s where it gets interesting—we check the outcome and implement fallback logic:
- name: Handle Claude Review Failure
if: steps.claude-review.outcome == 'failure'
run: |
echo "Claude Code review failed. Creating a comment requesting manual review."
# Use GitHub CLI to post a comment
gh pr comment "${{ github.event.pull_request.number }}" \
--body "Automated code review encountered an error. Manual review requested."
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
This is the essence of graceful degradation: when automated review fails, we don’t silently skip it. We create an explicit signal so the team knows to investigate. The PR doesn’t merge without review; it just gets human review instead of automated review.
The Psychology of Failures: Why Teams Ignore Error Handling
Before we implement error handling, let’s be honest about why teams often skip it. Error handling feels optional. Your workflow works fine 99% of the time. Adding retry logic and fallback paths feels like gold-plating—unnecessary complexity for edge cases that probably won’t happen.
This is a dangerous assumption. That 1% failure rate—when it happens at scale, it’s catastrophic. If you run 1000 workflows per day and each has a 1% failure rate, that’s 10 failures per day. 70 per week. 3600 per year. Even a 0.1% failure rate results in daily failures.
More importantly, failures compound. If your code review fails silently, a bug merges. That bug causes a production outage, which triggers an incident, which requires debugging, which delays feature development, which angers customers. The cost of that single silent failure: days of engineering time. The cost of building error handling: a few hours.
The ROI on error handling is massive. But it requires thinking probabilistically. You have to accept that your system will fail and plan for it. Most engineers aren’t trained to think this way—training emphasizes happy paths, not failure paths.
Here’s the shift in mindset: Error handling isn’t about preventing errors (you can’t). It’s about controlling what happens when they occur. Every path through your system should be thought through: success, transient failure, permanent failure. What happens in each case?
Implementing Retry Logic with Exponential Backoff
Most transient errors benefit from retries. But dumb retries—just hammering the API again immediately—are wasteful and might make rate limiting worse. Exponential backoff is the pattern where you wait longer between each retry: wait 1 second, then 2 seconds, then 4 seconds, then 8 seconds, and so on. This gives temporary problems time to resolve while respecting rate limits.
GitHub Actions doesn’t have built-in exponential backoff, but we can implement it using bash. Here’s a reusable function:
#!/bin/bash
retry_with_backoff() {
local max_attempts=5
local timeout=1
local attempt=1
local statusCode=0
while [ $attempt -le $max_attempts ]; do
echo "Attempt $attempt of $max_attempts..."
# Run the command
if "$@"; then
return 0
else
statusCode=$?
fi
if [ $attempt -lt $max_attempts ]; then
echo "Command failed. Waiting ${timeout}s before retry..."
sleep $timeout
timeout=$((timeout * 2))
fi
attempt=$((attempt + 1))
done
echo "Command failed after $max_attempts attempts."
return $statusCode
}
# Usage:
retry_with_backoff curl -X POST https://api.example.com/review \
-H "Authorization: Bearer $ANTHROPIC_API_KEY" \
-d '{"code": "..."}'
The function doubles the wait time after each failure: 1s → 2s → 4s → 8s → 16s. The "$@" syntax runs whatever command you pass it. So retry_with_backoff my-command arg1 arg2 will run my-command arg1 arg2 with retries.
In a GitHub Actions workflow, you’d use it like this:
- name: Download Retry Script
run: |
cat > retry.sh << 'EOF'
#!/bin/bash
retry_with_backoff() {
local max_attempts=5
local timeout=1
local attempt=1
while [ $attempt -le $max_attempts ]; do
if "$@"; then return 0; fi
if [ $attempt -lt $max_attempts ]; then
sleep $timeout
timeout=$((timeout * 2))
fi
attempt=$((attempt + 1))
done
return 1
}
retry_with_backoff "$@"
EOF
chmod +x retry.sh
- name: Run Claude Review with Retry
id: claude-review
run: |
./retry.sh curl -X POST \
-H "Authorization: Bearer ${{ secrets.ANTHROPIC_API_KEY }}" \
-H "Content-Type: application/json" \
-d @- \
https://api.anthropic.com/v1/messages << 'REQUEST'
{
"model": "claude-opus-4-1",
"max_tokens": 2048,
"messages": [{"role": "user", "content": "Review this code..."}]
}
REQUEST
continue-on-error: true
Notice the continue-on-error: true at the end. This means even if all retries fail, the workflow continues (we’ll handle the failure in the next step).
What Success Actually Looks Like
Before implementing retry strategies, let’s define what we’re aiming for. A “successful” workflow isn’t one that never fails—it’s one where:
- Failures are visible: When something goes wrong, someone knows about it
- Failures are handled gracefully: Transient errors are retried, permanent errors escalate
- Failures don’t cascade: A failure in one part doesn’t break unrelated parts
- Failures are informative: When things break, there’s enough context to understand why
This means your workflow design should answer:
- What happens if Claude Code times out?
- What happens if the API returns 429?
- What happens if the user’s API key is invalid?
- What happens if we can’t authenticate to GitHub?
- What happens if a file we’re trying to read doesn’t exist?
- What happens if we need to write to a directory but it’s not writable?
Most teams never ask these questions. They assume success. Then they’re shocked when something fails.
Respecting Rate Limits and the Retry-After Header
The Anthropic API will tell you when to retry using the Retry-After header. Ignoring this header and implementing your own backoff is stubborn and inefficient. Here’s how to respect it:
#!/bin/bash
call_claude_api() {
local url="$1"
local data="$2"
local attempt=1
local max_attempts=5
while [ $attempt -le $max_attempts ]; do
echo "Calling Claude API (attempt $attempt)..."
response=$(curl -s -w "\n%{http_code}" -X POST "$url" \
-H "Authorization: Bearer $ANTHROPIC_API_KEY" \
-H "Content-Type: application/json" \
-d "$data")
# Split response and status code
body=$(echo "$response" | head -n -1)
status=$(echo "$response" | tail -n 1)
if [ "$status" = "200" ]; then
echo "$body"
return 0
elif [ "$status" = "429" ]; then
# Rate limited - extract Retry-After header
retry_after=$(curl -s -I -X POST "$url" \
-H "Authorization: Bearer $ANTHROPIC_API_KEY" \
-H "Content-Type: application/json" \
-d "$data" | grep -i "retry-after" | cut -d' ' -f2 | tr -d '\r')
if [ -z "$retry_after" ]; then
retry_after=5 # Default to 5s if header is missing
fi
echo "Rate limited. Waiting ${retry_after}s..." >&2
sleep "$retry_after"
attempt=$((attempt + 1))
elif [ "$status" = "500" ] || [ "$status" = "502" ] || [ "$status" = "503" ]; then
# Server error - exponential backoff
wait_time=$((2 ** (attempt - 1)))
echo "Server error ($status). Waiting ${wait_time}s..." >&2
sleep "$wait_time"
attempt=$((attempt + 1))
else
# Client error - don't retry
echo "Client error ($status). Not retrying." >&2
echo "$body" >&2
return 1
fi
done
echo "Failed after $max_attempts attempts." >&2
return 1
}
This function is smarter about retries:
- 200: Success, return immediately.
- 429: Read the Retry-After header, wait that long, retry.
- 5xx errors: Exponential backoff (server errors often resolve themselves).
- 4xx errors: Don’t retry—they won’t succeed no matter how many times you try.
The key insight is that the Retry-After header is your source of truth for rate limits. Don’t invent your own backoff strategy; let the API tell you when to try again.
Understanding When to Retry vs. When to Fail Fast
This is the critical decision point that most implementations get wrong. You don’t retry everything. Some errors are worth retrying; others just waste time and resources.
Worth retrying:
- Network errors (connection refused, timeout, reset)
- Server errors (5xx status codes)
- Rate limit errors (429) when Retry-After header is present and reasonable
- Transient infrastructure issues
Not worth retrying:
- Authentication errors (401, 403)
- Malformed requests (400)
- Resource not found (404)
- Invalid API keys or tokens
The difference matters because retrying the wrong thing is expensive. If you’re retrying a 401 error 5 times, you’ve wasted 5 API calls that could have been used for something that works. You’ve delayed feedback to the user. And you’ve consumed GitHub Actions runner minutes (which have quotas).
A good rule: retry transient errors (those that happen due to temporary system state), but fail fast on permanent errors (those caused by incorrect configuration or logic).
#!/bin/bash
call_api_smart_retry() {
local max_attempts=5
local attempt=1
while [ $attempt -le $max_attempts ]; do
response=$(curl -s -w "\n%{http_code}" "$@")
status=$(echo "$response" | tail -n1)
body=$(echo "$response" | head -n-1)
case $status in
200|201|204)
# Success
echo "$body"
return 0
;;
400|401|403|404)
# Permanent error—don't retry
echo "Permanent error ($status): $body" >&2
return 1
;;
429)
# Rate limited—respect Retry-After
retry_after=$(echo "$response" | grep -i "retry-after" | cut -d' ' -f2)
sleep "${retry_after:-5}"
attempt=$((attempt + 1))
;;
500|502|503|504)
# Transient error—exponential backoff
if [ $attempt -lt $max_attempts ]; then
wait_time=$((2 ** (attempt - 1)))
sleep "$wait_time"
attempt=$((attempt + 1))
else
echo "Transient error ($status) after $max_attempts attempts" >&2
return 1
fi
;;
*)
# Unknown error—give up
echo "Unknown error ($status): $body" >&2
return 1
;;
esac
done
return 1
}
This function categorizes errors and makes the right decision about each one. Notice that 4xx errors (except 429) return immediately—there’s no point retrying. 5xx errors retry with exponential backoff. Rate limits retry respecting the header.
The Testing Nightmare: Actually Testing Error Paths
Here’s something that rarely gets discussed: how do you test your error handling? Most teams never do. They test the happy path (success), maybe test one failure case, but they never test:
- What happens when retries are exhausted?
- What happens when both retry and fallback fail?
- What happens when the fallback takes longer than the timeout?
- What happens when rate limiting hits?
You can’t reliably test these against a live API. But you can mock:
# test-error-handling.yml
name: Test Error Handling
on: [push]
jobs:
test-retry-logic:
runs-on: ubuntu-latest
steps:
# Start a mock server that simulates errors
- name: Start Mock API
run: |
# Start a server that returns 429 first, then 200
python3 << 'EOF'
from http.server import HTTPServer, BaseHTTPRequestHandler
import threading
import time
request_count = 0
class MockHandler(BaseHTTPRequestHandler):
def do_POST(self):
global request_count
request_count += 1
if request_count == 1:
# First request: rate limited
self.send_response(429)
self.send_header('Retry-After', '2')
self.end_headers()
else:
# Second request: success
self.send_response(200)
self.send_header('Content-Type', 'application/json')
self.end_headers()
self.wfile.write(b'{"success": true}')
def log_message(self, format, *args):
pass # Suppress logs
server = HTTPServer(('localhost', 8000), MockHandler)
thread = threading.Thread(target=server.serve_forever)
thread.daemon = True
thread.start()
time.sleep(1) # Give server time to start
EOF
&
# Wait for mock server to start
- name: Test Retry Logic
run: |
# This script retries after 429
./test-retry.sh https://automateanddeploy.com:8000/review
# Verify the retry succeeded
- name: Verify Result
run: |
if grep -q "success" retry_result.json; then
echo "✅ Retry logic works"
else
echo "❌ Retry logic failed"
exit 1
fi
This is more work than testing the happy path, but it’s essential. By mocking error responses, you can verify that your retry logic actually works before deploying.
Monitoring and Alerting: Knowing When Things Break
Error handling is only half the battle. You also need to know when errors occur. Even if your workflow handles failures gracefully, you need visibility into patterns.
- name: Monitor Workflow Health
if: always()
run: |
# Log workflow result for monitoring
curl -X POST https://monitoring.example.com/workflow-event \
-H "Authorization: Bearer ${{ secrets.MONITORING_TOKEN }}" \
-d '{
"workflow": "${{ github.workflow }}",
"status": "${{ job.status }}",
"run_id": "${{ github.run_id }}",
"attempt": ${{ github.run_attempt }},
"timestamp": "'$(date -u +%Y-%m-%dT%H:%M:%SZ)'",
"duration": '$WORKFLOW_DURATION'
}'
This sends data to your monitoring system about every workflow run. You can then:
- Set alerts: “If Claude Code reviews fail 3 times in a row, alert the team”
- Track trends: “Error rates are increasing on weekends—investigate”
- Correlate failures: “These reviews failed right after a deployment—related?”
- Understand patterns: “These specific error types fail 80% of the time and should be reclassified”
Over time, monitoring data informs your retry strategies. If you see that 429 errors consistently resolve within 2 seconds, you can tune your backoff. If you see that 500 errors never resolve and always require manual intervention, you can stop retrying and escalate immediately.
Understanding Cost: Retries and API Credits
This deserves its own section because retry costs are often invisible until they become catastrophic.
Each request to Claude Code’s API costs credits. Failed requests that timeout still cost credits. Retries that don’t respect rate limits compound the cost. Consider this scenario:
Initial request at capacity: fails with 429 (rate limited)
Retry without respecting Retry-After: fails, wastes credits
Retry again immediately: fails, wastes credits
Retry again: fails, wastes credits
...after 10 retries, you've burned 10x credits for a single request
Compare to:
Initial request: fails with 429, includes Retry-After: 5
Wait 5 seconds as instructed
Retry: succeeds, costs 1x credits
The difference is 10x. Multiply that by thousands of workflows, and you’re talking about massive cost differences. Smart retry strategies literally save money.
Moreover, rate-limited requests signal that you’re hitting API capacity. Respecting those signals isn’t just polite—it’s economically rational. If you’re rate-limited, retrying faster won’t help and will anger the API providers (and potentially trigger your own rate limits being reduced).
Building Resilience Across Your Infrastructure
Error handling in a single workflow is one thing. Building resilience across an entire CI/CD infrastructure is another. Here’s how to think about it systematically:
Layer 1: Individual Workflow Resilience (what we’ve covered)
- Timeouts prevent hanging
- Retries with backoff handle transient errors
- Graceful fallbacks when errors are permanent
- Monitoring to detect patterns
Layer 2: Cross-Workflow Communication
Multiple workflows might depend on each other. If one fails, others might be left in an inconsistent state. Use job outcomes to coordinate:
build-and-test:
runs-on: ubuntu-latest
outputs:
success: ${{ job.status }}
steps:
- uses: actions/checkout@v4
- run: npm test
deploy:
needs: build-and-test
if: needs.build-and-test.outputs.success == 'success'
runs-on: ubuntu-latest
steps:
- run: echo "Deploying..."
This ensures deployment only happens if tests pass. More importantly, it prevents a failed test from silently being ignored.
Layer 3: Failure Notifications
When something actually fails (not just retried but truly failed), humans need to know:
- name: Notify on Failure
if: failure()
uses: actions/github-script@v6
with:
script: |
github.rest.issues.createComment({
issue_number: context.issue.number,
owner: context.repo.owner,
repo: context.repo.repo,
body: `⚠️ Automated review failed. Manual review requested.\n\nError: ${{ job.status }}\nRun: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}`
})
This creates a visible signal in the PR that something went wrong, rather than letting a failed workflow disappear into logs.
Real-World Patterns: The Lifecycle of a Robust Workflow
Let’s trace what a real, production-grade workflow looks like end-to-end:
name: Claude Code Review with Resilience
on:
pull_request:
types: [opened, synchronize, reopened]
jobs:
automated-review:
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
# Get PR info (will be needed for error handling)
- name: Get PR Info
id: pr-info
uses: actions/github-script@v6
with:
script: |
const pr = context.payload.pull_request;
core.setOutput('number', pr.number);
core.setOutput('title', pr.title);
# Run review with explicit error handling
- name: Run Claude Code Review
id: review
uses: anthropics/claude-code-action@v1
with:
prompt: "Perform a security and code quality review of this PR"
api_key: ${{ secrets.ANTHROPIC_API_KEY }}
timeout-minutes: 10
continue-on-error: true
# Categorize the failure
- name: Categorize Failure
id: categorize
if: steps.review.outcome == 'failure'
run: |
# Check logs to determine failure type
if grep -q "429" ${{ job.logs_file }}; then
echo "type=rate-limited" >> $GITHUB_OUTPUT
elif grep -q "timeout" ${{ job.logs_file }}; then
echo "type=timeout" >> $GITHUB_OUTPUT
elif grep -q "401\|403" ${{ job.logs_file }}; then
echo "type=auth-error" >> $GITHUB_OUTPUT
else
echo "type=unknown" >> $GITHUB_OUTPUT
fi
# Handle rate limiting with smart retry
- name: Retry if Rate Limited
if: steps.categorize.outputs.type == 'rate-limited'
run: |
echo "Rate limited. Waiting 30 seconds before retry..."
sleep 30
# Retry logic here (could trigger another workflow)
# Handle auth errors with alert
- name: Alert on Auth Error
if: steps.categorize.outputs.type == 'auth-error'
uses: actions/github-script@v6
with:
script: |
github.rest.issues.createComment({
issue_number: ${{ steps.pr-info.outputs.number }},
owner: context.repo.owner,
repo: context.repo.repo,
body: '⚠️ Authentication error. Check API credentials.'
})
# Fallback: request manual review
- name: Fallback to Manual Review
if: steps.review.outcome == 'failure'
uses: actions/github-script@v6
with:
script: |
github.rest.issues.createComment({
issue_number: ${{ steps.pr-info.outputs.number }},
owner: context.repo.owner,
repo: context.repo.repo,
body: '⚠️ Automated review could not complete. Manual review requested.'
})
# Post success comment
- name: Post Review Results
if: steps.review.outcome == 'success'
uses: actions/github-script@v6
with:
script: |
github.rest.issues.createComment({
issue_number: ${{ steps.pr-info.outputs.number }},
owner: context.repo.owner,
repo: context.repo.repo,
body: '✅ Automated review completed.'
})
This workflow:
- Handles different error types differently
- Never silently fails (always comments on the PR)
- Retries when appropriate (rate limiting)
- Alerts for critical errors (auth)
- Falls back to manual review
- Has clear success signals
The Hidden Layer: Why Systems Fail and How to Predict It
Error handling in infrastructure is ultimately about understanding failure modes. Every system has ways it can fail, and most teams don’t think clearly about this until after they’ve experienced the failure in production.
A mature approach to error handling requires you to think probabilistically about failures and categorize them systematically. Not “what can go wrong”—too vague. Instead: “What are the specific ways this integration can fail, and what should happen when it does?”
For a GitHub Actions workflow calling Claude Code, the failure modes include:
Network layer: The request never reaches the API. Connection timeout, DNS failure, network unreachable. These happen regularly, often microscopically. A request that usually takes 50ms occasionally takes 3 seconds. Sometimes that timeout boundary is hit. Your retry strategy should handle this gracefully.
API layer: The request reaches the API but the response indicates an error. 429 (rate limited), 500 (internal error), 503 (overloaded). These are “expected failures”—the API tells you explicitly something went wrong. Your code should read the response and categorize it correctly.
Silent failure: The request appears to succeed (HTTP 200) but the response is corrupted or incomplete. This is insidious because your monitoring might not detect it. You got a response, so you assume success. But the response doesn’t have the fields you expect. Or it has them but they’re malformed. Or the JSON is truncated. These require defensive parsing and validation.
Cascade failures: Failure in one step cascades to others. The code review fails silently (no comment on the PR). The PR merges anyway. The tests fail. You don’t know the review failed because there was no visible signal. This is why every failure path needs explicit communication—a comment, a log, an alert—something that makes the failure visible.
Compound failures: Multiple independent systems fail simultaneously. GitHub is down, Claude API is rate-limited, and your monitoring service is also having issues. Your workflow can’t use any of them. What’s the graceful degradation? You have to plan for this before it happens.
The probability of each failure type informs your strategy. If you’re seeing 429 rate limits once a week, that’s worth handling with intelligent retries. If you’re seeing them 10 times a day, you have a scaling problem that retries won’t solve. If you’ve never seen a network failure in production, but your error logs show your firewall drops 0.01% of outbound connections, you should still handle it—because at scale, 0.01% happens daily.
This thinking—probabilistic risk assessment—is what separates resilient systems from systems that happen to work until they don’t.
The Organizational Impact of Reliable Automation
Here’s something that’s often overlooked in discussions about error handling: when your automation is reliable, it changes how your organization operates. Developers stop being skeptical of automated workflows. They don’t add manual checkpoints “just in case.” They trust the system because the system has proven itself trustworthy through repeated handling of failures gracefully.
This trust is valuable. Teams that build reliable CI/CD pipelines ship faster. They’re confident in automated deployments because they’ve seen the system handle transient failures and recover. They don’t spend time debugging mysterious workflow failures because the workflows are transparent about what happened. The logs tell the story. The monitoring shows the trends. When something does fail, it’s recoverable because they’ve designed for it.
But this only happens if your error handling is systematic and well-designed. Random timeouts that sometimes work, retries that hammer the API, silent failures that nobody notices—these erode trust. Teams stop using automation. They do things manually to be safe. You’ve lost the efficiency gains.
Conversely, when error handling is solid, teams lean on automation. It becomes a multiplier for your engineering effectiveness. One person can orchestrate thousands of tests, thousands of reviews, thousands of deployments through well-designed automation. That’s the power of systems that fail gracefully.
Testing Your Error Handling Under Load
Here’s a practice that separates teams with truly resilient systems from teams that think their systems are resilient: chaos testing. You intentionally break things and verify your error handling works.
This doesn’t require much. Set up a staging environment. Point your workflows at a mock API that randomly returns errors. Run your entire workflow suite. Watch what breaks. Fix it. Repeat. You’re learning what your actual system does when things fail, rather than guessing based on code review.
The cheapest way to do this: write a simple HTTP server that returns errors based on environment variables. Run it locally. Point your CI/CD at it. See what your workflows do. Many teams discover their error handling doesn’t work as well as they thought. Maybe retries happen too fast. Maybe the error categorization misses a case. Maybe alerts aren’t firing. You find out in a safe environment before production surprises you.
This kind of testing feels tedious. But every minute you spend testing error handling prevents hours of debugging when the errors happen in production. And they will happen. Murphy’s Law isn’t a suggestion.
Wrapping Up: Building Resilience Across Your Infrastructure
Resilient systems don’t try to prevent all errors—they can’t. Instead, they prepare for errors when they come. You’ve learned how to categorize errors, implement smart retry logic that respects rate limits, build graceful fallbacks, and monitor for failures.
The key insight: error handling isn’t a feature to add at the end. It’s foundational. Every workflow that calls Claude Code or any external API needs:
- Timeouts to prevent hanging forever
- Proper categorization to distinguish transient from permanent errors
- Smart retries that respect API signaling (Retry-After header)
- Exponential backoff to avoid hammering a struggling API
- Graceful fallbacks so failures don’t cascade
- Visibility so you know when things break
- Monitoring so you can improve over time
Your workflow will survive—not because it never fails, but because when it does fail, it fails gracefully and bounces back.
The systems you build with these principles become antifragile. They don’t just survive failures; they improve from them. Each failure teaches you something about your system. Each error recovery demonstrates that your design is sound. Over time, your error handling becomes so solid that you stop worrying about whether your workflows will fail. You trust that they will, and you trust that they’ll recover. That’s the mark of truly professional infrastructure.
Start implementing these patterns today. Begin with timeouts and basic error categorization. Add smart retries. Layer in monitoring. Then expand gradually as you learn what your system actually needs. The investment you make in error handling now will pay dividends for years as your automation becomes the trusted backbone of your development process.
-iNet