All Articles Claude Code

Claude Code GitHub Actions Error Handling

Your workflow is humming along beautifully. Tests pass, PRs get reviewed in seconds, and Claude Code is automatically fixing bugs before they hit production.

Your workflow is humming along beautifully. Tests pass, PRs get reviewed in seconds, and Claude Code is automatically fixing bugs before they hit production. Then, at 3 AM on a Saturday, the API goes offline for 45 minutes—and suddenly your entire automation pipeline fails. Your workflow crashes. No retry happens. No fallback kicks in. You wake up to 200 failed actions and a backlog of work that could have been handled gracefully.

The difference between a resilient system and a brittle one often comes down to one thing: what happens when things break. In this guide, we’ll walk through the reality of running Claude Code in GitHub Actions—the ways it fails, why it fails, and most importantly, how to build workflows that don’t just break, but bounce back. -iNet

Why Error Handling Matters in GitHub Actions

Here’s the uncomfortable truth: any system that calls an external API will eventually encounter errors. The Claude Code API is incredibly reliable, but it’s still subject to rate limits, occasional outages, network timeouts, and the random hiccups that affect every cloud service. Your GitHub Actions workflow sits at the intersection of two different systems—GitHub’s runners and Anthropic’s API—which means twice as many things that can go wrong.

Without proper error handling, a single API error kills your entire workflow. Your code reviews don’t happen. Your documentation doesn’t update. Your automated fixes don’t run. And worse, you often don’t know why until you dig through logs hours later. You might miss a critical security issue because the automated review crashed silently. You might let a buggy commit merge because the linter never ran.

With proper error handling, you get graceful degradation. Transient errors retry automatically. Persistent failures fall back to safer alternatives. You get visibility into what happened. And your team keeps moving forward instead of being blocked.

The good news? Building resilient workflows isn’t complicated. It requires understanding what can fail, categorizing those failures, and implementing strategies that match each failure type. The operational benefit isn’t just technical—it’s about trust. When your CI/CD system is reliable, your developers rely on it. It becomes part of your development velocity instead of a source of friction and frustration.

Understanding Error Types in Claude Code Workflows

Not all errors are created equal. A timeout is different from a rate limit, which is different from an authentication failure. Your retry strategy needs to match the error type, or you’ll just waste time and money retrying things that won’t succeed.

Transient Errors are temporary problems that usually resolve themselves if you retry:

  • 429 (Too Many Requests): You’ve hit a rate limit. The API tells you to wait, then try again. This is common during traffic spikes or when multiple workflows hit the API simultaneously.
  • 502/503 (Bad Gateway / Service Unavailable): The API is temporarily down or unreachable. A load balancer is struggling. A deployment is rolling out. Wait a bit, it’ll come back.
  • timeout errors: The request took too long. Network hiccups, temporary load spikes, or just the random latency you get with cloud services. Waiting and retrying usually succeeds.
  • Connection resets: The connection dropped mid-request. TCP hiccup. DNS flake. Retry it.
  • 500 (Internal Server Error): Something went wrong on the server side, but it might be transient. A worker crashed but others are fine. A cache invalidation is in progress.

These errors are not your fault. You didn’t send a bad request. Your credentials are fine. You just have to wait and try again.

Permanent Errors won’t go away no matter how many times you retry:

  • 400 (Bad Request): Your request was malformed. You sent invalid JSON, missing required fields, or a body that doesn’t match the schema. Fix the input, not the retry logic. Retrying the same bad request 100 times won’t magically fix it.
  • 401 (Unauthorized): Your API key is invalid, missing, or expired. Check your credentials. Rotate your token. Retrying won’t help if the key itself is wrong.
  • 403 (Forbidden): Your API key lacks permissions for this operation. Your organization has rate limits or usage restrictions. Contact Anthropic support to elevate permissions.
  • 404 (Not Found): The endpoint doesn’t exist. Check your URL. Maybe you’re targeting the wrong API version or the endpoint was deprecated.

These errors mean: stop trying. Fix the underlying problem, or acknowledge that this operation will never succeed.

The Rate Limit Error deserves special mention because it’s in a weird middle ground. It’s technically “permanent” in that moment (the API will reject your request right now), but it’s “transient” in that waiting a bit and retrying will succeed. This is where the Retry-After header becomes your best friend.

The Anthropic API includes a Retry-After header in 429 responses that tells you exactly how many seconds to wait. This is gold. When you see 429, look at that header, wait that long, then retry. No guessing. No exponential backoff arithmetic. Just follow the header. If you ignore it and implement your own backoff, you might retry too quickly and hit rate limits again.

The Cost of Retry Strategies

Before diving into code, let’s understand why retry strategy matters economically. Each request to Claude Code’s API costs API credits. Dumb retries that don’t respect rate limits waste those credits. If you’re retry-looping rapidly against a rate limit, you might burn through your monthly credits on retries that will never succeed.

Exponential backoff (waiting 1 second, then 2, then 4, then 8) is efficient because it:

  1. Respects the API: You’re not hammering it while it’s already struggling
  2. Saves credits: Failed requests that time out waste less than rapid retry spam
  3. Gives transient problems time to recover: A 503 that lasts 3 seconds recovers while you’re sleeping for 4 seconds
  4. Is humane to runners: GitHub Actions runners that spend time sleeping don’t burn compute

But there’s a catch: exponential backoff can be too patient. If a transient error lasts 30 seconds, waiting 1→2→4→8→16 seconds (31 seconds total) might timeout your workflow. That’s why the Retry-After header is a lifesaver—it tells you exactly how long the API needs, not a guess.

Setting Up Error Handling in Your Workflow

Let’s build a real workflow that handles errors gracefully. We’ll start with the skeleton, then add error handling layer by layer.

name: Claude Code with Error Handling

on:
  pull_request:
    types: [opened, synchronize]

jobs:
  code-review:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Run Claude Code Review
        uses: anthropics/claude-code-action@v1
        with:
          prompt: "Review this pull request for code quality and security issues."
          api_key: ${{ secrets.ANTHROPIC_API_KEY }}

This is a starting point, but it has no error handling. If the Claude Code action fails, the workflow stops. The job is marked as failed. All downstream steps don’t run. Let’s add a timeout first—this prevents hanging workflows that might be waiting on a network issue:

- name: Run Claude Code Review
  uses: anthropics/claude-code-action@v1
  with:
    prompt: "Review this pull request for code quality and security issues."
    api_key: ${{ secrets.ANTHROPIC_API_KEY }}
  timeout-minutes: 10

Now we’re saying: “This step should complete in 10 minutes or fail.” But what happens when it fails? Nothing yet. Let’s add a continue-on-error flag so the workflow doesn’t stop:

- name: Run Claude Code Review
  id: claude-review
  uses: anthropics/claude-code-action@v1
  with:
    prompt: "Review this pull request for code quality and security issues."
    api_key: ${{ secrets.ANTHROPIC_API_KEY }}
  timeout-minutes: 10
  continue-on-error: true

Now the workflow keeps going even if Claude Code fails. We’ve added an id: claude-review so we can reference whether this step succeeded or failed in later steps. Here’s where it gets interesting—we check the outcome and implement fallback logic:

- name: Handle Claude Review Failure
  if: steps.claude-review.outcome == 'failure'
  run: |
    echo "Claude Code review failed. Creating a comment requesting manual review."
    # Use GitHub CLI to post a comment
    gh pr comment "${{ github.event.pull_request.number }}" \
      --body "Automated code review encountered an error. Manual review requested."
  env:
    GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}

This is the essence of graceful degradation: when automated review fails, we don’t silently skip it. We create an explicit signal so the team knows to investigate. The PR doesn’t merge without review; it just gets human review instead of automated review.

The Psychology of Failures: Why Teams Ignore Error Handling

Before we implement error handling, let’s be honest about why teams often skip it. Error handling feels optional. Your workflow works fine 99% of the time. Adding retry logic and fallback paths feels like gold-plating—unnecessary complexity for edge cases that probably won’t happen.

This is a dangerous assumption. That 1% failure rate—when it happens at scale, it’s catastrophic. If you run 1000 workflows per day and each has a 1% failure rate, that’s 10 failures per day. 70 per week. 3600 per year. Even a 0.1% failure rate results in daily failures.

More importantly, failures compound. If your code review fails silently, a bug merges. That bug causes a production outage, which triggers an incident, which requires debugging, which delays feature development, which angers customers. The cost of that single silent failure: days of engineering time. The cost of building error handling: a few hours.

The ROI on error handling is massive. But it requires thinking probabilistically. You have to accept that your system will fail and plan for it. Most engineers aren’t trained to think this way—training emphasizes happy paths, not failure paths.

Here’s the shift in mindset: Error handling isn’t about preventing errors (you can’t). It’s about controlling what happens when they occur. Every path through your system should be thought through: success, transient failure, permanent failure. What happens in each case?

Implementing Retry Logic with Exponential Backoff

Most transient errors benefit from retries. But dumb retries—just hammering the API again immediately—are wasteful and might make rate limiting worse. Exponential backoff is the pattern where you wait longer between each retry: wait 1 second, then 2 seconds, then 4 seconds, then 8 seconds, and so on. This gives temporary problems time to resolve while respecting rate limits.

GitHub Actions doesn’t have built-in exponential backoff, but we can implement it using bash. Here’s a reusable function:

#!/bin/bash

retry_with_backoff() {
  local max_attempts=5
  local timeout=1
  local attempt=1
  local statusCode=0

  while [ $attempt -le $max_attempts ]; do
    echo "Attempt $attempt of $max_attempts..."

    # Run the command
    if "$@"; then
      return 0
    else
      statusCode=$?
    fi

    if [ $attempt -lt $max_attempts ]; then
      echo "Command failed. Waiting ${timeout}s before retry..."
      sleep $timeout
      timeout=$((timeout * 2))
    fi

    attempt=$((attempt + 1))
  done

  echo "Command failed after $max_attempts attempts."
  return $statusCode
}

# Usage:
retry_with_backoff curl -X POST https://api.example.com/review \
  -H "Authorization: Bearer $ANTHROPIC_API_KEY" \
  -d '{"code": "..."}'

The function doubles the wait time after each failure: 1s → 2s → 4s → 8s → 16s. The "$@" syntax runs whatever command you pass it. So retry_with_backoff my-command arg1 arg2 will run my-command arg1 arg2 with retries.

In a GitHub Actions workflow, you’d use it like this:

- name: Download Retry Script
  run: |
    cat > retry.sh << 'EOF'
    #!/bin/bash
    retry_with_backoff() {
      local max_attempts=5
      local timeout=1
      local attempt=1
      while [ $attempt -le $max_attempts ]; do
        if "$@"; then return 0; fi
        if [ $attempt -lt $max_attempts ]; then
          sleep $timeout
          timeout=$((timeout * 2))
        fi
        attempt=$((attempt + 1))
      done
      return 1
    }
    retry_with_backoff "$@"
    EOF
    chmod +x retry.sh

- name: Run Claude Review with Retry
  id: claude-review
  run: |
    ./retry.sh curl -X POST \
      -H "Authorization: Bearer ${{ secrets.ANTHROPIC_API_KEY }}" \
      -H "Content-Type: application/json" \
      -d @- \
      https://api.anthropic.com/v1/messages << 'REQUEST'
    {
      "model": "claude-opus-4-1",
      "max_tokens": 2048,
      "messages": [{"role": "user", "content": "Review this code..."}]
    }
    REQUEST
  continue-on-error: true

Notice the continue-on-error: true at the end. This means even if all retries fail, the workflow continues (we’ll handle the failure in the next step).

What Success Actually Looks Like

Before implementing retry strategies, let’s define what we’re aiming for. A “successful” workflow isn’t one that never fails—it’s one where:

  1. Failures are visible: When something goes wrong, someone knows about it
  2. Failures are handled gracefully: Transient errors are retried, permanent errors escalate
  3. Failures don’t cascade: A failure in one part doesn’t break unrelated parts
  4. Failures are informative: When things break, there’s enough context to understand why

This means your workflow design should answer:

  • What happens if Claude Code times out?
  • What happens if the API returns 429?
  • What happens if the user’s API key is invalid?
  • What happens if we can’t authenticate to GitHub?
  • What happens if a file we’re trying to read doesn’t exist?
  • What happens if we need to write to a directory but it’s not writable?

Most teams never ask these questions. They assume success. Then they’re shocked when something fails.

Respecting Rate Limits and the Retry-After Header

The Anthropic API will tell you when to retry using the Retry-After header. Ignoring this header and implementing your own backoff is stubborn and inefficient. Here’s how to respect it:

#!/bin/bash

call_claude_api() {
  local url="$1"
  local data="$2"
  local attempt=1
  local max_attempts=5

  while [ $attempt -le $max_attempts ]; do
    echo "Calling Claude API (attempt $attempt)..."

    response=$(curl -s -w "\n%{http_code}" -X POST "$url" \
      -H "Authorization: Bearer $ANTHROPIC_API_KEY" \
      -H "Content-Type: application/json" \
      -d "$data")

    # Split response and status code
    body=$(echo "$response" | head -n -1)
    status=$(echo "$response" | tail -n 1)

    if [ "$status" = "200" ]; then
      echo "$body"
      return 0
    elif [ "$status" = "429" ]; then
      # Rate limited - extract Retry-After header
      retry_after=$(curl -s -I -X POST "$url" \
        -H "Authorization: Bearer $ANTHROPIC_API_KEY" \
        -H "Content-Type: application/json" \
        -d "$data" | grep -i "retry-after" | cut -d' ' -f2 | tr -d '\r')

      if [ -z "$retry_after" ]; then
        retry_after=5  # Default to 5s if header is missing
      fi

      echo "Rate limited. Waiting ${retry_after}s..." >&2
      sleep "$retry_after"
      attempt=$((attempt + 1))
    elif [ "$status" = "500" ] || [ "$status" = "502" ] || [ "$status" = "503" ]; then
      # Server error - exponential backoff
      wait_time=$((2 ** (attempt - 1)))
      echo "Server error ($status). Waiting ${wait_time}s..." >&2
      sleep "$wait_time"
      attempt=$((attempt + 1))
    else
      # Client error - don't retry
      echo "Client error ($status). Not retrying." >&2
      echo "$body" >&2
      return 1
    fi
  done

  echo "Failed after $max_attempts attempts." >&2
  return 1
}

This function is smarter about retries:

  1. 200: Success, return immediately.
  2. 429: Read the Retry-After header, wait that long, retry.
  3. 5xx errors: Exponential backoff (server errors often resolve themselves).
  4. 4xx errors: Don’t retry—they won’t succeed no matter how many times you try.

The key insight is that the Retry-After header is your source of truth for rate limits. Don’t invent your own backoff strategy; let the API tell you when to try again.

Understanding When to Retry vs. When to Fail Fast

This is the critical decision point that most implementations get wrong. You don’t retry everything. Some errors are worth retrying; others just waste time and resources.

Worth retrying:

  • Network errors (connection refused, timeout, reset)
  • Server errors (5xx status codes)
  • Rate limit errors (429) when Retry-After header is present and reasonable
  • Transient infrastructure issues

Not worth retrying:

  • Authentication errors (401, 403)
  • Malformed requests (400)
  • Resource not found (404)
  • Invalid API keys or tokens

The difference matters because retrying the wrong thing is expensive. If you’re retrying a 401 error 5 times, you’ve wasted 5 API calls that could have been used for something that works. You’ve delayed feedback to the user. And you’ve consumed GitHub Actions runner minutes (which have quotas).

A good rule: retry transient errors (those that happen due to temporary system state), but fail fast on permanent errors (those caused by incorrect configuration or logic).

#!/bin/bash
call_api_smart_retry() {
  local max_attempts=5
  local attempt=1

  while [ $attempt -le $max_attempts ]; do
    response=$(curl -s -w "\n%{http_code}" "$@")
    status=$(echo "$response" | tail -n1)
    body=$(echo "$response" | head -n-1)

    case $status in
      200|201|204)
        # Success
        echo "$body"
        return 0
        ;;
      400|401|403|404)
        # Permanent error—don't retry
        echo "Permanent error ($status): $body" >&2
        return 1
        ;;
      429)
        # Rate limited—respect Retry-After
        retry_after=$(echo "$response" | grep -i "retry-after" | cut -d' ' -f2)
        sleep "${retry_after:-5}"
        attempt=$((attempt + 1))
        ;;
      500|502|503|504)
        # Transient error—exponential backoff
        if [ $attempt -lt $max_attempts ]; then
          wait_time=$((2 ** (attempt - 1)))
          sleep "$wait_time"
          attempt=$((attempt + 1))
        else
          echo "Transient error ($status) after $max_attempts attempts" >&2
          return 1
        fi
        ;;
      *)
        # Unknown error—give up
        echo "Unknown error ($status): $body" >&2
        return 1
        ;;
    esac
  done

  return 1
}

This function categorizes errors and makes the right decision about each one. Notice that 4xx errors (except 429) return immediately—there’s no point retrying. 5xx errors retry with exponential backoff. Rate limits retry respecting the header.

The Testing Nightmare: Actually Testing Error Paths

Here’s something that rarely gets discussed: how do you test your error handling? Most teams never do. They test the happy path (success), maybe test one failure case, but they never test:

  • What happens when retries are exhausted?
  • What happens when both retry and fallback fail?
  • What happens when the fallback takes longer than the timeout?
  • What happens when rate limiting hits?

You can’t reliably test these against a live API. But you can mock:

# test-error-handling.yml
name: Test Error Handling

on: [push]

jobs:
  test-retry-logic:
    runs-on: ubuntu-latest
    steps:
      # Start a mock server that simulates errors
      - name: Start Mock API
        run: |
          # Start a server that returns 429 first, then 200
          python3 << 'EOF'
          from http.server import HTTPServer, BaseHTTPRequestHandler
          import threading
          import time

          request_count = 0

          class MockHandler(BaseHTTPRequestHandler):
            def do_POST(self):
              global request_count
              request_count += 1

              if request_count == 1:
                # First request: rate limited
                self.send_response(429)
                self.send_header('Retry-After', '2')
                self.end_headers()
              else:
                # Second request: success
                self.send_response(200)
                self.send_header('Content-Type', 'application/json')
                self.end_headers()
                self.wfile.write(b'{"success": true}')

            def log_message(self, format, *args):
              pass  # Suppress logs

          server = HTTPServer(('localhost', 8000), MockHandler)
          thread = threading.Thread(target=server.serve_forever)
          thread.daemon = True
          thread.start()
          time.sleep(1)  # Give server time to start
          EOF
        &

      # Wait for mock server to start
      - name: Test Retry Logic
        run: |
          # This script retries after 429
          ./test-retry.sh https://automateanddeploy.com:8000/review

      # Verify the retry succeeded
      - name: Verify Result
        run: |
          if grep -q "success" retry_result.json; then
            echo "✅ Retry logic works"
          else
            echo "❌ Retry logic failed"
            exit 1
          fi

This is more work than testing the happy path, but it’s essential. By mocking error responses, you can verify that your retry logic actually works before deploying.

Monitoring and Alerting: Knowing When Things Break

Error handling is only half the battle. You also need to know when errors occur. Even if your workflow handles failures gracefully, you need visibility into patterns.

- name: Monitor Workflow Health
  if: always()
  run: |
    # Log workflow result for monitoring
    curl -X POST https://monitoring.example.com/workflow-event \
      -H "Authorization: Bearer ${{ secrets.MONITORING_TOKEN }}" \
      -d '{
        "workflow": "${{ github.workflow }}",
        "status": "${{ job.status }}",
        "run_id": "${{ github.run_id }}",
        "attempt": ${{ github.run_attempt }},
        "timestamp": "'$(date -u +%Y-%m-%dT%H:%M:%SZ)'",
        "duration": '$WORKFLOW_DURATION'
      }'

This sends data to your monitoring system about every workflow run. You can then:

  • Set alerts: “If Claude Code reviews fail 3 times in a row, alert the team”
  • Track trends: “Error rates are increasing on weekends—investigate”
  • Correlate failures: “These reviews failed right after a deployment—related?”
  • Understand patterns: “These specific error types fail 80% of the time and should be reclassified”

Over time, monitoring data informs your retry strategies. If you see that 429 errors consistently resolve within 2 seconds, you can tune your backoff. If you see that 500 errors never resolve and always require manual intervention, you can stop retrying and escalate immediately.

Understanding Cost: Retries and API Credits

This deserves its own section because retry costs are often invisible until they become catastrophic.

Each request to Claude Code’s API costs credits. Failed requests that timeout still cost credits. Retries that don’t respect rate limits compound the cost. Consider this scenario:

Initial request at capacity: fails with 429 (rate limited)
Retry without respecting Retry-After: fails, wastes credits
Retry again immediately: fails, wastes credits
Retry again: fails, wastes credits
...after 10 retries, you've burned 10x credits for a single request

Compare to:

Initial request: fails with 429, includes Retry-After: 5
Wait 5 seconds as instructed
Retry: succeeds, costs 1x credits

The difference is 10x. Multiply that by thousands of workflows, and you’re talking about massive cost differences. Smart retry strategies literally save money.

Moreover, rate-limited requests signal that you’re hitting API capacity. Respecting those signals isn’t just polite—it’s economically rational. If you’re rate-limited, retrying faster won’t help and will anger the API providers (and potentially trigger your own rate limits being reduced).

Building Resilience Across Your Infrastructure

Error handling in a single workflow is one thing. Building resilience across an entire CI/CD infrastructure is another. Here’s how to think about it systematically:

Layer 1: Individual Workflow Resilience (what we’ve covered)

  • Timeouts prevent hanging
  • Retries with backoff handle transient errors
  • Graceful fallbacks when errors are permanent
  • Monitoring to detect patterns

Layer 2: Cross-Workflow Communication
Multiple workflows might depend on each other. If one fails, others might be left in an inconsistent state. Use job outcomes to coordinate:

build-and-test:
  runs-on: ubuntu-latest
  outputs:
    success: ${{ job.status }}
  steps:
    - uses: actions/checkout@v4
    - run: npm test

deploy:
  needs: build-and-test
  if: needs.build-and-test.outputs.success == 'success'
  runs-on: ubuntu-latest
  steps:
    - run: echo "Deploying..."

This ensures deployment only happens if tests pass. More importantly, it prevents a failed test from silently being ignored.

Layer 3: Failure Notifications
When something actually fails (not just retried but truly failed), humans need to know:

- name: Notify on Failure
  if: failure()
  uses: actions/github-script@v6
  with:
    script: |
      github.rest.issues.createComment({
        issue_number: context.issue.number,
        owner: context.repo.owner,
        repo: context.repo.repo,
        body: `⚠️ Automated review failed. Manual review requested.\n\nError: ${{ job.status }}\nRun: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}`
      })

This creates a visible signal in the PR that something went wrong, rather than letting a failed workflow disappear into logs.

Real-World Patterns: The Lifecycle of a Robust Workflow

Let’s trace what a real, production-grade workflow looks like end-to-end:

name: Claude Code Review with Resilience

on:
  pull_request:
    types: [opened, synchronize, reopened]

jobs:
  automated-review:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    permissions:
      contents: read
      pull-requests: write

    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      # Get PR info (will be needed for error handling)
      - name: Get PR Info
        id: pr-info
        uses: actions/github-script@v6
        with:
          script: |
            const pr = context.payload.pull_request;
            core.setOutput('number', pr.number);
            core.setOutput('title', pr.title);

      # Run review with explicit error handling
      - name: Run Claude Code Review
        id: review
        uses: anthropics/claude-code-action@v1
        with:
          prompt: "Perform a security and code quality review of this PR"
          api_key: ${{ secrets.ANTHROPIC_API_KEY }}
        timeout-minutes: 10
        continue-on-error: true

      # Categorize the failure
      - name: Categorize Failure
        id: categorize
        if: steps.review.outcome == 'failure'
        run: |
          # Check logs to determine failure type
          if grep -q "429" ${{ job.logs_file }}; then
            echo "type=rate-limited" >> $GITHUB_OUTPUT
          elif grep -q "timeout" ${{ job.logs_file }}; then
            echo "type=timeout" >> $GITHUB_OUTPUT
          elif grep -q "401\|403" ${{ job.logs_file }}; then
            echo "type=auth-error" >> $GITHUB_OUTPUT
          else
            echo "type=unknown" >> $GITHUB_OUTPUT
          fi

      # Handle rate limiting with smart retry
      - name: Retry if Rate Limited
        if: steps.categorize.outputs.type == 'rate-limited'
        run: |
          echo "Rate limited. Waiting 30 seconds before retry..."
          sleep 30
          # Retry logic here (could trigger another workflow)

      # Handle auth errors with alert
      - name: Alert on Auth Error
        if: steps.categorize.outputs.type == 'auth-error'
        uses: actions/github-script@v6
        with:
          script: |
            github.rest.issues.createComment({
              issue_number: ${{ steps.pr-info.outputs.number }},
              owner: context.repo.owner,
              repo: context.repo.repo,
              body: '⚠️ Authentication error. Check API credentials.'
            })

      # Fallback: request manual review
      - name: Fallback to Manual Review
        if: steps.review.outcome == 'failure'
        uses: actions/github-script@v6
        with:
          script: |
            github.rest.issues.createComment({
              issue_number: ${{ steps.pr-info.outputs.number }},
              owner: context.repo.owner,
              repo: context.repo.repo,
              body: '⚠️ Automated review could not complete. Manual review requested.'
            })

      # Post success comment
      - name: Post Review Results
        if: steps.review.outcome == 'success'
        uses: actions/github-script@v6
        with:
          script: |
            github.rest.issues.createComment({
              issue_number: ${{ steps.pr-info.outputs.number }},
              owner: context.repo.owner,
              repo: context.repo.repo,
              body: '✅ Automated review completed.'
            })

This workflow:

  • Handles different error types differently
  • Never silently fails (always comments on the PR)
  • Retries when appropriate (rate limiting)
  • Alerts for critical errors (auth)
  • Falls back to manual review
  • Has clear success signals

The Hidden Layer: Why Systems Fail and How to Predict It

Error handling in infrastructure is ultimately about understanding failure modes. Every system has ways it can fail, and most teams don’t think clearly about this until after they’ve experienced the failure in production.

A mature approach to error handling requires you to think probabilistically about failures and categorize them systematically. Not “what can go wrong”—too vague. Instead: “What are the specific ways this integration can fail, and what should happen when it does?”

For a GitHub Actions workflow calling Claude Code, the failure modes include:

Network layer: The request never reaches the API. Connection timeout, DNS failure, network unreachable. These happen regularly, often microscopically. A request that usually takes 50ms occasionally takes 3 seconds. Sometimes that timeout boundary is hit. Your retry strategy should handle this gracefully.

API layer: The request reaches the API but the response indicates an error. 429 (rate limited), 500 (internal error), 503 (overloaded). These are “expected failures”—the API tells you explicitly something went wrong. Your code should read the response and categorize it correctly.

Silent failure: The request appears to succeed (HTTP 200) but the response is corrupted or incomplete. This is insidious because your monitoring might not detect it. You got a response, so you assume success. But the response doesn’t have the fields you expect. Or it has them but they’re malformed. Or the JSON is truncated. These require defensive parsing and validation.

Cascade failures: Failure in one step cascades to others. The code review fails silently (no comment on the PR). The PR merges anyway. The tests fail. You don’t know the review failed because there was no visible signal. This is why every failure path needs explicit communication—a comment, a log, an alert—something that makes the failure visible.

Compound failures: Multiple independent systems fail simultaneously. GitHub is down, Claude API is rate-limited, and your monitoring service is also having issues. Your workflow can’t use any of them. What’s the graceful degradation? You have to plan for this before it happens.

The probability of each failure type informs your strategy. If you’re seeing 429 rate limits once a week, that’s worth handling with intelligent retries. If you’re seeing them 10 times a day, you have a scaling problem that retries won’t solve. If you’ve never seen a network failure in production, but your error logs show your firewall drops 0.01% of outbound connections, you should still handle it—because at scale, 0.01% happens daily.

This thinking—probabilistic risk assessment—is what separates resilient systems from systems that happen to work until they don’t.

The Organizational Impact of Reliable Automation

Here’s something that’s often overlooked in discussions about error handling: when your automation is reliable, it changes how your organization operates. Developers stop being skeptical of automated workflows. They don’t add manual checkpoints “just in case.” They trust the system because the system has proven itself trustworthy through repeated handling of failures gracefully.

This trust is valuable. Teams that build reliable CI/CD pipelines ship faster. They’re confident in automated deployments because they’ve seen the system handle transient failures and recover. They don’t spend time debugging mysterious workflow failures because the workflows are transparent about what happened. The logs tell the story. The monitoring shows the trends. When something does fail, it’s recoverable because they’ve designed for it.

But this only happens if your error handling is systematic and well-designed. Random timeouts that sometimes work, retries that hammer the API, silent failures that nobody notices—these erode trust. Teams stop using automation. They do things manually to be safe. You’ve lost the efficiency gains.

Conversely, when error handling is solid, teams lean on automation. It becomes a multiplier for your engineering effectiveness. One person can orchestrate thousands of tests, thousands of reviews, thousands of deployments through well-designed automation. That’s the power of systems that fail gracefully.

Testing Your Error Handling Under Load

Here’s a practice that separates teams with truly resilient systems from teams that think their systems are resilient: chaos testing. You intentionally break things and verify your error handling works.

This doesn’t require much. Set up a staging environment. Point your workflows at a mock API that randomly returns errors. Run your entire workflow suite. Watch what breaks. Fix it. Repeat. You’re learning what your actual system does when things fail, rather than guessing based on code review.

The cheapest way to do this: write a simple HTTP server that returns errors based on environment variables. Run it locally. Point your CI/CD at it. See what your workflows do. Many teams discover their error handling doesn’t work as well as they thought. Maybe retries happen too fast. Maybe the error categorization misses a case. Maybe alerts aren’t firing. You find out in a safe environment before production surprises you.

This kind of testing feels tedious. But every minute you spend testing error handling prevents hours of debugging when the errors happen in production. And they will happen. Murphy’s Law isn’t a suggestion.

Wrapping Up: Building Resilience Across Your Infrastructure

Resilient systems don’t try to prevent all errors—they can’t. Instead, they prepare for errors when they come. You’ve learned how to categorize errors, implement smart retry logic that respects rate limits, build graceful fallbacks, and monitor for failures.

The key insight: error handling isn’t a feature to add at the end. It’s foundational. Every workflow that calls Claude Code or any external API needs:

  1. Timeouts to prevent hanging forever
  2. Proper categorization to distinguish transient from permanent errors
  3. Smart retries that respect API signaling (Retry-After header)
  4. Exponential backoff to avoid hammering a struggling API
  5. Graceful fallbacks so failures don’t cascade
  6. Visibility so you know when things break
  7. Monitoring so you can improve over time

Your workflow will survive—not because it never fails, but because when it does fail, it fails gracefully and bounces back.

The systems you build with these principles become antifragile. They don’t just survive failures; they improve from them. Each failure teaches you something about your system. Each error recovery demonstrates that your design is sound. Over time, your error handling becomes so solid that you stop worrying about whether your workflows will fail. You trust that they will, and you trust that they’ll recover. That’s the mark of truly professional infrastructure.

Start implementing these patterns today. Begin with timeouts and basic error categorization. Add smart retries. Layer in monitoring. Then expand gradually as you learn what your system actually needs. The investment you make in error handling now will pay dividends for years as your automation becomes the trusted backbone of your development process.


-iNet

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.