All Articles Claude Code

Claude Code for Kubernetes Manifest Review

Kubernetes manifests are deceptively easy to write and deceptively easy to get wrong. You can have a deployment that technically runs but starves for CPU.

Kubernetes manifests are deceptively easy to write and deceptively easy to get wrong. You can have a deployment that technically runs but starves for CPU. Or a pod that’s secure on paper but exposes privileges you didn’t realize. Or a liveness probe that restarts your container every 30 seconds because the threshold is just wrong. Or a scaling policy that works perfectly for your current load but fails catastrophically when traffic spikes. Or resource requests that look reasonable until you multiply them across your production cluster and suddenly you’ve overbought hardware by a factor of three.

Standard YAML linters catch syntax. They don’t catch logic. They don’t catch the subtle security misconfiguration or the resource request that’s way too high for your actual application behavior. They can’t tell you whether your probe thresholds make sense for the latency profile of your application. That’s where Claude Code comes in. We’re building a manifest review system that understands Kubernetes best practices at a deep level and can spot problems that would take a human expert hours to catch—or worse, problems that won’t show up until you’re in production dealing with a full-page outage at 2 AM.

This isn’t about validating YAML structure. This is about understanding your workload’s actual requirements and whether your manifest actually meets them. This is about having a knowledgeable peer review your deployments before they hit production, asking intelligent questions and catching edge cases that static linters completely miss. This is about building confidence that your Kubernetes infrastructure won’t let you down.

The Hidden Costs of Manifest Mistakes: Why This Matters

Kubernetes is uniquely unforgiving when it comes to deployment mistakes. Unlike application code where bugs might affect a few users temporarily, Kubernetes mistakes can take down entire services. A single misconfigured manifest can cause cascading failures that disrupt every service depending on it. When you’re running business-critical workloads on Kubernetes, the stakes are genuinely high.

Real production incidents caused by manifest problems are shockingly common. A company deploys an update that doesn’t include resource limits. During a traffic spike, a single pod consumes all available memory, the node becomes unstable, and Kubernetes evicts other pods trying to recover. Those pods were running critical services—payment processing, authentication, fraud detection—now they’re down. The company loses millions in customer impact within minutes. Trust is damaged. Incident response costs are enormous. And it all traces back to a missing line in a manifest that nobody thought to check.

Another company sets a liveness probe threshold too aggressively. Pods restart constantly because the probe fails during normal operation. The service becomes flaky. Customers get intermittent errors. They lose confidence in the product. The team spends days debugging something that was simply a manifest misconfiguration. Meanwhile, this could have been caught in code review if anyone had known what to look for.

A third company uses the latest image tag without realizing that every time the tag gets rebuilt, the deployment changes. They deploy version 1.0. Somewhere in CI/CD, the image gets rebuilt and retagged. Without anyone noticing, all their pods get replaced with completely different code. The behavior changes. Performance suffers. Debugging becomes a nightmare because they don’t realize the code changed at all.

These aren’t hypothetical scenarios. These happen regularly at scale because manifests look simple—they’re just YAML, how hard can they be?—but they encode complex requirements about how your application should behave in production. Getting those requirements right is genuinely difficult. You need to understand resource utilization profiles. You need to understand startup latency. You need to understand failure modes. Most developers know their application code well, but manifest requirements? That’s a blind spot for many teams.

This is where Claude Code creates value. It can reason about the entire system: the manifest, the application code, the infrastructure. It can spot when resource limits are too low for the actual application. It can catch when probe configurations don’t match the application’s startup characteristics. It can understand the interactions between different parts of the system and flag potential failure modes that only experts would catch. The investment in manifest review pays dividends through reduced incidents, faster troubleshooting when problems do occur, and increased confidence in deployments.

More importantly, Claude can explain its reasoning. It’s not just flagging problems; it’s teaching you Kubernetes best practices along the way. You learn why certain configurations matter, what can go wrong if you ignore them, and how to think about Kubernetes deployments strategically rather than just copying examples.

Resource Management Philosophy: Thinking Beyond Numbers

Before diving into implementation, let’s understand the deeper principles behind resource management. Numbers in Kubernetes manifests aren’t arbitrary. They’re statements about what your application needs to run safely and efficiently.

When you say a container needs 256Mi of memory, you’re making a promise: “This application can do its job with 256Mi.” When you say it limits itself to 512Mi, you’re making a guarantee: “This application will never consume more than 512Mi under any circumstances we can imagine.”

These commitments have direct consequences. Resource requests determine where the scheduler places your pod. If you request 256Mi and the node only has 200Mi available, the pod won’t schedule at all. If you request too little, the scheduler might co-locate your pod with other resource-hungry pods, causing competition and performance degradation.

Resource limits determine what happens under stress. If you set no limit, a runaway application can consume all resources on the node, causing the Kubernetes system itself to become unstable. If you set limits too low, your application gets throttled, leading to mysterious timeouts and performance issues.

Understanding your application’s actual resource profile requires measurement, not guessing. The teams with the best deployments don’t just copy example manifests. They actually run their applications, measure how much CPU and memory they use under realistic load, add safety margin, and configure limits based on reality.

Claude Code helps with this investigation. It can look at an application and reason about typical resource requirements for that type of application. A web service that handles HTTP requests might need 128-256Mi per instance. A database might need gigabytes. A background worker might need very little. By understanding the application, Claude can flag when resource requests seem misaligned with what the app likely needs.

Probe Configuration Philosophy: Detecting When Things Are Wrong

Health probes are where many manifests fail, because the thresholds are often guesses rather than measurements. A readiness probe that fails after 5 seconds of no response assumes your application becomes ready within 5 seconds. But what if your application takes 30 seconds to start? Traffic gets routed to pods that aren’t ready.

A liveness probe that restarts a container after 30 seconds of no response will create cascading failures if your application occasionally takes 31 seconds to respond. The container restarts, the pod is temporarily unavailable, traffic gets routed elsewhere, and users experience flakiness.

Getting probes right requires understanding your application’s actual behavior. How long does startup take? In the worst case? What’s the typical latency under normal operation? During peak load? What endpoints are safe to hit that indicate readiness without side effects?

Claude Code can help with this reasoning. By understanding the application, it can suggest reasonable probe configurations and flag when they seem misaligned with how long the app actually takes to start or respond.

The Problem with Most K8s Reviews: Standard Approaches Fall Short

Before we build the solution, let’s be clear about what we’re solving. Typical approaches all have significant gaps that leave real problems undetected.

Generic linters like kubeval and kube-score excel at flagging obvious issues like missing resource limits or improper field formats. They’ll catch syntax errors and obviously wrong configurations. But they miss context-specific problems. A 256Mi memory limit is flagged as “low” by some linters, but if your application is a lightweight sidecar, 256Mi is actually appropriate. If it’s a database, it’s catastrophically insufficient. The linter can’t know which you’re building.

Manual review catches everything—theoretically. In practice, it takes hours per manifest and depends entirely on reviewer expertise and attention. Reviews require someone who understands Kubernetes deeply, understands your specific application requirements, understands your infrastructure constraints, and is attentive enough to catch subtle problems when reviewing their tenth manifest of the day. That person exists, but they’re expensive. And human reviewers get tired, miss details, sometimes rubber-stamp PRs because they trust the developer or because they’ve been reviewing for too long.

Policy-as-code tools like Kyverno and OPA are great for enforcement. They can say “all deployments must have resource limits” or “all containers must run as non-root.” But they require pre-defined policies, and they can’t reason about workload-specific tradeoffs. Is it acceptable to violate a policy in this specific case? Policy-as-code says no. But a human reviewer might say “yes, this is acceptable because…”

What we need is something that understands your specific workload and asks intelligent questions. Is this database really going to fit in 256Mi of memory, or will it crash under real load? Does this service actually need 4 CPU cores, or is that overkill based on actual traffic patterns? Why is there no readiness probe? How will the scheduler handle these resource requests across your cluster? What happens to persistent data during pod restarts? Can a single node failure cascade into complete service failure, or is your replica strategy sufficient?

Claude Code can reason about these things because it understands both the manifest syntax and the code it’s running. It can correlate the deployment specification with application requirements. It can spot patterns that indicate problems. When it finds issues, it doesn’t just flag them; it explains them and suggests fixes. It acts like a knowledgeable colleague reviewing your work, asking clarifying questions and helping you build better systems.

Why Manifest Review Matters: Understanding Your Blast Radius

The blast radius of a bad Kubernetes deployment is enormous. Let’s be very specific about what can go wrong and why it matters.

A manifest without resource limits means one pod can consume all CPU on a node, evicting other critical workloads. Your payment processing service gets evicted by a runaway analytics pod. Customers trying to pay get connection timeouts. Your business grinds to a halt. This happens in minutes without your knowledge.

A pod without a readiness probe means traffic gets sent to pods that aren’t ready to handle requests. If your application takes 30 seconds to start and you don’t have a readiness probe, you’re sending traffic to pods during their startup window. Those requests fail. Cascading failures through your system follow. Clients see flaky, unreliable service. Some clients retry immediately, which creates thundering herd problems.

Missing security contexts mean a compromised container gains access to the host filesystem. If your container gets exploited, the attacker can see and modify other containers’ data. They can steal secrets. They can install rootkits. The blast radius expands far beyond your single pod.

Latest image tags mean deployments change unpredictably. You can’t reproduce issues reliably. You can’t know with confidence what code is running in production. Rolling back becomes problematic because you don’t know exactly what you’re rolling back to. Version management becomes impossible.

Unbounded memory limits mean a single pod causes node overload and cascading failures. Memory pressure causes the OS to swap. Swapping causes massive latency. Your entire node becomes sluggish, affecting all pods on that node. As those pods start failing, their load redistributes to other nodes, causing a cascading domino effect of failure propagating through your cluster.

Each of these is preventable with a good manifest review. The earlier you catch them, the cheaper the fix. A developer can update a manifest at code review time—that’s minutes of work. In production, you’re dealing with incident response, customer impact, wasted infrastructure costs, and trust damage. We’re talking about preventing millions of dollars in losses through disciplined review.

Building the Manifest Analyzer: A Comprehensive System

The manifest review system works by analyzing YAML, extracting key characteristics, checking against best practices, and generating detailed recommendations. Let’s build it piece by piece, understanding not just the code but the thinking behind each component.

Phase 1: Parse and Extract Manifest Metadata

Start by reading the manifest and extracting key information. This extraction phase is crucial because it gives us structured data to work with. We’re not trying to be exhaustive—we’re extracting the information that matters for security and reliability analysis.

When you parse a manifest, you’re looking for several key categories. First, the basic identity: what kind of Kubernetes object is this, what’s it named, where does it live in the cluster. Second, the container configuration: what images are we running, what resources are they asking for and limiting themselves to. Third, the health and readiness signals: do we have probes configured, and are they reasonable. Fourth, security posture: are we running as root, do we have capabilities restricted. Fifth, storage: are we using persistent volumes, are they configured correctly. Sixth, networking and environment: what ports are exposed, what environment variables are set.

interface ManifestMetadata {
  kind: string; // Deployment, StatefulSet, DaemonSet, etc.
  name: string;
  namespace: string;
  containerImages: string[];
  replicas?: number;
  resources: {
    requests?: { cpu: string; memory: string };
    limits?: { cpu: string; memory: string };
  };
  probes: {
    liveness?: { initialDelaySeconds: number; periodSeconds: number };
    readiness?: { initialDelaySeconds: number; periodSeconds: number };
  };
  securityContext?: any;
  volumes?: string[];
  env?: string[];
}

async function parseManifest(yamlContent: string): Promise<ManifestMetadata> {
  const yaml = require("js-yaml");
  const manifest = yaml.load(yamlContent);

  const metadata: ManifestMetadata = {
    kind: manifest.kind,
    name: manifest.metadata.name,
    namespace: manifest.metadata.namespace || "default",
    containerImages: manifest.spec.template.spec.containers.map(
      (c: any) => c.image,
    ),
    replicas: manifest.spec.replicas || 1,
    resources: {
      requests: manifest.spec.template.spec.containers[0]?.resources?.requests,
      limits: manifest.spec.template.spec.containers[0]?.resources?.limits,
    },
    probes: extractProbes(manifest),
    securityContext: manifest.spec.template.spec.securityContext,
    volumes: manifest.spec.template.spec.volumes?.map((v: any) => v.name),
    env: manifest.spec.template.spec.containers[0]?.env?.map(
      (e: any) => e.name,
    ),
  };

  return metadata;
}

function extractProbes(manifest: any) {
  const container = manifest.spec.template.spec.containers[0];
  return {
    liveness: container?.livenessProbe,
    readiness: container?.readinessProbe,
  };
}

This extraction gives us all the key information we need to reason about the manifest. We’re capturing resource requests, probe configuration, security settings, storage, and environment variables. Everything else is detail—important detail, but not the core signal.

Phase 2: Run Checks Against Best Practices

Now we check the manifest against known best practices and gather issues. These checks represent distilled knowledge from thousands of Kubernetes deployments. Each check answers a specific question: Is this aspect of the manifest configured correctly for production?

The checks work at different abstraction levels. Some are hard rules: you must have resource limits. Some are soft recommendations: you should have a readiness probe. Some are conditional: if you’re stateful, you need PersistentVolumes. The severity level reflects the actual risk of ignoring the check.

interface ManifestIssue {
  severity: "critical" | "high" | "medium" | "low";
  category: string;
  title: string;
  description: string;
  suggestion: string;
}

async function checkManifestBestPractices(
  metadata: ManifestMetadata,
): Promise<ManifestIssue[]> {
  const issues: ManifestIssue[] = [];

  // Check 1: Resource limits - can't run safely without them
  if (!metadata.resources.limits) {
    issues.push({
      severity: "high",
      category: "Resources",
      title: "No resource limits specified",
      description:
        "Without resource limits, this pod can consume unlimited CPU and memory, potentially starving other workloads on the node. A single pod can cause cascading failures across your entire cluster.",
      suggestion:
        "Add resources.limits.cpu and resources.limits.memory to constrain resource consumption. Start with conservative estimates and tune based on actual monitoring.",
    });
  }

  // Check 2: Resource requests - scheduler can't work well without them
  if (!metadata.resources.requests) {
    issues.push({
      severity: "medium",
      category: "Resources",
      title: "No resource requests specified",
      description:
        "The scheduler won't know how much resources this pod needs, leading to inefficient placement. Pods might be scheduled on nodes that don't have enough actual resources available.",
      suggestion: "Add resources.requests.cpu and resources.requests.memory. These should reflect realistic requirements based on application profiling.",
    });
  }

  // Check 3: Readiness probe - traffic routing depends on this
  if (!metadata.probes.readiness) {
    issues.push({
      severity: "high",
      category: "Health",
      title: "No readiness probe configured",
      description:
        "Without a readiness probe, traffic may be sent to pods that aren't ready to handle requests. If your app takes time to initialize, traffic during startup will fail, cascading through your system.",
      suggestion:
        "Add a readinessProbe to detect when the pod is ready for traffic. Configure thresholds based on your actual startup time and normal operation patterns.",
    });
  }

  // Check 4: Liveness probe - dead pods should be replaced
  if (!metadata.probes.liveness) {
    issues.push({
      severity: "medium",
      category: "Health",
      title: "No liveness probe configured",
      description:
        "If the application hangs or deadlocks, Kubernetes won't automatically restart it. You'll have zombie pods consuming resources but not serving traffic.",
      suggestion: "Add a livenessProbe to detect and recover from application failures. Be conservative with thresholds to avoid false-positive restarts.",
    });
  }

  // Check 5: Security context - root = massive blast radius
  if (!metadata.securityContext?.runAsNonRoot) {
    issues.push({
      severity: "high",
      category: "Security",
      title: "Container may run as root",
      description:
        "Running as root increases the blast radius of security vulnerabilities. If your container is compromised, attackers gain root access to the host. This is a critical security risk.",
      suggestion:
        "Set securityContext.runAsNonRoot: true and specify a non-root user. Use a dedicated user for your application, not root.",
    });
  }

  // Check 6: Image tag - reproducibility depends on this
  const hasLatestTag = metadata.containerImages.some((img) => img.endsWith(":latest"));
  if (hasLatestTag) {
    issues.push({
      severity: "medium",
      category: "Deployment",
      title: "Using :latest image tag",
      description:
        "The :latest tag means deployments can change unpredictably when the image is rebuilt. You can't reproduce issues. You can't know what code is running in production. Rolling back becomes problematic.",
      suggestion:
        "Use explicit version tags like :1.2.3 or :2026-03-17 to ensure reproducible deployments. This also makes tracking what's in production straightforward.",
    });
  }

  return issues;
}

These checks cover the most common problems we see in production deployments. Each issue has severity, a clear description of why it matters, and actionable suggestions. Notice that the descriptions explain the actual business impact, not just the technical problem. This helps developers understand why they should care, not just that they should fix it.

Phase 3: Ask Claude for Domain-Specific Analysis

For issues that require understanding the actual application, send the manifest to Claude. Claude can look at the specific numbers in your manifest and reason about whether they make sense for the type of application you’re running.

async function runClaudeAnalysis(
  yamlContent: string,
  metadata: ManifestMetadata,
): Promise<string> {
  const { spawn } = require("child_process");

  return new Promise((resolve, reject) => {
    const process = spawn("claude", ["--json"], {
      stdio: ["pipe", "pipe", "pipe"],
    });

    let output = "";

    process.stdout.on("data", (data: Buffer) => {
      output += data.toString();
    });

    process.on("close", (code: number) => {
      if (code !== 0) {
        reject(new Error("Claude Code failed"));
      } else {
        try {
          const result = JSON.parse(output);
          resolve(result.response);
        } catch (e) {
          reject(new Error("Failed to parse Claude response"));
        }
      }
    });

    const prompt = `You are a Kubernetes expert reviewing a deployment manifest. Analyze this manifest for production readiness and potential issues that static linters would miss.

Manifest:
${yamlContent}

Workload Details:
- Kind: ${metadata.kind}
- Name: ${metadata.name}
- Replicas: ${metadata.replicas}
- Container images: ${metadata.containerImages.join(", ")}
- Resource requests: ${JSON.stringify(metadata.resources.requests)}
- Resource limits: ${JSON.stringify(metadata.resources.limits)}
- Has readiness probe: ${!!metadata.probes.readiness}
- Has liveness probe: ${!!metadata.probes.liveness}

Please analyze:
1. Are resource requests/limits appropriate for this workload type? Consider typical resource utilization patterns.
2. Is the probe configuration adequate? Will probes detect real problems?
3. Are there any security concerns beyond the standard checks?
4. Will this manifest work well in a production cluster under realistic load?
5. Are there any scaling or resilience issues?

Provide specific, actionable recommendations with reasoning about why each matters.`;

    process.stdin.write(JSON.stringify({ prompt, context: yamlContent }));
    process.stdin.end();
  });
}

Claude can reason about things like: “This is a database deployment with 256Mi memory limit, but databases typically require at least 512Mi to fit their working set in memory without thrashing.” Or “This service has 2 replicas but no pod disruption budgets, so a single node failure takes down 50% of capacity.” Or “The startup time based on the probe configuration seems optimistic for a JVM application.” Static linters can’t make those inferences.

Phase 4: Generate Comprehensive Report

Combine all findings into a comprehensive report that gives developers actionable next steps.

interface ManifestReview {
  manifest: string;
  metadata: ManifestMetadata;
  staticIssues: ManifestIssue[];
  claudeAnalysis: string;
  summary: {
    totalIssues: number;
    criticalIssues: number;
    riskScore: number; // 0-100
  };
}

async function reviewManifest(yamlContent: string): Promise<ManifestReview> {
  console.log("📋 Parsing manifest...");
  const metadata = await parseManifest(yamlContent);

  console.log("🔍 Checking best practices...");
  const staticIssues = await checkManifestBestPractices(metadata);

  console.log("🤖 Running Claude analysis...");
  const claudeAnalysis = await runClaudeAnalysis(yamlContent, metadata);

  const criticalIssues = staticIssues.filter(
    (i) => i.severity === "critical",
  ).length;

  return {
    manifest: yamlContent,
    metadata,
    staticIssues,
    claudeAnalysis,
    summary: {
      totalIssues: staticIssues.length,
      criticalIssues,
      riskScore: calculateRiskScore(staticIssues),
    },
  };
}

function calculateRiskScore(issues: ManifestIssue[]): number {
  let score = 0;
  for (const issue of issues) {
    switch (issue.severity) {
      case "critical":
        score += 25;
        break;
      case "high":
        score += 15;
        break;
      case "medium":
        score += 8;
        break;
      case "low":
        score += 2;
        break;
    }
  }
  return Math.min(100, score);
}

Phase 5: Format and Output the Report

Generate a detailed, readable report that developers can actually use:

function formatReview(review: ManifestReview): string {
  let report = "";

  report += `# Kubernetes Manifest Review\n`;
  report += `**Manifest**: ${review.metadata.name} (${review.metadata.kind})\n`;
  report += `**Namespace**: ${review.metadata.namespace}\n`;
  report += `**Risk Score**: ${review.summary.riskScore}/100\n\n`;

  if (review.summary.criticalIssues > 0) {
    report += `## ⚠️  Critical Issues (${review.summary.criticalIssues})\n`;
    report += `These issues should be fixed before deploying to production.\n\n`;
    const critical = review.staticIssues.filter((i) => i.severity === "critical");
    for (const issue of critical) {
      report += `### ${issue.title}\n`;
      report += `${issue.description}\n`;
      report += `**Fix**: ${issue.suggestion}\n\n`;
    }
  }

  if (review.staticIssues.length > review.summary.criticalIssues) {
    report += `## Other Issues\n`;
    const others = review.staticIssues.filter((i) => i.severity !== "critical");
    for (const issue of others) {
      report += `- **[${issue.severity}]** ${issue.title}: ${issue.description}\n`;
    }
    report += "\n";
  }

  report += `## Claude Analysis\n`;
  report += review.claudeAnalysis;

  return report;
}

Real-World Example: Catching Common Mistakes

Let’s say you have a simple web server deployment that looks fine at first glance:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web-server
  namespace: default
spec:
  replicas: 2
  selector:
    matchLabels:
      app: web-server
  template:
    metadata:
      labels:
        app: web-server
    spec:
      containers:
        - name: web
          image: myregistry.azurecr.io/web-server:latest
          ports:
            - containerPort: 8080
          env:
            - name: PORT
              value: "8080"

The review would find:

  • Critical: No resource limits. A traffic spike could cause this pod to consume all CPU on its node, evicting other workloads.
  • Critical: Running as root. If the web server is compromised, attackers get root access to the host.
  • High: No readiness probe. Traffic gets sent to pods during startup, causing cascading failures.
  • High: No liveness probe. Hung processes stay running and consume resources forever.
  • Medium: Using :latest tag. Deployments change unpredictably. You can’t reproduce issues. Rolling back is problematic.
  • Claude finds: With only 2 replicas and no pod disruption budgets, a single node failure takes down your entire web service. For a production web service, you need at least 3 replicas with pod disruption budgets.

The report would provide specific fixes for each issue, explaining why they matter and how to implement them.

What a Solid Manifest Looks Like

For contrast, here’s what a production-ready manifest looks like. Every field has a reason:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web-server
  namespace: production
  labels:
    app: web-server
    version: "1.2.3"
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  selector:
    matchLabels:
      app: web-server
  template:
    metadata:
      labels:
        app: web-server
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "9090"
    spec:
      serviceAccountName: web-server
      securityContext:
        runAsNonRoot: true
        runAsUser: 1000
        fsGroup: 1000
      containers:
        - name: web
          image: myregistry.azurecr.io/web-server:1.2.3
          imagePullPolicy: IfNotPresent
          ports:
            - name: http
              containerPort: 8080
              protocol: TCP
            - name: metrics
              containerPort: 9090
              protocol: TCP
          env:
            - name: PORT
              value: "8080"
            - name: LOG_LEVEL
              value: "info"
          resources:
            requests:
              cpu: 100m
              memory: 256Mi
            limits:
              cpu: 500m
              memory: 512Mi
          readinessProbe:
            httpGet:
              path: /health/ready
              port: http
            initialDelaySeconds: 5
            periodSeconds: 10
            timeoutSeconds: 3
            failureThreshold: 3
          livenessProbe:
            httpGet:
              path: /health/live
              port: http
            initialDelaySeconds: 15
            periodSeconds: 20
            timeoutSeconds: 3
            failureThreshold: 2
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities:
              drop:
                - ALL
          volumeMounts:
            - name: tmp
              mountPath: /tmp
            - name: cache
              mountPath: /app/cache
      volumes:
        - name: tmp
          emptyDir: {}
        - name: cache
          emptyDir:
            sizeLimit: 100Mi
      affinity:
        podAntiAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              podAffinityTerm:
                labelSelector:
                  matchExpressions:
                    - key: app
                      operator: In
                      values:
                        - web-server
                topologyKey: kubernetes.io/hostname

This manifest demonstrates:
– Explicit version tags (no :latest, no risk of unpredictable changes)
– Resource requests AND limits (scheduler can make intelligent decisions, runaway pods are contained)
– Both readiness and liveness probes (traffic only goes to healthy pods, dead pods are replaced)
– Security context running as non-root (blast radius is minimized)
– Read-only root filesystem (file-based attacks are harder)
– Capability dropping (exploits have fewer tools available)
– Pod anti-affinity for distribution (single node failure doesn’t take down the service)
– Proper labels and annotations (monitoring, debugging, automation can all work)
– Meaningful replica count for HA (redundancy is built in)

Every change here has a reason. Every field prevents a real production problem that has happened to someone.

Integration Patterns: Making This Part of Your Workflow

Pattern 1: Local Review Before Committing

claude code review-k8s-manifest deployment.yaml

Get feedback locally before committing to git. Catch issues before they even hit a PR. This is the fastest feedback loop.

Pattern 2: CI/CD Gate

# GitHub Actions example
- name: Review Kubernetes Manifest
  run: |
    claude code review-k8s-manifest k8s/deployment.yaml > review.txt
    if grep -q "Critical:" review.txt; then
      echo "Critical issues found in manifest"
      cat review.txt
      exit 1
    fi

Block PRs with critical manifest issues. No one can merge a manifest that the review system flagged. This is enforcement at the point where it’s cheapest to fix.

Pattern 3: Continuous Validation

# Scan all manifests in cluster
find k8s/ -name "*.yaml" -type f | while read manifest; do
  claude code review-k8s-manifest "$manifest"
done

Audit all deployments for compliance. Run this weekly or daily depending on how frequently you deploy. Over time, you build historical records of how your manifests evolve.

Why This Matters: Building Systems You Can Trust

Kubernetes deployments are complex. A manifest that looks correct might have subtle issues that only appear under load. Without proper resource requests, pods get randomly evicted when the cluster is under pressure. Without readiness probes, traffic goes to unhealthy pods. Without security contexts, a compromised container can access the host. Without proper replica counts and anti-affinity, a single node failure cascades into complete service failure.

The teams with the most reliable Kubernetes deployments aren’t the ones with the most complex manifests. They’re the ones with manifests that have been thoroughly reviewed and follow best practices consistently. They’ve invested in understanding not just how to write Kubernetes manifests, but how to write them well.

Claude Code can automate that review process. It can catch issues that would take hours to find manually. It can explain not just what’s wrong but why it matters and how to fix it. It can embed best practices into your development process so every manifest that ships is better than the last.

Specific Configuration Patterns: Learning From Success and Failure

Over time, you discover patterns in Kubernetes that work well and patterns that cause problems. Learning these patterns helps you recognize good and bad configurations instantly.

The Stateless Service Pattern (Works Well)

Stateless services that can scale horizontally are Kubernetes’ sweet spot. Multiple replicas, no persistent state, configuration via environment variables, health probes for routing. These deployments are simple, reliable, and scale beautifully.

The Database Pattern (Hard to Get Right)

Database deployments are fundamentally different from stateless services. They need persistent storage, careful coordination of replicas, specific resource guarantees, and careful handling of upgrades. StatefulSets exist specifically for this. Yet many teams try to run databases as Deployments with PersistentVolumes, which works in simple cases but fails catastrophically in real-world failure scenarios.

The Batch Job Pattern (Needs Care)

Batch jobs—one-off tasks that run for a while and finish—need different configurations than long-running services. They don’t need readiness probes because they’re not routing traffic. They don’t need multiple replicas because they’re not redundant. They do need backoff policies and retry logic for when they fail. Using the wrong pattern leads to jobs that hang or repeatedly fail.

The Sidecar Pattern (Powerful But Dangerous)

Sidecar containers—additional containers in the same pod—enable powerful patterns but create complications. Multiple containers competing for resources. Multiple containers with different startup times. Coordinating health checks across containers becomes complex. Powerful when used intentionally, problematic when added without consideration.

Claude Code can understand these patterns and recognize when a manifest violates the pattern it appears to follow. A database running as a Deployment is a red flag. A stateless service with persistent volumes is suspicious. A batch job with 5 replicas suggests a misunderstanding of what you’re building.

Advanced Manifest Analysis: Multi-Object Reviews

Real Kubernetes deployments often involve multiple manifests working together—a Deployment with associated Services, ConfigMaps, PersistentVolumeClaims, and Ingress rules. Claude Code can analyze the entire ensemble and spot consistency issues that single-manifest reviews would miss.

async function reviewManifestEnsemble(manifestFiles: string[]): Promise<any> {
  const manifests = {};

  // Load all manifests
  for (const file of manifestFiles) {
    const content = fs.readFileSync(file, "utf8");
    manifests[file] = yaml.load(content);
  }

  // Cross-reference checks
  const issues = [];

  // Check 1: Service selectors match deployment labels
  const deployments = Object.entries(manifests)
    .filter(([_, m]) => m.kind === "Deployment")
    .map(([_, m]) => m);

  const services = Object.entries(manifests)
    .filter(([_, m]) => m.kind === "Service")
    .map(([_, m]) => m);

  for (const service of services) {
    const matchingDeployments = deployments.filter(
      (d) =>
        JSON.stringify(d.spec.selector.matchLabels) ===
        JSON.stringify(service.spec.selector),
    );

    if (matchingDeployments.length === 0) {
      issues.push({
        severity: "high",
        message: `Service ${service.metadata.name} has no matching deployment`,
        suggestion: "Ensure service selectors match deployment labels",
      });
    }
  }

  // Check 2: Persistent volumes are mounted and available
  const pvcs = Object.values(manifests).filter((m) => m.kind === "PersistentVolumeClaim");

  for (const deployment of deployments) {
    for (const volume of deployment.spec.template.spec.volumes || []) {
      if (volume.persistentVolumeClaim) {
        const matchingPvc = pvcs.find(
          (p) =>
            p.metadata.name ===
            volume.persistentVolumeClaim.claimName,
        );

        if (!matchingPvc) {
          issues.push({
            severity: "critical",
            message: `Deployment references non-existent PVC: ${volume.persistentVolumeClaim.claimName}`,
            suggestion: "Create the referenced PVC or fix the reference",
          });
        }
      }
    }
  }

  return issues;
}

This multi-manifest analysis catches coordination problems that wouldn’t be visible when reviewing manifests in isolation. A Service that doesn’t match any Deployment is a deployment error that static analysis will never catch because it’s looking at incomplete information.

Continuous Manifest Auditing: Staying Healthy Over Time

Kubernetes isn’t static. Over time, best practices change, your team learns new patterns, and configurations drift. Build continuous auditing to catch these changes:

#!/bin/bash
# Script to continuously audit deployed manifests

REVIEW_INTERVAL=604800  # Weekly

while true; do
  echo "Starting manifest audit..."

  # Get all deployed manifests
  kubectl get all --all-namespaces -o yaml > deployed-manifests.yaml

  # Review with Claude Code
  claude code review-k8s-manifest deployed-manifests.yaml > audit-$(date +%Y-%m-%d).txt

  # Email results to team
  if grep -q "critical" audit-$(date +%Y-%m-%d).txt; then
    echo "Critical issues found!" | mail -s "K8s Manifest Audit" [email protected]
    cat audit-$(date +%Y-%m-%d).txt | mail -a "audit-$(date +%Y-%m-%d).txt" [email protected]
  fi

  sleep $REVIEW_INTERVAL
done

Schedule this to run weekly or daily depending on your deployment frequency. Over time, the audit logs become a historical record of how your manifests evolved and where you fixed issues. You can track progress: “Last quarter we had 47 critical issues across our manifests. This quarter it’s down to 12. We’re making progress.”

Common Pitfalls Teams Fall Into

Pitfall 1: Cargo Cult Configuration
Teams copy manifests from examples and commit them without understanding each field. Claude Code can explain what each setting does and why it matters. When you understand your manifests, you can modify them with confidence. This pitfall is particularly insidious because cargo cult manifests work initially—until load patterns change or failure scenarios surface. A manifest that works fine under light test load might completely fail under production traffic because the resource configuration was borrowed from a different type of workload. The developer who created it doesn’t know what each field does, so when issues arise, debugging becomes extremely difficult. They don’t know whether the problem is in the manifest, the application code, the Kubernetes version, or infrastructure constraints. Claude Code prevents this by explaining not just what each setting does, but why it exists and what happens when you modify it. Understanding your manifests means you can adapt them when your circumstances change.

Pitfall 2: Copy-Paste Without Customization
A manifest works for one service and gets copied for another without considering different resource requirements. A web service might run fine with 100m CPU. A machine learning model might need 4000m. Claude Code can catch when resource limits don’t match the workload. This happens constantly in large organizations. A deployment manifest is successful, so the team treats it as a template. They copy it for the next project without thinking about whether the requirements are similar. Resource constraints that are appropriate for a stateless microservice are catastrophically wrong for a stateful service like a database. Memory limits that are fine for a REST API are insufficient for a data processing job that loads files into memory. What you get is a growing collection of manifests across your infrastructure, all based on the same template, all configured for the wrong workload. When something breaks, you have an archeological problem: which of these manifests is responsible, and what was it supposed to do anyway? Claude Code helps by analyzing whether the resource configuration actually matches the apparent workload type. If you declare it’s a web service but request 8GB of memory and 4 CPU cores, Claude flags the mismatch.

Pitfall 3: Over-Optimization
Teams set limits too low trying to save resources, causing performance problems and timeouts. Claude Code can flag when limits seem too tight for the application, based on the image and typical application behavior. The temptation to over-optimize is real. Your company is paying for cloud resources. Someone looks at the bill and says “reduce resource limits and save money.” So teams start setting memory limits to the absolute minimum their application can run with—not the minimum it should run with, but the minimum it can theoretically fit in. They set CPU limits so low that response times increase 10x because the container is constantly being throttled. The application technically works, but it’s slow and unreliable. Users experience timeouts. Customers complain. Incident response costs exceed the monthly savings. Claude Code prevents this by understanding typical resource utilization for different application types. It can flag when you’re trying to run a Python data science application on 256Mi of memory when typical models need 2-4GB. It understands that a Java application needs a certain amount of heap space regardless of how much memory you wish you could give it.

Pitfall 4: Security Checkbox Syndrome
Teams add security settings because “we should” without understanding them. Claude Code can explain the security implications and whether settings actually help. Security settings are easy to add but easy to add wrongly. You see a tutorial recommending readOnlyRootFilesystem: true so you add it, only to discover that your application needs to write temporary files and now it fails. You add allowPrivilegeEscalation: false everywhere, including services that actually need to change user permissions as part of their startup. You drop all Linux capabilities, breaking services that need specific capabilities for their core function. The problem is that security settings come with tradeoffs, and those tradeoffs aren’t always obvious. Running with a read-only root filesystem is more secure but requires your application to be designed for it (write temp files to emptyDir volumes, not to /tmp on the root filesystem). Dropping capabilities is more secure but breaks services that legitimately need them. Claude Code helps by explaining the security implications of each setting, warning when a setting might be incompatible with your application, and suggesting alternatives when a security setting is too restrictive for your use case.

Pitfall 5: Ignoring the Blast Radius
A single misconfigured manifest can take down an entire service. Claude Code helps you understand the impact of each setting and whether it could cause cascading failures. Blast radius is a concept that’s easy to understand in theory but hard to reason about in practice. You reduce the number of replicas from 3 to 1 to save resources. Technically the application still works. But the blast radius just expanded enormously. A single node failure or a pod restart now means total service unavailability. A security update that requires rolling out new pods is now a scary operation because you have zero redundancy. Claude Code helps by understanding the replica strategy, checking for pod disruption budgets, analyzing affinity settings, and asking intelligent questions: “You have 2 replicas with pod anti-affinity set to preferred. In practice, this means both pods might run on the same node. If that node fails, your service is down. Is that acceptable?” This kind of reasoning prevents the configuration decisions that seem fine in isolation but create dangerous failure modes when viewed holistically.

Production Considerations: Making This Operational

In production, manifest review becomes continuous. Your cluster is running manifests that were deployed weeks or months ago. Have those best practices drifted? Are there newer, better approaches? Claude Code can periodically audit deployed manifests and flag things that need updating.

Build manifest review into your release process. Don’t just validate that YAML is syntactically correct. Validate that it’s architecturally sound, secure, and resource-efficient. Make this a team norm. When every manifest that ships has been reviewed by Claude, your Kubernetes infrastructure becomes more reliable, more secure, and easier to operate.

The teams with the most reliable Kubernetes deployments aren’t the ones with perfect manifests written on the first try. They’re the ones that review continuously, learn from mistakes, and apply those lessons systematically across their entire fleet.


-iNet | Building smarter Kubernetes deployments.

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.