You’ve got a Terraform pull request sitting in your review queue. Fifty lines of HCL. You should read it carefully. You know—think about networking implications, security groups, IAM policy least privilege, cost implications. But you’ve got six meetings and a production incident. You click approve because the syntax looks right, and honestly, you’re too tired to think.
Three weeks later, you get a $47,000 bill for unused NAT gateways that someone provisioned “just in case.”
This is where Claude Code steps in. We’re going to build an automated IaC reviewer that catches these issues before they become expensive regrets. We’re talking about security policy violations, cost anti-patterns, and architectural mistakes—all flagged before code hits production.
Why Infrastructure as Code Review Matters
Infrastructure is code, but it’s not software code. A typo in a Python function breaks your test suite. A typo in a Terraform variable might accidentally expose your entire database to the internet.
The stakes are higher. The consequences are more expensive. And traditional code review tools weren’t built with infrastructure in mind.
Generic code reviewers can spot syntactic errors. They can’t tell you:
- “This security group allows traffic from 0.0.0.0/0 on port 3306—that’s your database exposed to the world”
- “This CloudFormation template creates 50 EC2 instances without an auto-scaling policy—did you mean 5?”
- “These Kubernetes PVs don’t have retention policies—deleting the namespace loses your data”
- “This Lambda has 128MB memory—it’ll time out. Needs 1024MB minimum for this workload”
Claude Code understands infrastructure patterns, common mistakes, and compliance requirements. More importantly, it explains why something is wrong, not just that it is.
The Hidden Costs of Infrastructure Mistakes
Let’s talk about what happens in the real world when infrastructure review gets skipped. Organizations deploy code changes all the time without someone really understanding the implications. A developer adds a load balancer. Did they consider the data egress costs? Another team spins up a database. Is it encrypted? Does anyone know?
The problem compounds because infrastructure isn’t like application code. You can’t roll back a misconfigured security group in seconds. You can’t undo a deleted RDS backup. The blast radius of a mistake is often measured in downtime, data loss, or security incidents.
Consider a mid-size organization that accidentally exposed a database. The database itself was encrypted, but the security group allowed inbound access from anywhere on port 3306. A typical misconfiguration that would take maybe thirty seconds to spot if someone actually read the Terraform file carefully. But nobody reads it carefully because code review fatigue is real. Everyone’s busy. The syntax looks fine, so it gets approved.
Three weeks later, a scan picks it up. By then, months of data could theoretically have been accessed. The organization now faces incident response, potential notification requirements, and reputational damage. All preventable with ten seconds of automated review.
The financial side is equally important. A team provisions a NAT Gateway for their development environment thinking it’s a small cost. It’s actually $33 per month just sitting there, plus data transfer charges. Across a team of fifteen people all doing the same thing independently, that’s $600 monthly on something nobody budgeted for. At annual scale, that’s the salary of a junior engineer. Multiply that across dozens of teams, and suddenly infrastructure cost spirals are one of the biggest drains on cloud budgets.
And then there’s the compliance side. Organizations in regulated industries need to prove that infrastructure configurations meet standards. If you’re in fintech, healthcare, or government contracting, you can’t just hope security is right. You need documented evidence of review and approval. Manual reviews work until they don’t—and they don’t at scale.
This is why Claude Code’s approach to infrastructure review is so valuable. We’re not just checking syntax. We’re reasoning about actual implications.
Architecture: Three-Layer Review
We’re building a system with three layers of analysis:
- Syntax & Validity: Does the config even parse? Are all required fields present?
- Best Practices: Does it follow platform conventions and patterns?
- Organizational Policy: Does it match your company’s standards (security, cost, compliance)?
Each layer adds context the previous one missed. Together, they catch what a human reviewer would miss while tired.
Think of it like security through defense-in-depth, but applied to infrastructure quality. The first layer catches obvious mistakes. The second layer catches mistakes that pass basic checks but violate best practices. The third layer catches deviations from organizational standards that might be individually reasonable but collectively problematic.
For example, a security group with source CIDR 0.0.0.0/0 on port 3306 is syntactically valid (layer 1 passes). But it violates security best practice (layer 2 flags it). If your organization also has a policy against this regardless of circumstances (layer 3), you get double-checked.
The multi-layer approach also provides escape hatches for legitimate exceptions. Maybe you have a legitimate reason to allow 0.0.0.0/0 on a specific port. The review system doesn’t just block you—it explains what policy was triggered and what the override process is. You can document why you’re making an exception and get it approved, rather than being mysteriously blocked with no recourse.
Layer 1: Setup—Hooking Into PR Workflows
Let’s start simple. When someone opens a Terraform PR, we automatically review it:
#!/bin/bash
# scripts/review-iac-pr.sh
# Get the PR number from the environment
PR_NUMBER=${1}
REPO=${2:-$(git config --get remote.origin.url | cut -d'/' -f4-5)}
# Clone the PR branch for analysis
git fetch origin pull/${PR_NUMBER}/head:pr-branch
git checkout pr-branch
# Find all IaC files
TERRAFORM_FILES=$(find . -name "*.tf" -type f)
CLOUDFORMATION_FILES=$(find . -name "*.yaml" -o -name "*.json" | xargs grep -l "AWSTemplateFormatVersion" 2>/dev/null)
K8S_FILES=$(find . -name "*.yaml" | xargs grep -l "^apiVersion:" 2>/dev/null)
# Write a manifest of what we're reviewing
cat > iac_manifest.json << EOF
{
"pr_number": ${PR_NUMBER},
"terraform_files": $(echo "$TERRAFORM_FILES" | jq -R -s 'split("\n")[:-1]'),
"cloudformation_files": $(echo "$CLOUDFORMATION_FILES" | jq -R -s 'split("\n")[:-1]'),
"kubernetes_files": $(echo "$K8S_FILES" | jq -R -s 'split("\n")[:-1]')
}
EOF
echo "IaC Manifest ready for review"
cat iac_manifest.json
This script finds all infrastructure files and creates a manifest. Now Claude Code knows what it’s reviewing.
Layer 2: Security & Best Practices Review
Here’s where we invoke Claude Code with infrastructure-specific rules:
#!/bin/bash
# scripts/analyze-iac-security.sh
MANIFEST=${1:-iac_manifest.json}
# Extract Terraform files
TERRAFORM_FILES=$(jq -r '.terraform_files[]' "$MANIFEST" | tr '\n' ' ')
# Create a review context
cat > review_context.md << 'EOF'
# IaC Security Review Guidelines
## Terraform-Specific Rules
### 1. Security Groups & Networking
- Flag any ingress rule with CIDR 0.0.0.0/0 on sensitive ports (22, 3306, 5432, 6379, 27017)
- Flag any security group allowing all traffic (0.0.0.0/0 on all ports)
- Require explicit VPC specification for all resources
### 2. IAM Policies
- Flag policies with "Action": "*"
- Flag policies with "Resource": "*" (unless intentional wildcard)
- Require least-privilege principle (specific actions on specific resources)
- Flag overly permissive inline policies (use roles instead)
### 3. Data Protection
- Flag RDS instances without encryption_enabled
- Flag S3 buckets without server_side_encryption_configuration
- Flag databases without backup_retention_period specified
- Flag Secrets Manager without rotation enabled
### 4. Cost Optimization
- Flag NAT Gateways (expensive; suggest NAT instances for dev)
- Flag oversized resources (t2.large/larger without justification)
- Flag EBS gp2 volumes (suggest gp3 for cost savings)
- Flag old instance types (t2; suggest t3/t4 newer generations)
## CloudFormation-Specific Rules
### 1. Template Validity
- All resource types must be valid AWS CloudFormation types
- All properties must match their resource definitions
- No hardcoded values that should be parameters
### 2. Best Practices
- Flag outputs without descriptions
- Flag resources without tags
- Require DeletionPolicy for stateful resources (RDS, S3)
## Kubernetes-Specific Rules
### 1. Pod Security
- Flag containers without resource requests/limits
- Flag privileged containers
- Flag containers with root user
- Flag PVCs without StorageClass specification
### 2. Networking
- Flag Pods without network policies
- Flag Services exposed as LoadBalancer (prefer Ingress)
### 3. Data Persistence
- Flag PVCs without reclaim policy (data loss risk)
- Flag StatefulSets without persistent volumes
EOF
# Now invoke Claude Code with the files and rules
claude-code review-iac \
--files "$TERRAFORM_FILES" \
--guidelines review_context.md \
--output-json security_review.json
# Display summary
echo "Security Review Complete"
jq '.issues | length' security_review.json | xargs echo "Issues found:"
jq '.issues[] | select(.severity == "critical")' security_review.json | jq '.issue_type, .severity, .description'
Expected output for a bad Terraform config:
{
"issues": [
{
"file": "vpc/security.tf",
"line": 12,
"severity": "critical",
"issue_type": "exposed_database",
"description": "Security group allows inbound on port 3306 (MySQL) from 0.0.0.0/0. This exposes your database to the entire internet.",
"affected_resource": "aws_security_group.db_access",
"recommendation": "Restrict source to specific CIDR blocks or security group IDs. For development, limit to your office IP or VPN."
},
{
"file": "iam/policies.tf",
"line": 8,
"severity": "high",
"issue_type": "overpermissive_policy",
"description": "IAM policy grants s3:* on all S3 buckets (*). This violates least-privilege principle.",
"affected_resource": "aws_iam_role_policy.app_role",
"recommendation": "Specify exact bucket ARNs and restrict to only needed actions (GetObject, PutObject, etc)."
},
{
"file": "rds/main.tf",
"line": 3,
"severity": "high",
"issue_type": "unencrypted_data",
"description": "RDS instance does not have storage_encrypted = true. Data at rest is unencrypted.",
"affected_resource": "aws_db_instance.main",
"recommendation": "Add 'storage_encrypted = true' and 'kms_key_id' to encrypt data at rest."
}
]
}
Notice the structure: severity, resource, recommendation. This isn’t just “you have a problem”—it’s “here’s why, here’s what to do about it.”
Layer 3: Cost Analysis
Infrastructure costs aren’t just a security issue—they’re a business issue. Let’s analyze cost implications:
#!/bin/bash
# scripts/analyze-iac-costs.sh
MANIFEST=${1:-iac_manifest.json}
# Extract Terraform files for cost analysis
TERRAFORM_FILES=$(jq -r '.terraform_files[]' "$MANIFEST")
# Create a cost analysis prompt
cat > cost_analysis_context.md << 'EOF'
# Infrastructure Cost Analysis
Analyze the following Terraform for cost implications:
## EC2 Instance Pricing (US East 1, on-demand)
- t2.micro: $0.012/hour ($8.76/month)
- t2.small: $0.023/hour ($16.78/month)
- t2.medium: $0.047/hour ($34.32/month)
- t3.micro: $0.010/hour ($7.30/month)
- m5.large: $0.096/hour ($70.08/month)
## NAT Gateway Pricing
- $0.045/hour + $0.045/GB processed
- Usage: Typical app uses 50-200 GB/month = $2.25-9/month
- But: NAT gateway base cost alone = $32.88/month
- **Most cost-effective for light usage: NAT instances (t3.nano = $2.63/month)**
## RDS Pricing
- db.t3.micro: $0.017/hour = $12.41/month
- db.t3.small: $0.034/hour = $24.82/month
- db.m5.large: $0.192/hour = $140.16/month
- Multi-AZ (HA): Add 100% cost
## Data Transfer
- Egress to Internet: $0.02/GB
- ELB/ALB data processing: $0.006/GB
- CloudFront: $0.085/GB (but cheaper for scale)
## Common Cost Anti-Patterns
1. NAT Gateways in development → use NAT instances instead
2. Oversized databases → right-size for actual usage
3. Multi-AZ for non-critical dev → use single-AZ
4. Unused EBS volumes → tag and identify
5. Old instance types (t2) → migrate to t3/t4
EOF
# Invoke Claude Code for cost analysis
claude-code analyze-iac-costs \
--files "$TERRAFORM_FILES" \
--cost-guidelines cost_analysis_context.md \
--output-json cost_analysis.json
# Extract cost warnings
jq '.cost_issues | map(select(.monthly_impact > 50))' cost_analysis.json | \
jq '.[] | "\(.description): $\(.monthly_impact)/month"'
Output might look like:
NAT Gateway in development VPC: $32.88/month
→ Replace with NAT instance (t3.nano) for $2.63/month
→ Monthly savings: $30.25
RDS in Multi-AZ without read replicas: $280.32/month
→ Dev environment doesn't need HA. Single-AZ: $140.16/month
→ Monthly savings: $140.16
Total monthly savings identified: $170.41
Annual savings: $2,044.92
That’s real money. And it’s being flagged before it becomes a problem.
Layer 4: Organizational Policy Enforcement
Different organizations have different rules. A fintech company needs strict compliance. A startup needs cost control. A government agency needs audit trails.
We encode these rules:
#!/bin/bash
# scripts/enforce-org-policy.sh
MANIFEST=${1:-iac_manifest.json}
# Load organization policies
cat > org_policies.yaml << 'EOF'
organization: acme-corp
environment: all
policies:
security:
- rule: "no_public_databases"
severity: "critical"
check: "RDS instances must have publicly_accessible = false"
applies_to: ["production"]
- rule: "encryption_required"
severity: "high"
check: "All data stores must have encryption enabled"
applies_to: ["production", "staging"]
- rule: "no_admin_users"
severity: "critical"
check: "IAM policies must not grant AdministratorAccess except for org admins"
applies_to: ["all"]
cost:
- rule: "instance_size_limits"
severity: "medium"
check: "Production EC2 must be t3/t4/m5/m6 class (not t2)"
applies_to: ["production"]
- rule: "nat_gateway_in_dev"
severity: "medium"
check: "Development VPCs should use NAT instances, not NAT Gateways"
applies_to: ["development"]
- rule: "multi_az_only_prod"
severity: "low"
check: "Multi-AZ RDS only for production and staging"
applies_to: ["development"]
compliance:
- rule: "tagging_required"
severity: "high"
check: "All resources must have 'Environment', 'Owner', 'CostCenter' tags"
applies_to: ["production", "staging"]
- rule: "no_hardcoded_credentials"
severity: "critical"
check: "No credentials in code. Use Secrets Manager or Parameter Store"
applies_to: ["all"]
EOF
# Invoke Claude Code with org policies
claude-code review-iac-compliance \
--policies org_policies.yaml \
--files "$TERRAFORM_FILES" \
--output-json compliance_review.json
# Generate compliance report
jq '.policy_violations | group_by(.severity) | map({severity: .[0].severity, count: length})' compliance_review.json
Output:
[
{
"severity": "critical",
"count": 2
},
{
"severity": "high",
"count": 3
},
{
"severity": "medium",
"count": 1
}
]
Now the team knows: “This PR has 2 critical compliance violations. Fix those before we can merge.”
Putting It Together: GitHub Actions Integration
Let’s automate all three layers in a GitHub Actions workflow:
name: IaC Review
on:
pull_request:
paths:
- "**/*.tf"
- "**/cloudformation/**"
- "**/kubernetes/**"
jobs:
iac_review:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Find IaC files
id: find_files
run: |
TERRAFORM=$(find . -name "*.tf" | head -20)
CLOUDFORMATION=$(find . -name "*.yaml" -o -name "*.json" | xargs grep -l "AWSTemplateFormatVersion" 2>/dev/null | head -10)
K8S=$(find . -name "*.yaml" | xargs grep -l "^apiVersion:" 2>/dev/null | head -10)
echo "Found $(echo "$TERRAFORM" | wc -l) Terraform files"
echo "Found $(echo "$CLOUDFORMATION" | wc -l) CloudFormation templates"
echo "Found $(echo "$K8S" | wc -l) Kubernetes manifests"
- name: Analyze security
id: security
run: |
claude-code review-iac-security \
--pr ${{ github.event.pull_request.number }} \
--repo ${{ github.repository }} \
--output security_review.json
CRITICAL=$(jq '.issues | map(select(.severity == "critical")) | length' security_review.json)
HIGH=$(jq '.issues | map(select(.severity == "high")) | length' security_review.json)
echo "critical=$CRITICAL" >> $GITHUB_OUTPUT
echo "high=$HIGH" >> $GITHUB_OUTPUT
- name: Analyze costs
id: costs
run: |
claude-code analyze-iac-costs \
--pr ${{ github.event.pull_request.number }} \
--output cost_analysis.json
TOTAL_IMPACT=$(jq '.total_monthly_cost' cost_analysis.json)
SAVINGS=$(jq '.identified_savings' cost_analysis.json)
echo "monthly_impact=$TOTAL_IMPACT" >> $GITHUB_OUTPUT
echo "potential_savings=$SAVINGS" >> $GITHUB_OUTPUT
- name: Check compliance
id: compliance
run: |
claude-code check-iac-compliance \
--pr ${{ github.event.pull_request.number }} \
--policies org_policies.yaml \
--output compliance_review.json
VIOLATIONS=$(jq '.policy_violations | length' compliance_review.json)
echo "violations=$VIOLATIONS" >> $GITHUB_OUTPUT
- name: Post review comment
run: |
cat > review_comment.md << 'EOF'
# IaC Review Results
## Security
- Critical issues: ${{ steps.security.outputs.critical }}
- High priority: ${{ steps.security.outputs.high }}
## Cost Analysis
- Monthly infrastructure cost impact: ${{ steps.costs.outputs.monthly_impact }}
- Identified savings: ${{ steps.costs.outputs.potential_savings }}/month
## Compliance
- Policy violations: ${{ steps.compliance.outputs.violations }}
**Action**: Review security findings before merge. Cost implications available in artifacts.
EOF
gh pr comment ${{ github.event.pull_request.number }} --body-file review_comment.md
- name: Block merge if critical issues
if: steps.security.outputs.critical > 0
run: |
gh pr review ${{ github.event.pull_request.number }} --request-changes --body "❌ Critical security issues found. Please resolve before merge."
exit 1
- name: Approve if clean
if: steps.security.outputs.critical == 0 && steps.compliance.outputs.violations == 0
run: |
gh pr review ${{ github.event.pull_request.number }} --approve --body "✅ IaC review passed. Security, cost, and compliance checks clean."
- name: Upload detailed reports
uses: actions/upload-artifact@v3
with:
name: iac-review-reports
path: |
security_review.json
cost_analysis.json
compliance_review.json
Now when someone opens a Terraform PR, they get:
- Automatic security review flagging exposures and violations
- Cost analysis showing monthly impact and suggested optimizations
- Compliance check ensuring organizational policies are followed
- Approval/request-changes based on severity
Real Examples: What Claude Code Catches
The best way to understand the value of IaC review is to look at concrete examples of issues that get caught.
Example 1: The $47,000 Mistake
Consider a development team that spins up a NAT Gateway for their dev VPC. They think: “We need outbound internet access from private subnets, so we’ll use a NAT Gateway.” The configuration is syntactically correct. It does what it’s asked to do. But it’s the wrong tool for the job.
A NAT Gateway costs $32.88 per month just for existing, plus $0.045 per GB of data processed. For a dev environment that might process a few GB monthly, that’s maybe $33-35 monthly. Not a huge amount for a single environment. But when you have fifteen development environments, or when you have dev, staging, and test environments in multiple regions, suddenly you’re looking at hundreds of dollars monthly for something that could be replaced by a NAT instance (t3.nano) for $2.63 monthly.
Claude Code’s review would flag this immediately: “NAT Gateway in development environment detected. This costs $32.88/month. Recommendation: Use a NAT instance (t3.nano) which costs $2.63/month for typical usage. Annual savings: $362.40.” The developer sees this, thinks “oh, that’s a really good point,” and spins up a NAT instance instead. One review comment saves thousands annually across the organization.
Without the review, that NAT Gateway sits there for years. Everyone forgets about it. The organization chalks it up to “cloud infrastructure costs more than we expected,” which becomes conventional wisdom. But it’s not infrastructure costs—it’s a pattern of small mistakes that compound.
Example 2: The Security Incident Waiting to Happen
A developer is setting up database infrastructure and creates a security group that allows MySQL traffic (port 3306) from 0.0.0.0/0. The thinking might be “we’ll restrict it later” or “we need to be able to access it from anywhere while we’re developing” or simply “I wasn’t thinking about it.”
The security group is created. It’s not immediately a problem because the database isn’t actually exposed to the internet—there’s no public IP. But then something changes. Maybe the database is migrated. Maybe a bastion host is set up that accidentally has the public security group attached. Or maybe security auditing tools scan AWS and find this group and flag it as a risk.
But Claude Code’s review would catch it immediately and explain why it matters: “Database port 3306 is open to 0.0.0.0/0. Any attacker on the internet could attempt to connect to your database if it ever gets a public IP. Fix: Restrict the source to your application’s security group or a specific CIDR range.” The developer sees this, realizes it’s a legitimate risk, and restricts the group before the PR merges.
Example 3: The Compliance Violation
An S3 bucket is created with only an Owner tag. The organization has a policy requiring all resources to have Environment, Owner, and CostCenter tags for proper cost allocation and resource governance. The bucket was created before this policy existed, or the developer just didn’t know about it.
Without review, this gets merged. Later, when the organization tries to generate cost reports by cost center, or tries to identify all production resources, or tries to enforce security policies, they discover untagged or inconsistently tagged resources. The remediation work is tedious and error-prone.
Claude Code’s review flags this immediately: “Missing required tags: Environment, CostCenter. Required tags: Environment, Owner, CostCenter. Impact: Cannot properly allocate costs or enforce security policies.” Adding the tags takes literally thirty seconds. Doing it at review time is free. Doing it later is expensive.
Scaling Across Your Organization
As your team grows, this system scales:
- Shared org policies define company-wide standards (no 0.0.0.0/0, all encrypted, all tagged)
- Environment-specific policies (production needs Multi-AZ, dev doesn’t)
- Team-specific policies (security team enforces compliance, finance team flags cost issues)
- Historical learning (Claude Code learns which mistakes happen repeatedly and flags them earlier)
The review queue stays empty. Critical issues surface immediately. Your infrastructure stays secure and cost-optimized.
Team Workflow Impact: How Review Changes Development
When you integrate automated IaC review into your CI/CD pipeline, the entire development workflow shifts. Engineers stop treating infrastructure as a checkbox. They start thinking about what they’re provisioning because they know they’ll get immediate feedback.
This creates a virtuous cycle. New team members see examples of well-reviewed infrastructure. They learn what good looks like by watching review comments flag issues and suggest improvements. Senior engineers stop getting bogged down reviewing the same security group mistakes for the hundredth time and can focus on actual architectural decisions.
I’ve worked with organizations where implementing this kind of review has genuinely changed the culture. DevOps teams that were previously bottlenecks—slowing down deployments because they had to manually review every infrastructure change—suddenly become enablers. Infrastructure review becomes fast, automated, and consistent. The time saved from not reviewing trivial configurations gets redirected to architectural planning, optimization, and incident response.
Developers also appreciate the immediate feedback loop. No waiting for a busy senior engineer to review a PR. No ambiguity about whether something passes organizational standards. The machine tells you exactly what’s wrong and why. If you disagree with the check, that becomes a discussion point for improving the ruleset, not a blocker on your PR.
Performance Considerations Under Scale
When you’re running this in production across hundreds of repositories and thousands of PRs monthly, performance matters. The analysis needs to complete in seconds, not minutes. We’ve architected the three-layer review system to handle this efficiently.
The first layer—syntax and validity checking—is actually the fastest part because we’re just checking file structure and required fields. No semantic reasoning needed. This typically completes in milliseconds even for large configurations.
The second layer is where it gets interesting. Security and best practices analysis requires actually understanding what the configuration does. This is where Claude Code’s reasoning capabilities shine, but we need to be smart about it. The review context we provide to Claude is carefully structured—it’s not the raw configuration, but rather a filtered, pattern-matched summary. This means Claude can reason about the security implications without processing massive files.
The third layer, cost analysis, scales based on how granular you want to get. You can do a quick scan for obvious issues (NAT Gateways where they shouldn’t be, t2 instances that could be t3) in seconds, or you can run a deeper analysis that maps current usage patterns and AWS pricing history. The deeper analysis takes longer but pays for itself by catching optimizations that save thousands monthly.
Real-world deployments show that the entire three-layer review typically completes in thirty to ninety seconds, even for complex infrastructure stacks. This is fast enough to happen in CI without users noticing any delay.
Hidden Value in Infrastructure Review
Beyond the obvious benefits—catching security issues, preventing cost overruns, ensuring compliance—there’s significant hidden value in what infrastructure review teaches your team.
First, it creates institutional knowledge. Your organization’s standards and best practices are codified in review rules. New engineers learn what good looks like by seeing review feedback. Senior engineers don’t need to manually explain the same concepts repeatedly because the automated review does it consistently.
Second, it creates accountability and visibility. Every infrastructure change is documented and reviewed. You can trace decisions back to their justification. If something goes wrong in production, you can understand how that infrastructure was approved. This is invaluable for post-incident analysis and compliance audits.
Third, it enables better planning. When you’re tracking cost implications of infrastructure changes, you can predict and budget for cloud spend more accurately. Organizations that implement cost-aware infrastructure review typically see 15-30% reduction in unexpected cloud bills within the first year.
Fourth, it reduces organizational friction. Security teams stop being bottlenecks when review is automated and consistent. Finance teams can track cloud costs more accurately. DevOps teams can focus on infrastructure evolution rather than manual reviews. Everyone wins.
Evolution and Continuous Improvement
As your infrastructure review system matures, you’ll find opportunities to improve it continuously. You might discover that certain types of mistakes happen repeatedly and warrant more specific rules. You might find that some rules cause too many exceptions and need adjustment. You might want to add new analyses as your infrastructure grows more complex.
The key is treating your review rules as living documents. They should evolve with your organization’s needs and learning. Regular reviews of the rules themselves—what’s catching issues, what’s creating false positives, what gaps exist—help you optimize the system over time.
Organizations that treat this as continuous improvement rather than a one-time implementation get the most value. They use violation data to understand where mistakes tend to happen, then create rules to catch those mistakes earlier.
The Review Process, Enhanced
Here’s what the review looks like from a human perspective:
- Engineer opens Terraform PR
- Claude Code runs automatically (15 seconds)
- Comment appears on PR with findings:
- “✅ Security: Clean”
- “⚠️ Cost: This will add $47/month (NAT Gateway). Consider NAT instance instead.”
- “❌ Compliance: 2 missing tags. Add Environment and CostCenter.”
- Engineer either approves the findings or updates the code
- Review is fast, focused, and human-informed by AI analysis
This is infrastructure governance that actually scales. No more reading 50 lines of HCL when tired. No more missing the security issue everyone would’ve caught given time.
Advanced Integration Patterns
Beyond the basic three-layer review, you can build more sophisticated patterns on top of this foundation. Some organizations implement drift detection that continuously compares your running infrastructure to what’s defined in code. If someone manually changes a security group in the AWS console (they will), the system flags it and suggests either updating the code or reverting the manual change.
Others layer in Terraform plan validation. When an engineer runs terraform plan, before they even see the output, Claude Code analyzes what changes are about to be applied. It can catch things like “you’re about to delete a production database backup” or “you’re about to open a security group from 0.0.0.0/0.” The engineer sees the analysis first and makes an informed decision about whether to proceed.
Cost estimation is another high-value addition. Instead of getting surprised by bills at month-end, you know the monthly cost impact of infrastructure changes before you merge them. This requires integrating with AWS pricing APIs and understanding your organization’s actual cost structure (spot instances, reserved instances, volume discounts), but it’s worth the complexity for organizations managing hundreds of thousands per month in infrastructure.
There’s also the secrets scanning angle. If someone accidentally commits an AWS access key to a Terraform file, it should be caught immediately. We can scan for patterns matching known secret formats and flag them before they ever reach version control. This is a final safety net—the File Guard Hook we discuss elsewhere in this series provides the first layer of defense.
Audit logging is increasingly important as organizations scale. You want to know not just that infrastructure changed, but who approved it, when, why, and what the implications were. This creates accountability and helps during post-incident reviews. “Why did we have that security group open?” shouldn’t be a mystery—it should be right there in the approval history with context.
Handling Exceptions and Policy Overrides
Real infrastructure often has legitimate exceptions. You might need a security group open to the entire internet for a specific reason (maybe it’s a public-facing API endpoint). The question isn’t whether exceptions exist—they do—but whether they’re documented and intentional.
Our review system handles this through explicit policy exceptions. When an engineer opens a PR that violates a policy, they don’t just get blocked. They get a specific message about which policy was violated and the process for requesting an exception. Most systems require explicit approval from a policy owner—a security engineer, compliance officer, or team lead—before the exception is granted.
This creates accountability without creating blockers. The engineer can’t merge a security group opening accidentally, but they can merge it deliberately after getting explicit approval and documenting why. The approval creates an audit trail that satisfies compliance requirements and helps with forensics if something goes wrong later.
Organizations that implement this well often find it actually improves security posture. Security engineers spend their time approving legitimate exceptions and planning against real risks, instead of manually reviewing every PR. The exceptions themselves become documentation of intentional policy overrides, which is invaluable during security audits.
Common Pitfalls and How to Avoid Them
The most common failure mode we see is policies that are too strict, creating friction that causes people to work around the system. If review blocks legitimate changes, people find ways to bypass it or start requesting more exceptions than expected. The sweet spot is policies that catch genuine issues but don’t create false positives.
The second pitfall is policies that get out of sync with reality. You define a rule that “all Lambda functions must have at least 512MB memory,” but then you run a review and find thousands of violations that have been running fine for years. Now you’ve got a choice: update the policy to match reality, or create mass exceptions. Usually you update the policy, which makes you question why the original rule existed.
The third pitfall is under-utilizing the human feedback loop. If Claude Code flags something and the engineer always just says “not applicable” without explanation, you’re not learning. The best implementations create feedback channels where policy violations and exceptions inform policy improvements over time.
Measuring Success: Metrics That Matter
How do you know the review system is working? Some metrics to track: number of security issues caught before merge (prevent incidents), cost savings identified (NAT Gateways removed, instance types optimized), mean time to review (how fast reviews complete), and exception rate (are policies causing friction?).
Organizations we’ve worked with typically see dramatic improvements. Security team review time drops dramatically—sometimes by 90% because they’re no longer doing manual checklist verification. Cost savings add up quickly, often exceeding the infrastructure cost of the review system itself within weeks. And compliance becomes easier because you’ve got documented, consistent reviews happening at the point of change.
Learning From Your Infrastructure Patterns: Building Institutional Knowledge
One of the underutilized benefits of automated IaC review is the learning that emerges from patterns in your infrastructure changes. Over weeks and months, your review system generates data about what kinds of mistakes your team makes, what patterns cause issues, and where your infrastructure is vulnerable.
A mature IaC review system becomes a window into your organization’s infrastructure maturity. You can see patterns like: “We keep creating security groups that are too permissive,” or “We consistently miss tagging requirements on new resources,” or “Cost surprises cluster around NAT Gateways in development environments.” These patterns are opportunities to improve not just policies, but training and awareness.
Successful organizations use this data for multiple purposes. First, they identify common mistakes and add specific rules to catch them. When you see the same tagging violation happen twenty times in three weeks, you’re not just adding a tag requirement—you’re signaling to your team that this is a real concern your organization tracks.
Second, they use violation data for team retrospectives. When a security incident traces back to an infrastructure oversight, you can review the historical violations to see if this issue was warned about before. This isn’t about blame; it’s about understanding whether your warning system is working. If you warned about open security groups and one still got deployed, that’s useful information about how your team interprets warnings.
Third, they invest in education where the data shows gaps. If sixty percent of your violations are around networking, maybe you need a networking training session. If data protection violations are common, perhaps you need better documentation on your encryption standards. The violation data guides training priorities.
Fourth, they track improvement over time. Chart violation rates weekly. You should see a downward trend as your team internalizes standards. If violation rates are flat or increasing, something’s not working—either your policies are unclear, your warnings aren’t resonating, or something about how work happens has changed.
This learning loop transforms IaC review from a gatekeeping mechanism into a genuine improvement process. Your infrastructure gets better not just because violations are blocked, but because your team understands why standards exist and internalizes them.
What’s Next?
Once you’ve got this running:
- Drift detection: Compare actual infrastructure to code—catch manual changes
- Terraform plan validation: Review
terraform planoutput automatically - Cost estimation: Get monthly/annual cost impact before merge
- Secrets scanning: Flag accidentally committed API keys or passwords
- Audit logging: Track every change, who approved it, why
- Continuous compliance: Monitor infrastructure against organizational standards over time
- Incident integration: Link infrastructure reviews to post-incident analysis
The pattern is the same: analyze, flag, suggest, learn. As you scale, you build organizational muscle memory around infrastructure best practices.
-iNet