All Articles Claude Code

Claude Code Model Selection: Choosing the Right Brain

So you've got Claude Code running. Nice. But here's the thing—not every task needs the same brain power.

So you’ve got Claude Code running. Nice. But here’s the thing—not every task needs the same brain power. Sure, you could throw Opus at everything, but that’s like using a Tesla to grab milk from the corner store. It works, but your wallet will hate you.

Let me walk you through how Claude Code handles model selection, when to pick which model, and how to make smart decisions that balance capability with cost. This isn’t just about saving money—it’s about shipping faster by using the right tool at the right moment.

The Three Models: Quick Rundown

You’ve got three neural networks at your fingertips in Claude Code. Each trades power for speed and cost.

Opus 4.6: The heavyweight. This is the most capable Claude model—deeper reasoning, better code architecture analysis, handles complex multi-step problems that would trip up smaller models. Best for when you’re cracking hard problems. Think: designing microservices architectures, security audits, novel algorithm design, cross-domain analysis. Opus doesn’t just understand code; it understands the system behind the code.

Sonnet 4.6: The balanced choice. This is the sweet spot for most daily work. Fast enough to not waste your time, capable enough to handle real development tasks without bailing out on you midway. It’s the model that handles 95% of feature development, debugging, test writing, and refactoring. It’s your workhorse.

Haiku 4.5: The speedster. Fastest and cheapest. Great for bulk operations, simple tasks, and spinning up multiple subagents in parallel without burning through your quota. Haiku is your parallel processor—when you need 50 things done fast, this is your guy.

Here’s the thing nobody tells you: Haiku is genuinely good. It’s not a toy. It handles most coding tasks better than you’d expect. Where it starts sweating is complex architecture decisions and multi-step reasoning across 5+ cascading logic branches.

Understanding Model Trade-Offs

Before we dive into configuration, let’s be clear about what you’re trading:

Speed vs. Capability: Haiku responds in 2-3 seconds. Opus takes 15-30 seconds. For simple tasks, Haiku is instant; for complex reasoning, Opus’s extra thinking is worth the wait.

Cost vs. Quality: Haiku costs ~3-4x less per token than Opus. But Opus catches subtle bugs Haiku misses. On a $200 security audit? Opus pays for itself. On generating boilerplate? Haiku’s overkill.

Parallelism vs. Depth: You can run 10 Haiku tasks in parallel cheaper than 1 Opus task sequentially. Sometimes width beats depth. Sometimes you need the depth.

Context Efficiency: Haiku handles shorter contexts better. If your task fits in 2000 tokens, Haiku’s perfect. If it needs 10,000 tokens of context for Opus to reason properly, that’s a choice you’re making.

The goal isn’t to always use the most powerful model. It’s to match capability to task.

Model Configuration: Where to Set It

You can configure your default model in two places:

Option 1: settings.json

The primary way is through Claude Code’s settings file. You’ll find it here (path varies by OS):

macOS/Linux:

~/.claude/settings.json

Windows:

%APPDATA%\Claude\settings.json

Inside, you’ll set your default model like this:

{
  "default_model": "claude-opus-4-20250514",
  "extended_thinking": {
    "enabled": false,
    "budget_tokens": 10000
  },
  "performance": {
    "timeout_seconds": 120,
    "max_retries": 3
  }
}

The default_model key controls what Claude Code uses when you don’t explicitly specify one. Set it once, and that becomes your baseline. Change it, and restart Claude Code for it to take effect.

Important: Settings load at startup. If you change settings.json while Claude Code is running, it won’t pick up the changes. Restart the CLI or reload the configuration explicitly.

Option 2: CLI Flags

When you invoke Claude Code from the terminal, you can override the default:

claude code --model claude-opus-4-20250514 "Analyze this production bug"

This single invocation uses Opus, but your settings.json remains unchanged. Handy for one-off tasks where you need more firepower. You can also use shorthand:

claude code --model opus "Task here"
claude code --model sonnet "Task here"
claude code --model haiku "Task here"

Claude Code interprets the shorthand and expands it to the full model name.

The Model Names (2026 Era)

Here are the exact model identifiers you’ll use:

Opus 4.6:   claude-opus-4-20250514
Sonnet 4.6: claude-sonnet-4-20250514
Haiku 4.5:  claude-haiku-4-5-20251001

These are version-pinned. Anthropic updates models quarterly. If these dates feel stale by the time you read this, check the official documentation. The pattern stays the same, but the dates change.

When to Use Each Model: Decision Framework

Let me give you a framework that actually works in practice:

Use Opus When:

  • Complex Architecture: Designing multi-service systems, refactoring large codebases, evaluating architectural tradeoffs between microservices vs. monolith
  • Hard Problem-Solving: Multi-step reasoning, complex algorithms, ambiguous requirements that need unpacking, constraint satisfaction problems
  • Code Review: Analyzing existing code for subtle bugs, security vulnerabilities, performance problems, design flaws
  • Novel Patterns: Implementing something your team hasn’t done before, novel authentication flows, custom DSLs
  • High Stakes: Production database migrations, security audits, decisions that affect architecture for years
  • Cost isn’t a constraint: You’re billing the work to a client or the decision doesn’t move the needle on your budget

Example task:

Opus is your pick: "I'm redesigning our API from REST to GraphQL.
We have 40+ endpoints. Walk me through the migration strategy,
schema design, potential pitfalls, and rollback plan."

Use Sonnet When:

  • Standard Development: Most day-to-day coding tasks, writing features, fixing bugs
  • Code Generation: Writing boilerplate, tests, documentation, scaffolding
  • Debugging: Most debugging scenarios (unless it’s architectural)
  • Balanced Speed/Quality: You want results in seconds, not minutes
  • Most Production Work: This should be your default for real projects
  • Feature Development: Implementing requirements, extending existing systems

Example task:

Sonnet is your pick: "Write a Jest test suite for this React
component. Cover happy path, edge cases, and error states."

Use Haiku When:

  • Bulk Operations: Running the same task 50 times, need speed over perfection
  • Simple Tasks: Single-function rewrites, linting fixes, boilerplate generation, string transformations
  • Subagent Work: Spinning up many agents in parallel (you’ll learn about this below)
  • Rapid Iteration: You’re in discovery mode, want fast feedback loops, trying things out
  • Budget-Conscious: You’re on a token quota or cost matters
  • Mechanical Transformations: Format conversions, code style updates, simple refactoring

Example task:

Haiku is your pick: "Here's a list of 20 Python files.
Add type hints to the function signatures in each."

Configuring Models for Subagents

Here’s where it gets powerful. When you dispatch subagents in Claude Code, you can set their model independently:

/dispatch story-architect --model claude-haiku-4-5-20251001 "Outline a cyberpunk thriller"

This is game-changing for parallel work. Want to spin up 5 research agents simultaneously? Use Haiku for all of them, save tokens, still get useful output.

Real workflow example:

# Cheap parallel research phase
/dispatch research-validator --model claude-haiku-4-5-20251001 "Validate tech facts"
/dispatch world-builder --model claude-haiku-4-5-20251001 "Expand magic system"
/dispatch character-forge --model claude-haiku-4-5-20251001 "Develop antagonist"

# Then have Opus orchestrate results
/dispatch prose-generator --model claude-opus-4-20250514 "Write opening chapter"

You’re leveraging the speed/cost of Haiku for parallel work, then hitting the hard synthesis task with Opus’s reasoning. This pattern—parallel Haiku, Opus synthesis—is how you scale smart.

Extended Thinking and Output Quality

There’s another lever you can pull: extended thinking. This lets Claude spend more tokens on internal reasoning before responding.

In your settings:

{
  "extended_thinking": {
    "enabled": true,
    "budget_tokens": 10000
  }
}

What happens: Claude internally reasons for up to 10,000 tokens (your budget), then gives you a refined answer. You don’t see this reasoning unless you ask, but the quality of the final answer improves.

Extended thinking with Haiku: Sometimes better than regular Opus. You’re trading speed for depth. For a complex problem, Haiku + 10k thinking tokens can outthink Sonnet.

Extended thinking with Opus: Overkill for simple tasks, brilliant for hard problems. You’re getting deep reasoning plus deep capability.

The trade-off is latency. Extended thinking takes longer. But for complex decisions, it’s worth it.

Example with extended thinking:

Config: claude-opus-4-20250514 + extended_thinking (10k tokens)
Task: "Here's a 5000-line codebase. Identify the top 3 architectural
issues and propose refactoring."

Result: Opus takes 20-30 seconds, but gives you deeply-reasoned
architectural insights instead of surface-level suggestions.

When extended thinking helps:

  • Complex trade-off analysis
  • Security vulnerability hunting
  • Architectural decisions
  • Novel algorithm design
  • Evaluating multiple competing approaches

When it’s overkill:

  • Writing simple functions
  • Fixing typos
  • Generating boilerplate
  • Writing documentation
  • Quick debugging of known issues

Cost-Per-Task Analysis: Real Numbers

Let me break down what you’re actually paying (as of March 2026):

Input Token Cost

  • Opus: ~$3 per million input tokens
  • Sonnet: ~$3 per million input tokens
  • Haiku: ~$0.80 per million input tokens

Output Token Cost

  • Opus: ~$15 per million output tokens
  • Sonnet: ~$15 per million output tokens
  • Haiku: ~$4 per million output tokens

Real Task Estimates

Task 1: Write a Test Suite (3000 input tokens, 2000 output tokens)

Haiku:   $0.003 + $0.008 = $0.011
Sonnet:  $0.009 + $0.030 = $0.039
Opus:    $0.009 + $0.030 = $0.039

Haiku is 3-4x cheaper and delivers 90% of the quality.

Task 2: Architectural Review (5000 input tokens, 4000 output)

Haiku:   $0.004 + $0.016 = $0.020
Sonnet:  $0.015 + $0.060 = $0.075
Opus:    $0.015 + $0.060 = $0.075

Haiku works for many reviews. But Opus catches subtle issues Haiku misses—security problems, scalability bottlenecks, coupling issues that cause problems years later.

Task 3: Complex Multi-Step Design (10000 input, 8000 output)

Haiku:   $0.008 + $0.032 = $0.040
Sonnet:  $0.030 + $0.120 = $0.150
Opus:    $0.030 + $0.120 = $0.150

This is where Opus earns its cost. Haiku struggles with multi-step reasoning across many decision branches.

Smart Model Selection Strategy

Here’s a decision tree I actually use:

Is this a quick question?
├─ Yes → Is it architectural?
│         ├─ Yes → Opus
│         └─ No → Haiku (maybe Sonnet if unsure)
└─ No → Is this a one-off or repeated?
        ├─ One-off → Opus if hard, Sonnet if moderate, Haiku if simple
        └─ Repeated → Always Haiku (you'll run it 50+ times)

Is this going to production?
├─ Yes → Never default to Haiku. Use Sonnet minimum.
└─ No → Optimize for cost. Haiku until it fails.

Am I spinning up subagents?
├─ Yes → All Haiku in parallel, Opus for synthesis
└─ No → Pick based on task complexity

Configuration Cheat Sheet

Here’s a production-ready settings.json:

{
  "default_model": "claude-sonnet-4-20250514",
  "extended_thinking": {
    "enabled": true,
    "budget_tokens": 5000
  },
  "performance": {
    "timeout_seconds": 120,
    "max_retries": 3,
    "fast_mode_enabled": true
  },
  "models": {
    "heavy_lifting": "claude-opus-4-20250514",
    "daily_work": "claude-sonnet-4-20250514",
    "bulk_operations": "claude-haiku-4-5-20251001"
  }
}

Then override at the command line when needed:

# Heavy architectural work
claude code --model claude-opus-4-20250514 "Design decision..."

# Quick question
claude code --model claude-haiku-4-5-20251001 --fast "Quick fix?"

# Subagent dispatch
/dispatch character-forge --model claude-haiku-4-5-20251001 "Task..."

Performance Benchmarks: Real-World Timing

Let me give you actual timing data so you can predict latency for your workflow:

Response Latency (measured in March 2026)

Simple Task (300 input tokens, generate 200 output):

Haiku:   ~2-3 seconds
Sonnet:  ~3-5 seconds
Opus:    ~4-7 seconds

Medium Task (2000 input, 1000 output):

Haiku:   ~4-6 seconds
Sonnet:  ~6-10 seconds
Opus:    ~8-15 seconds

Complex Task (5000 input, 3000 output):

Haiku:   ~8-12 seconds
Sonnet:  ~12-20 seconds
Opus:    ~15-30 seconds

With Extended Thinking (5000 input, 3000 output, 5k thinking budget):

Haiku:   ~15-20 seconds
Sonnet:  ~20-30 seconds
Opus:    ~30-45 seconds

Real talk: If you need results in under 5 seconds, Haiku is non-negotiable. Sonnet if you can tolerate 10 seconds. Opus only if the quality gap justifies the wait.

What Makes Models Different

Here’s the real talk about why models differ. Opus has roughly 10x more parameters than Haiku—think of parameters as the “knowledge slots” the model can use. More slots means more sophisticated reasoning. But bigger also means slower and more expensive. Sonnet splits the difference.

During training, Opus got explicit emphasis on reasoning and edge cases. It saw more diverse training data and spent proportionally more compute on hard problems. Haiku was optimized for common, straightforward tasks. This is why Opus excels at architectural thinking and Haiku excels at quick generation.

The key architectural difference: larger models have deeper attention layers. Opus has ~32 layers, Sonnet has ~20, Haiku has ~13. Attention layers are how the model tracks relationships between distant parts of your code or prompt. A complex refactoring task with 5000 tokens is easier for Opus to reason about than Haiku, even if Haiku could technically process all the tokens.

The practical implication: there’s no “best” model in the abstract. There’s only “best for this specific task at this specific moment.” The art is matching the capability profile to what you actually need.

Real-World Use Cases: Case Studies

Case Study 1: Startup Bootstrapping a Backend

Scenario: Early-stage startup, 3 developers, $0 budget for AI tools (free tier).

Model Strategy:

  • Haiku for 70% of tasks (feature scaffolding, simple bugs, test writing)
  • Sonnet for 30% of tasks (complex logic, debugging bottlenecks)
  • Opus never (not in the budget)

Result: Shipped features 40% faster than without Claude Code. Stayed within free tier the whole time. Quality was solid because the team knew when Haiku wasn’t enough and escalated to Sonnet.

Lesson: Haiku is underrated for fast-moving teams. Know its limits and escalate intelligently.

Case Study 2: Enterprise Code Audit

Scenario: Large bank needs security review of 200,000-line trading system.

Model Strategy:

  • Opus for architectural review (identify weak points)
  • Sonnet for targeted review (each identified module)
  • Haiku for compliance checklist validation

Result: Found 7 critical vulnerabilities that manual review missed. Cost: ~$200. Value: prevented a potential breach.

Lesson: Opus earns its cost on high-stakes work. Parallelize cheaper models where you can.

Case Study 3: Legacy Codebase Modernization

Scenario: 15-year-old PHP app, needs migration to Python FastAPI. 200 endpoints.

Model Strategy:

  • Haiku for endpoint-by-endpoint translation (bulk task)
  • Sonnet for refactoring output (quality pass)
  • Opus for architectural decisions (how to structure new system)

Result: 3-month project instead of 12 months. Fewer bugs because Sonnet caught issues Haiku missed.

Lesson: Use model strength matching. Haiku for mechanical work, Sonnet for quality, Opus for thinking.

Wrapping Up

Here’s the real talk: Sonnet should be your default. It’s fast, capable, and a sane middle ground. It handles 95% of real work without hesitation.

Use Opus for the genuinely hard stuff—architecture reviews, complex refactoring, novel problem-solving. Use Haiku for everything else when speed matters or you’re running in bulk. Use extended thinking for problems that benefit from deeper reasoning. Use fast mode when you want answers now over perfect answers.

The smartest engineers I know? They’re ruthless about using Haiku for 80% of work and reserving Opus for the 20% that actually needs it.

Start with Sonnet as your default. Run a few tasks. When you hit friction, switch up—Haiku if you need speed, Opus if Sonnet’s struggling. After a week, you’ll have intuition for which tool fits which job.

The art of model selection is knowing when you’re overthinking it. Most decisions are clear once you understand the trade-offs. Go pick the right brain for your task. Your token budget—and your deadlines—will thank you.

Monitoring and Adjusting Over Time

Set a monthly review cadence. Check your actual token usage:

claude code --show-usage

# Output: Claude Code estimated monthly cost: $47.80
#   Opus:   30% ($14.34)
#   Sonnet: 50% ($23.90)
#   Haiku:  20% ($9.56)

Ask yourself: Is Opus usage justified? Should it be 10%? Is Haiku usage enough—can we shift more work there? Are we getting good quality from our choices? Adjust your default_model accordingly based on results.

Real Patterns from Production Teams

After talking to dozens of teams using Claude Code in production, patterns emerge:

Pattern 1: Haiku for Exploration, Sonnet for Production

  • Start every task with Haiku to explore the problem
  • If Haiku struggles, escalate to Sonnet
  • Move to production only if genuinely simple
  • Result: 30% faster iteration, higher quality production code

Pattern 2: Opus on Mondays, Sonnet Rest of Week

  • Monday architecture sessions with Opus (weekly planning)
  • Tuesday-Friday execution with Sonnet
  • Saves ~40% on token costs while maintaining quality
  • Distributes expensive thinking upfront

Pattern 3: Dual-Model Validation

  • Write code with Sonnet
  • Validate with Opus (check for subtle issues)
  • Only costs the validation phase as Opus
  • Catches bugs Sonnet misses without paying Opus rates everywhere

Model Selection for Different Team Sizes

The dynamics change depending on your team. Here’s how to think about it:

Solo Developer

You’re maximizing your own productivity. Spend time benchmarking different models on your most common tasks. You probably want Sonnet as default with strategic Opus escalations. Your time is valuable; don’t waste it on slow feedback loops just to save a few cents.

Small Team (2-5 developers)

Establish conventions. Pick a default model for daily work and document when to escalate. A simple model policy document saves everyone from second-guessing their choices:

## Model Policy

- Default: Sonnet (daily work)
- Escalate to Opus: Architecture decisions, security reviews
- Use Haiku: Bulk operations, tests, boilerplate
- Fast mode: Never in production paths, OK for exploration

Medium Team (6-20 developers)

You need structure. Implement per-project model settings. Different projects have different needs—a legacy modernization might default to Haiku + Sonnet for QA, while greenfield system design needs Opus.

Large Team (20+ developers)

You need cost governance. Set up token budgets per team and track usage. Most teams benefit from tiered approach: Haiku/Sonnet by default, Opus on-demand with lightweight approval.

The Hidden Cost: Context Window Exhaustion

There’s a cost nobody talks about: running out of context. If you use Opus for everything, you’ll hit the 200K context limit on complex conversations. Haiku and Sonnet give you more room to grow.

Real example:

  • Scenario: Iterating on a complex system design with extended thinking
  • Using Opus: Extended thinking uses 15K tokens, your question uses 5K, response uses 8K. You’re at 28K already. Do this 6-7 times and you’re context-exhausted
  • Using Sonnet: Same structure but better pacing. You can iterate 10+ times

This compounds when working with large codebases. Sonnet leaves you more breathing room.

Seasonal Adjustments

Model selection isn’t static. Adjust for context:

Crunch Time (deadline in days)

  • Upgrade to Opus across the board
  • Speed of insight matters more than cost
  • Extended thinking on complex decisions
  • Fast mode for brainstorming

Normal Development

  • Sonnet default
  • Opus on escalation
  • Haiku for bulk work

Post-Release (Optimization Phase)

  • Haiku for refactoring
  • Sonnet for review
  • Opus for architectural improvements

-iNet

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.