Here’s a dirty secret about building with AI: a single prompt, no matter how carefully crafted, hits a ceiling. You can tune it, refine it, add examples, chain-of-thought your way into oblivion—and eventually you’ll still be fighting the same fundamental constraint. One model call, one context window, one perspective.
Multi-agent workflows blow that ceiling off. Instead of cramming everything into a single monolithic prompt, you decompose work across multiple specialized agents, each with its own context, instructions, and focus. One agent plans. Another writes code. A third reviews it. A fourth validates the output against your test suite. They pass context between each other, check each other’s work, and produce results that no single call could match.
But here’s the thing most tutorials won’t tell you: the hard part isn’t the agents. It’s everything between them. The orchestration layer—how agents communicate, how failures cascade, how you prevent context from exploding into an unmanageable mess—that’s where multi-agent systems live or die.
Let’s get into the architecture patterns, the implementation details, and the hidden failure modes that’ll bite you if you don’t see them coming.
Why Multi-Agent Instead of One Big Prompt
Before we build anything, let’s be honest about when you actually need multi-agent workflows. If your task fits comfortably in a single context window and doesn’t require specialized perspectives, a well-crafted prompt is simpler, cheaper, and faster. Don’t architect a distributed system when a function call will do.
You need multi-agent when:
- The task exceeds a single context window. Analyzing a 200-file codebase, processing a 500-page document, or generating a 50,000-word manuscript. No single call handles that.
- The task requires competing perspectives. Code generation and code review are fundamentally different cognitive modes. Asking one agent to do both is like asking a writer to proofread their own work—they’ll miss the same mistakes every time.
- Quality gates matter. If you need validation checkpoints between stages, you need separate agents to enforce them. An agent can’t objectively evaluate its own output.
- You need parallelism. Processing 100 documents simultaneously. Running security analysis, performance analysis, and code review concurrently. A single agent processes sequentially.
- Different tasks need different cost profiles. Complex architectural decisions need Opus. Reformatting boilerplate needs Haiku. Orchestrating the pipeline needs Sonnet. Model mixing saves you 60-80% on costs without sacrificing quality where it matters.
If two or more of those apply to your use case, you’re in multi-agent territory. Let’s talk about how to structure it.
The Four Architecture Patterns
Every multi-agent system I’ve built or seen falls into one of four patterns—or a hybrid of them. Understanding these patterns is the difference between a system that scales and one that collapses under its own complexity.
1. The Orchestrator Pattern
One central agent—the orchestrator—decomposes the task, dispatches subtasks to worker agents, collects results, and synthesizes the final output. The workers don’t talk to each other. All communication flows through the orchestrator.
This is your workhorse pattern. It handles 70% of multi-agent use cases. The orchestrator maintains the big picture while workers focus on narrow, well-defined subtasks.
When to use it: General-purpose task decomposition, research tasks, code generation with review, content pipelines.
The trap: The orchestrator’s context window becomes the bottleneck. If you’re aggregating results from 20 worker agents, the orchestrator needs to hold all those results simultaneously. Design your workers to return concise, structured outputs—not raw prose dumps.
2. The Pipeline Pattern
Agents execute sequentially, each taking the output of the previous agent as input. Agent A produces a plan. Agent B takes that plan and writes code. Agent C takes that code and reviews it. Agent D takes the review and fixes the issues. Linear. Predictable. Easy to debug.
When to use it: Content pipelines (outline, draft, edit, format), code pipelines (plan, implement, test, review), any workflow with clear sequential stages.
The trap: Pipeline latency is the sum of all stages. If each agent takes 10 seconds and you have 6 stages, every run takes a minute. Identify which stages can run in parallel and break them out.
3. The Parallel Fan-Out Pattern
One agent splits a task into independent subtasks, dispatches them all simultaneously, and another agent aggregates the results. Think MapReduce for AI. Process 50 documents in parallel, then merge the summaries.
When to use it: Batch processing, document analysis at scale, running multiple analysis perspectives simultaneously, any task where subtasks are independent.
The trap: Aggregation is harder than it looks. 50 parallel agents produce 50 outputs that might overlap, contradict, or use inconsistent formatting. Your aggregator agent needs explicit instructions for deduplication, conflict resolution, and normalization.
4. The Hierarchical Pattern
A tree of agents. A top-level orchestrator dispatches to mid-level coordinators, which dispatch to leaf-level workers. Military command structure. You use this when the task is too complex for a single orchestrator to manage but too interconnected for pure parallelism.
When to use it: Large-scale systems (analyzing an entire monorepo, generating a full application), multi-domain tasks that require domain-specific coordination.
The trap: Communication overhead grows exponentially with depth. Keep your hierarchy to 2-3 levels max. If you need more, your task decomposition is probably wrong.
Agent Communication: The Context Handoff Problem
Now we get to the part that kills most multi-agent systems. How do agents actually talk to each other?
The naive approach is to pass the entire context from one agent to the next. Agent A dumps its full output into Agent B’s prompt. This works for two agents. It does not work for six. By agent four, you’ve blown your context window with accumulated garbage from agents one through three—most of which agent four doesn’t need.
Here’s the principle: each agent should receive the minimum context it needs to do its job, and nothing more.
That means you need a context protocol. Every agent produces a structured output with two sections: a result (concise, structured, what the next agent needs) and a trace (detailed reasoning, stored separately for debugging). Only the result gets passed forward. The trace gets logged.
In practice, this looks like:
client = anthropic.Anthropic()
def call_agent(system_prompt: str, user_message: str, model: str = "claude-sonnet-4-20250514") -> dict:
"""Call a Claude agent and return structured output."""
response = client.messages.create(
model=model,
max_tokens=4096,
system=system_prompt,
messages=[{"role": "user", "content": user_message}]
)
return json.loads(response.content[0].text)
def orchestrator_pipeline(task: str) -> dict:
"""Orchestrator pattern: plan -> execute -> validate."""
# Step 1: Planner agent decomposes the task (Opus for complex reasoning)
plan = call_agent(
system_prompt="""You are a task planner. Decompose the given task into
discrete subtasks. Return JSON:
{
"subtasks": [
{"id": "1", "description": "...", "dependencies": [], "model_tier": "haiku|sonnet|opus"},
],
"execution_order": [["1", "2"], ["3"]], // groups run in parallel
"context_requirements": {"1": ["input"], "3": ["1", "2"]}
}""",
user_message=task,
model="claude-opus-4-20250514"
)
results = {}
# Step 2: Execute subtasks respecting dependencies and parallelism
for parallel_group in plan["execution_order"]:
# In production, run these concurrently with asyncio/threading
for subtask_id in parallel_group:
subtask = next(s for s in plan["subtasks"] if s["id"] == subtask_id)
# Build context from dependencies only—not entire history
context_ids = plan["context_requirements"].get(subtask_id, [])
context = {cid: results[cid]["result"] for cid in context_ids if cid in results}
# Model mixing: use the tier the planner recommended
model_map = {
"haiku": "claude-haiku-4-20250514",
"sonnet": "claude-sonnet-4-20250514",
"opus": "claude-opus-4-20250514"
}
result = call_agent(
system_prompt=f"""You are a specialist agent. Complete this subtask.
Return JSON: {{"result": "concise output", "confidence": 0.0-1.0, "trace": "reasoning"}}""",
user_message=json.dumps({
"subtask": subtask["description"],
"context": context
}),
model=model_map[subtask["model_tier"]]
)
results[subtask_id] = result
# Step 3: Validator agent checks combined results (Opus for judgment)
validation = call_agent(
system_prompt="""You are a quality validator. Check these results for
correctness, completeness, and consistency. Return JSON:
{"passed": true/false, "issues": [...], "final_output": "..."}""",
user_message=json.dumps({k: v["result"] for k, v in results.items()}),
model="claude-opus-4-20250514"
)
return validation
A few things to notice here. The planner decides which model tier each subtask needs—that’s your cost optimization happening at the planning stage. Context is assembled per-agent from explicit dependency declarations, not by dumping everything forward. And the validator is a separate agent from the workers, so it can objectively evaluate their output.
The Pipeline Pattern in Practice
Pipelines are the most intuitive pattern, but they need quality gates between stages to prevent garbage from propagating. Here’s a content pipeline with built-in validation:
client = anthropic.Anthropic()
def content_pipeline(topic: str, style_guide: str) -> dict:
"""Pipeline pattern with quality gates between stages."""
# Stage 1: Research and outline (Opus — needs deep reasoning)
outline = call_agent(
system_prompt=f"""You are a research and outlining agent.
Given a topic, produce a detailed content outline.
Style guide: {style_guide}
Return JSON: {{
"title": "...",
"sections": [{{"heading": "...", "key_points": [...], "word_target": 300}}],
"total_word_target": 2500,
"sources_needed": [...]
}}""",
user_message=topic,
model="claude-opus-4-20250514"
)
# QUALITY GATE 1: Outline validation
gate_1 = call_agent(
system_prompt="""You are an outline validator. Check:
1. Does the outline cover the topic comprehensively?
2. Is the structure logical?
3. Are word targets realistic?
4. Are there any gaps or redundancies?
Return JSON: {"pass": true/false, "issues": [...], "suggestions": [...]}""",
user_message=json.dumps(outline),
model="claude-haiku-4-20250514" # Simple validation = cheap model
)
if not gate_1["pass"]:
# Feed issues back to outliner for revision
outline = call_agent(
system_prompt="You are a research and outlining agent. Revise this outline.",
user_message=json.dumps({"original": outline, "issues": gate_1["issues"]}),
model="claude-opus-4-20250514"
)
# Stage 2: Draft each section in parallel (Sonnet — balanced quality/speed)
sections = []
for section in outline["sections"]:
draft = call_agent(
system_prompt=f"""You are a prose writer. Write this section following
the style guide. Target {section['word_target']} words.
Style: {style_guide}
Return JSON: {{"content": "markdown text", "word_count": 0}}""",
user_message=json.dumps({
"heading": section["heading"],
"key_points": section["key_points"],
"article_context": outline["title"]
}),
model="claude-sonnet-4-20250514"
)
sections.append(draft)
# QUALITY GATE 2: Style and consistency check
combined_draft = "\n\n".join(s["content"] for s in sections)
gate_2 = call_agent(
system_prompt=f"""You are a style enforcer. Check this draft against
the style guide: {style_guide}
Check for: voice consistency, jargon level, paragraph length, flow between sections.
Return JSON: {{
"pass": true/false,
"style_violations": [...],
"continuity_issues": [...]
}}""",
user_message=combined_draft,
model="claude-sonnet-4-20250514"
)
# Stage 3: Final edit pass (Opus — needs holistic judgment)
final = call_agent(
system_prompt=f"""You are a final editor. Take this draft and style feedback,
produce the final polished article. Fix all noted issues.
Return the complete article as markdown text (not JSON).""",
user_message=json.dumps({
"draft": combined_draft,
"style_issues": gate_2.get("style_violations", []),
"continuity_issues": gate_2.get("continuity_issues", [])
}),
model="claude-opus-4-20250514"
)
return {"outline": outline, "gates": [gate_1, gate_2], "final_article": final}
Notice the quality gates. Gate 1 uses Haiku because outline validation is structurally simple—check if sections exist, check word targets, check for gaps. Gate 2 uses Sonnet because style enforcement requires more nuance. The final edit uses Opus because holistic judgment over a full article is genuinely complex reasoning. That’s model mixing in action. You’re spending Opus tokens only where Opus-level reasoning is required.
Model Mixing: The Cost Engineering Play
Let’s talk numbers. As of early 2026, the cost difference between Claude model tiers is substantial. Opus costs roughly 10-15x what Haiku costs per token. If you’re running hundreds of agent calls per pipeline, using Opus for everything will bankrupt your API budget before you ship.
The strategy is straightforward. Map each agent role to a cognitive complexity level:
Opus (complex reasoning):
- Task decomposition and planning
- Architectural decisions
- Final quality judgment
- Handling ambiguity and edge cases
- Creative synthesis across multiple inputs
Sonnet (balanced):
- Orchestration and coordination
- Prose generation
- Code generation
- Moderate analysis tasks
- Style enforcement
Haiku (simple/fast):
- Structural validation (does this JSON have required fields?)
- Classification and routing
- Formatting and normalization
- Simple extraction tasks
- Status checks and gate evaluations
A well-mixed pipeline might use Opus for 10-15% of calls, Sonnet for 40-50%, and Haiku for the rest. That cuts your API costs by 60-70% compared to Opus-everywhere, with negligible quality impact—because Haiku is genuinely excellent at simple tasks. You’re not sacrificing quality; you’re allocating it correctly.
Scaling: From 5 Agents to 500
Running five agents sequentially is trivial. Running 500 in parallel is an infrastructure problem. Here’s what changes at scale.
Rate Limits
Anthropic’s API has rate limits on both requests per minute and tokens per minute. When you fan out to 100 parallel agents, you’ll hit those limits immediately. You need:
- A request queue with configurable concurrency. Start with 10-20 concurrent requests and adjust based on your tier.
- Exponential backoff with jitter on 429 responses. Don’t retry immediately—you’ll just hit the limit harder.
- Token budgeting. Before dispatching 100 agents, estimate total token consumption. If it exceeds your per-minute limit, batch the dispatches across time windows.
Context Window Management
The number one scaling failure mode is context explosion. Each agent accumulates context. If downstream agents inherit upstream context, your token usage grows quadratically with pipeline depth.
Solutions:
- Summarization agents. Between pipeline stages, insert a summarizer that compresses the previous stage’s output to 20% of its size. Use Haiku for this—it’s fast and cheap.
- Context registries. Instead of passing context directly, store results in a shared registry (a dictionary, a database, a file system). Agents receive registry keys, not full content. They pull only what they need.
- Sliding windows. For long-running pipelines, only pass the last N results forward. Archive the rest. If a downstream agent needs historical context, it can request specific items from the registry.
Failure Handling
In a single-agent system, failure is binary: it worked or it didn’t. In a multi-agent system, failure is fractal. Agent 3 out of 20 fails. Do you retry it? Do you restart the pipeline? Do you skip it and let downstream agents work with partial results?
Design your failure modes explicitly:
- Retry with backoff. For transient failures (rate limits, timeouts), retry 2-3 times with exponential backoff.
- Fallback models. If Opus times out, fall back to Sonnet. The output might be lower quality, but it’s better than no output.
- Graceful degradation. If a non-critical agent fails (say, the “add humor” agent in a content pipeline), mark it as skipped and continue. Don’t let a comedy agent take down your entire pipeline.
- Circuit breakers. If more than 30% of parallel agents fail, stop dispatching and surface the error. Something systemic is wrong—maybe the API is down, maybe your prompts are malformed.
The Hidden Layer: Orchestration Design
Here’s the insight that separates toy demos from production systems: design your orchestration layer first, agents second.
Most people start by building individual agents—a code writer, a code reviewer, a test generator. They get each one working in isolation, then try to wire them together. This fails. The agents were designed to work alone. Their inputs and outputs don’t match. Their context assumptions conflict. Their error formats are incompatible.
Start with the orchestration contract instead:
from dataclasses import dataclass, field
from typing import Optional
from enum import Enum
class AgentStatus(Enum):
PENDING = "pending"
RUNNING = "running"
SUCCESS = "success"
FAILED = "failed"
SKIPPED = "skipped"
@dataclass
class AgentResult:
"""Universal output contract — every agent returns this shape."""
agent_id: str
status: AgentStatus
result: dict # Structured output — what downstream agents consume
trace: str # Reasoning log — stored for debugging, never forwarded
confidence: float # 0.0-1.0 — used by quality gates
token_usage: dict # {"input": N, "output": N} — for cost tracking
errors: list = field(default_factory=list)
@dataclass
class AgentTask:
"""Universal input contract — every agent receives this shape."""
task_id: str
description: str
context: dict # Only the data this agent needs
model: str = "claude-sonnet-4-20250514" # Default to balanced model
max_retries: int = 2
timeout_ms: int = 30000
fallback_model: Optional[str] = None # Cheaper model if primary fails
quality_threshold: float = 0.7 # Minimum confidence to accept result
class Orchestrator:
"""The control plane — manages agent lifecycle, not agent logic."""
def __init__(self):
self.results: dict[str, AgentResult] = {}
self.cost_tracker = {"input_tokens": 0, "output_tokens": 0}
def dispatch(self, task: AgentTask) -> AgentResult:
"""Send task to agent, handle retries, fallbacks, and tracking."""
for attempt in range(task.max_retries + 1):
try:
model = task.model if attempt == 0 else (task.fallback_model or task.model)
result = self._call_model(task, model)
if result.confidence < task.quality_threshold:
if attempt < task.max_retries:
continue # Retry — confidence too low
result.status = AgentStatus.FAILED
result.errors.append(f"Confidence {result.confidence} below threshold {task.quality_threshold}")
self.results[task.task_id] = result
self._track_cost(result)
return result
except Exception as e:
if attempt == task.max_retries:
return AgentResult(
agent_id=task.task_id,
status=AgentStatus.FAILED,
result={},
trace=str(e),
confidence=0.0,
token_usage={"input": 0, "output": 0},
errors=[str(e)]
)
def get_context_for(self, task_id: str, dependency_ids: list[str]) -> dict:
"""Pull only relevant results for a downstream agent."""
return {
dep_id: self.results[dep_id].result
for dep_id in dependency_ids
if dep_id in self.results and self.results[dep_id].status == AgentStatus.SUCCESS
}
def _call_model(self, task: AgentTask, model: str) -> AgentResult:
"""Actual model call — implement with your preferred client."""
# Implementation here
...
def _track_cost(self, result: AgentResult):
self.cost_tracker["input_tokens"] += result.token_usage.get("input", 0)
self.cost_tracker["output_tokens"] += result.token_usage.get("output", 0)
Every agent consumes an AgentTask and produces an AgentResult. Period. No exceptions. This contract is the foundation. Once you have it, individual agents become interchangeable components. You can swap models, add retry logic, insert quality gates, and parallelize dispatches—all at the orchestration level, without touching agent-specific code.
The get_context_for method is critical. It enforces the minimum-context principle. Instead of forwarding everything, each agent declares its dependencies, and the orchestrator assembles only those results into the context payload. Agent 5 doesn’t get buried under the output of agents 1-4 when it only needs agent 3’s results.
Quality Gates: The Trust-But-Verify Layer
Quality gates are the difference between “we generated some output” and “we generated output we can trust.” Here’s the pattern:
Between every pipeline stage, insert a lightweight validation agent. This agent has one job: determine whether the previous stage’s output meets the bar for the next stage to consume it. If it doesn’t, the gate either triggers a retry or escalates to a human.
The key insight about quality gates is model selection. Gates should almost always use a cheaper model than the stage they’re validating. Why? Validation is cognitively simpler than generation. Checking if code compiles is easier than writing the code. Checking if an outline covers all topics is easier than creating the outline. Use Haiku for structural validation. Use Sonnet for semantic validation. Save Opus for final holistic judgment.
Gate failures should be informative. Don’t just return pass/fail. Return the specific issues found, so the retry agent can address them directly instead of regenerating from scratch. A gate that says “the code doesn’t handle the error case on line 47” is infinitely more useful than one that says “the code has issues.”
Common Failure Modes and How to Fix Them
Let me save you some debugging hours. These are the failure modes I’ve seen kill multi-agent systems:
Context window overflow. Your orchestrator accumulates results from all agents and eventually blows its context limit. Fix: summarize aggressively between stages. Store full results externally. Only pass forward what the next agent actually needs.
Infinite retry loops. An agent consistently fails quality gates, gets retried, produces the same bad output, fails again. Fix: limit retries to 2-3 attempts. After that, either escalate to a higher-capability model, fall back to a simpler approach, or surface the failure to a human.
Format drift. Agent outputs gradually deviate from the expected schema as the pipeline runs. One agent returns a string where you expected an array. Another nests data one level deeper than expected. Fix: validate output schemas at every gate. Use typed contracts. Parse and validate JSON strictly.
Orchestrator bottleneck. All communication flows through one orchestrator whose context window becomes the constraint. Fix: use hierarchical orchestration. Split into domain-specific sub-orchestrators that only report summaries up to the top level.
Cost explosion. You’re using Opus for everything because “quality matters.” Your API bill is five figures and climbing. Fix: model mixing. Profile which agents actually need Opus-level reasoning. Most don’t. The Haiku/Sonnet/Opus split should be roughly 40/45/15 by call volume.
Putting It All Together
Multi-agent workflows with Claude are powerful, but they’re not magic. They’re engineering. The agents themselves are the easy part—they’re just prompt-model pairs with structured outputs. The hard part is everything else: how they communicate, how context flows, how failures propagate, how costs scale.
Start with the orchestration contract. Define your AgentTask and AgentResult shapes before you write a single agent prompt. Build quality gates as first-class pipeline components, not afterthoughts. Mix models aggressively—Haiku for validation, Sonnet for generation, Opus for judgment. Manage context like a precious resource, because it is.
And here’s the final hidden layer lesson: monitor everything. Log every agent call, every context payload, every gate decision, every retry. When your multi-agent pipeline produces garbage—and it will, at some point—those logs are the only way you’ll figure out where the chain broke. In a single-agent system, debugging is “read the prompt, read the output.” In a multi-agent system, debugging is “trace the context chain across six agents and find where the semantic corruption entered.”
Build the observability in from day one. Your future self will thank you.