Here’s the thing nobody tells you when you start building with Claude: the model itself can’t call your APIs. Not directly. Not from the chat interface, not from some magic plugin system that silently fires HTTP requests in the background.
Claude doesn’t have network access. It doesn’t have a fetch function. It can’t curl your weather service or hit your Stripe endpoint. And that’s actually by design—you don’t want an AI making arbitrary network calls without your knowledge. What you do want is a clean architecture where Claude decides what to call, and your infrastructure handles the how.
That’s what this guide is about. We’re going to build the bridge between Claude’s reasoning capabilities and the outside world. By the end, you’ll have real patterns for wrapping external APIs as tools, handling auth, transforming responses, managing errors, and keeping costs under control. Let’s get into it.
The Architecture You Actually Need
Before we write a single line of code, you need to understand the two primary integration patterns. Pick the wrong one and you’ll fight your own architecture for the life of the project.
Pattern 1: Tool Use via the Anthropic API
You define tools in your API request. Claude sees the tool definitions, decides when to use them, and returns a tool_use content block telling your application what to call and with what arguments. Your application executes the actual API call, then feeds the result back to Claude as a tool_result. Claude incorporates that data and continues reasoning.
This is the pattern for custom applications—backends, CLI tools, chatbots, agents. You own the execution loop.
Pattern 2: Model Context Protocol (MCP) Servers
MCP is a standardized protocol where you build a server that exposes external APIs as tools. Claude Desktop, Claude Code, and other MCP-compatible clients discover these tools automatically. The MCP server handles the actual API calls.
This is the pattern for extending Claude’s capabilities in existing clients without building a full application.
Which One Do You Pick?
If you’re building a product, an agent, or a custom backend: Tool Use via the API. You get full control over the execution loop, error handling, retries, and auth management.
If you’re augmenting your personal Claude workflow or building internal tools for a team: MCP. You write the server once, and it works across any MCP-compatible client.
Both patterns share the same core concept: Claude never calls APIs directly. You build the bridge. Let’s see how.
Tool Use: Wrapping External APIs
This is where most production integrations live. You define a tool schema that describes what the external API does, Claude decides when to invoke it, and your code handles the actual HTTP request.
Defining a Tool That Wraps a Weather API
Here’s a complete example wrapping the OpenWeatherMap API. Pay attention to the schema—Claude uses the description fields to decide when and how to use the tool, so they need to be precise.
client = anthropic.Anthropic()
# Define the tool schema — this tells Claude what's available
tools = [
{
"name": "get_weather",
"description": "Get current weather conditions for a specific city. Returns temperature, humidity, wind speed, and a text description. Use this when the user asks about current weather anywhere in the world.",
"input_schema": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. 'London' or 'San Francisco'"
},
"units": {
"type": "string",
"enum": ["metric", "imperial"],
"description": "Temperature units. Use 'imperial' for US users, 'metric' for everyone else."
}
},
"required": ["city"]
}
}
]
def call_weather_api(city: str, units: str = "metric") -> dict:
"""Actually hit the OpenWeatherMap API."""
api_key = os.environ["OPENWEATHER_API_KEY"]
resp = httpx.get(
"https://api.openweathermap.org/data/2.5/weather",
params={"q": city, "appid": api_key, "units": units}
)
resp.raise_for_status()
data = resp.json()
# Transform the response — give Claude only what it needs
return {
"city": data["name"],
"country": data["sys"]["country"],
"temperature": data["main"]["temp"],
"feels_like": data["main"]["feels_like"],
"humidity": data["main"]["humidity"],
"description": data["weather"][0]["description"],
"wind_speed": data["wind"]["speed"],
"units": units
}
Notice what we’re doing in call_weather_api: we’re not dumping the entire OpenWeatherMap response into Claude’s context. We’re transforming it. The raw response includes icon codes, internal IDs, coordinate data, and a bunch of other fields Claude doesn’t need. Every token you feed Claude costs money and burns context window. Be surgical about what you pass back.
The Execution Loop
Now here’s the part that trips people up. You don’t just send the message and get a final answer. Tool use requires a conversation loop: send the message, check if Claude wants to use a tool, execute the tool, feed the result back, and let Claude generate the final response.
def chat_with_tools(user_message: str) -> str:
"""Complete tool-use loop with Claude."""
messages = [{"role": "user", "content": user_message}]
# First API call — Claude may request a tool
response = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
tools=tools,
messages=messages
)
# Loop until Claude gives a final text response
while response.stop_reason == "tool_use":
# Find the tool use block
tool_block = next(
b for b in response.content if b.type == "tool_use"
)
tool_name = tool_block.name
tool_input = tool_block.input
# Execute the actual API call
if tool_name == "get_weather":
try:
result = call_weather_api(**tool_input)
tool_result = {
"type": "tool_result",
"tool_use_id": tool_block.id,
"content": str(result)
}
except httpx.HTTPStatusError as e:
tool_result = {
"type": "tool_result",
"tool_use_id": tool_block.id,
"content": f"Error: API returned {e.response.status_code}",
"is_error": True
}
# Feed the result back to Claude
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": [tool_result]})
response = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
tools=tools,
messages=messages
)
# Extract the final text response
return next(b.text for b in response.content if hasattr(b, "text"))
That while response.stop_reason == "tool_use" loop is the heart of it. Claude might chain multiple tool calls—check the weather in two cities, look up a database record, then call a calculation service. Your loop keeps running until Claude has everything it needs and produces a text response.
Authentication: The Part Everyone Gets Wrong
External APIs need auth. Your weather API needs an API key. Your internal services need bearer tokens. Your payment processor needs OAuth. Here’s how to handle each pattern without leaking credentials into Claude’s context.
API Keys
The simplest pattern. Keep keys in environment variables, never in tool definitions, never in system prompts.
# WRONG — don't put API keys anywhere Claude can see them
tools = [{
"name": "get_weather",
"description": "Get weather. API key is sk-abc123..." # NO. STOP.
}]
# RIGHT — keys live in your execution layer
def call_weather_api(city: str, units: str = "metric") -> dict:
api_key = os.environ["OPENWEATHER_API_KEY"] # Claude never sees this
resp = httpx.get(url, params={"q": city, "appid": api_key, "units": units})
return transform_response(resp.json())
Claude doesn’t need to know how auth works. It just describes what it wants—”get the weather in Tokyo”—and your execution layer handles the credentials. This is a fundamental boundary. Claude reasons about what to do. Your code handles how.
Bearer Tokens and OAuth
For APIs that require user-specific tokens (Slack, GitHub, Google), you’ll typically have an OAuth flow in your application that stores tokens per user. Your tool execution layer retrieves the right token at runtime:
def execute_tool(tool_name: str, tool_input: dict, user_id: str) -> dict:
"""Tool executor with per-user auth."""
# Retrieve the user's stored OAuth token
token = token_store.get(user_id, service=tool_name_to_service(tool_name))
if not token:
return {"error": "User hasn't connected this service yet"}
if token.is_expired:
token = refresh_oauth_token(token)
token_store.save(user_id, token)
headers = {"Authorization": f"Bearer {token.access_token}"}
# ... make the API call with proper auth
The key insight: auth is entirely your application’s responsibility. Claude doesn’t participate in OAuth flows, doesn’t store tokens, doesn’t refresh credentials. It just says “call this tool with these arguments” and trusts your infrastructure to handle the rest.
Error Handling: Because APIs Break
External APIs fail. They time out. They rate-limit you. They return 500s at the worst possible moment. Your tool execution layer needs to handle all of this gracefully, because if you just pass raw errors to Claude, it’ll either hallucinate a recovery or give the user a terrible experience.
The Resilient API Wrapper
Here’s a pattern I use in production for every external API integration. It handles retries, timeouts, rate limits, and gives Claude clean error information it can actually reason about.
from typing import Any
class ResilientAPIClient:
"""Wrapper for external API calls with retry logic and clean error handling."""
def __init__(self, base_url: str, api_key: str, max_retries: int = 3):
self.base_url = base_url
self.max_retries = max_retries
self.client = httpx.Client(
base_url=base_url,
headers={"Authorization": f"Bearer {api_key}"},
timeout=10.0 # Hard timeout — don't let Claude's user wait forever
)
def call(self, method: str, path: str, **kwargs) -> dict[str, Any]:
"""Make an API call with automatic retries and clean error surfaces."""
last_error = None
for attempt in range(self.max_retries):
try:
resp = self.client.request(method, path, **kwargs)
# Rate limited — back off and retry
if resp.status_code == 429:
retry_after = int(resp.headers.get("Retry-After", 2))
time.sleep(min(retry_after, 10)) # Cap at 10s
continue
resp.raise_for_status()
return {"success": True, "data": resp.json()}
except httpx.TimeoutException:
last_error = "The service is taking too long to respond"
time.sleep(2 ** attempt) # Exponential backoff
except httpx.HTTPStatusError as e:
status = e.response.status_code
if status >= 500:
last_error = "The service is experiencing issues"
time.sleep(2 ** attempt)
elif status == 404:
return {
"success": False,
"error": "Resource not found",
"details": "The requested item doesn't exist"
}
elif status == 403:
return {
"success": False,
"error": "Access denied",
"details": "Insufficient permissions for this operation"
}
else:
return {
"success": False,
"error": f"Request failed with status {status}",
"details": e.response.text[:200]
}
return {
"success": False,
"error": last_error or "Request failed after retries",
"retries_exhausted": True
}
There are a few things worth calling out here.
Clean error messages. We don’t pass raw stack traces to Claude. We give it structured error information—what happened, whether it’s retryable, what the user-facing implication is. Claude can then craft a helpful response like “I wasn’t able to check the weather right now—the service seems to be having issues. Want me to try again in a minute?” instead of dumping a traceback.
Capped retries and backoff. Three retries with exponential backoff handles transient failures without making the user wait forever. The 10-second cap on rate-limit waits prevents one slow API from blocking an entire conversation.
Structured returns. Every response from call() has a consistent shape—success boolean plus either data or error. This makes your tool execution layer predictable and easy to reason about.
Data Transformation: The Hidden Superpower
This is the part most tutorials skip, and it’s arguably the most important part of the whole integration. The data you get from external APIs is almost never in the right shape for Claude.
A Stripe payment intent response is 50+ fields. A GitHub issue has nested objects three levels deep. A database query might return 500 rows. You can’t dump all of that into Claude’s context and expect good results. You’ll burn tokens, confuse the model, and get worse answers.
Transform Before You Pass
The rule is simple: give Claude exactly the information it needs to answer the user’s question, and nothing more.
def transform_stripe_payment(raw: dict) -> dict:
"""Transform a Stripe PaymentIntent into what Claude actually needs."""
return {
"payment_id": raw["id"],
"amount": raw["amount"] / 100, # Stripe uses cents
"currency": raw["currency"].upper(),
"status": raw["status"],
"customer_email": raw.get("receipt_email", "unknown"),
"created": raw["created"], # Unix timestamp
"description": raw.get("description", "No description"),
"last_four": raw.get("payment_method_details", {})
.get("card", {})
.get("last4", "N/A")
}
# We're dropping: metadata, charges array, shipping details,
# payment method fingerprint, radar risk data, and 40+ other fields
# Claude doesn't need any of that to answer "what's the status of payment X?"
Summarize Large Result Sets
When an API returns a list of results, don’t send all of them. Summarize, paginate, or filter.
def transform_database_results(rows: list[dict], query_context: str) -> dict:
"""Summarize large result sets instead of dumping everything."""
if len(rows) <= 10:
return {"results": rows, "total": len(rows)}
return {
"total_results": len(rows),
"showing": "first 10",
"results": rows[:10],
"summary": {
"date_range": f"{rows[-1]['created_at']} to {rows[0]['created_at']}",
"unique_users": len(set(r["user_id"] for r in rows)),
},
"note": "More results available. Ask the user if they want to narrow the search."
}
That note field is a nice trick. You’re giving Claude a hint about what to do next—suggest the user refine their query instead of trying to process hundreds of records.
Rate Limiting and Cost Management
Here’s where you need to think about two things at once: the rate limits of your external APIs, and the token cost of feeding responses back to Claude.
External API Rate Limits
Most APIs have rate limits. If you’re wrapping a free-tier weather API that allows 60 calls per minute, you need to track that in your execution layer—not in Claude.
from collections import defaultdict
class RateLimiter:
def __init__(self):
self.windows: dict[str, list[float]] = defaultdict(list)
def check(self, api_name: str, max_calls: int, window_seconds: int) -> bool:
now = time.time()
calls = self.windows[api_name]
# Remove expired entries
self.windows[api_name] = [t for t in calls if now - t < window_seconds]
if len(self.windows[api_name]) >= max_calls:
return False
self.windows[api_name].append(now)
return True
rate_limiter = RateLimiter()
def execute_tool_with_limits(tool_name: str, tool_input: dict) -> dict:
limits = {"get_weather": (60, 60), "search_database": (100, 60)}
if tool_name in limits:
max_calls, window = limits[tool_name]
if not rate_limiter.check(tool_name, max_calls, window):
return {
"success": False,
"error": "Rate limit reached for this service",
"retry_after_seconds": window
}
return execute_tool(tool_name, tool_input)
Token Cost Awareness
Every character in a tool result costs tokens. A single bloated API response can cost more than the rest of the conversation combined. Here’s a rough guide:
- Keep tool results under 2,000 tokens (~1,500 words) per call
- If you’re returning structured data, use compact formats—skip unnecessary whitespace in JSON
- Strip HTML, markdown formatting, and boilerplate from API responses
- For large text content (articles, documents), summarize rather than include the full text
A Stripe webhook payload can be 4,000+ tokens raw. After transformation? Under 200 tokens. That’s a 20x cost reduction on every single tool call.
MCP: The Other Path
If you’re building an MCP server instead of a custom application, the concepts are identical—you still wrap APIs, transform data, handle errors. The difference is the transport layer and discovery mechanism.
Here’s what a minimal MCP server looks like for the same weather API:
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("weather-service")
@mcp.tool()
async def get_weather(city: str, units: str = "metric") -> str:
"""Get current weather conditions for a city.
Args:
city: City name like 'London' or 'Tokyo'
units: 'metric' for Celsius, 'imperial' for Fahrenheit
"""
# Same API call, same transformation, same error handling
try:
result = await call_weather_api(city, units)
return f"Weather in {result['city']}: {result['temperature']}°, " \
f"{result['description']}, humidity {result['humidity']}%"
except Exception as e:
return f"Could not fetch weather: {str(e)}"
if __name__ == "__main__":
mcp.run()
Same patterns, different packaging. The MCP framework handles the protocol details—discovery, transport, message framing. You focus on wrapping the API and transforming the data.
Real-World Integration: Putting It All Together
Let’s walk through a realistic scenario. You’re building a customer support agent that needs to:
- Look up customer information from your database
- Check their recent orders via your order management API
- Process refunds through Stripe
That’s three external APIs, three tool definitions, three sets of auth, and a conversation loop that might chain all three in a single interaction.
The architecture looks like this:
User asks: “I need a refund for order #4521”
Claude’s reasoning chain:
- “I need to look up order #4521” → calls
get_ordertool - “The order is $47.99 for user [email protected], delivered 3 days ago” → reasons about refund eligibility
- “This is within the refund window, let me process it” → calls
process_refundtool - “Refund of $47.99 initiated, confirmation ID ref_abc123” → generates response
Your application handles all three API calls, manages the Stripe API key, validates the refund amount, and feeds clean results back to Claude. Claude never touches a network socket. It reasons about what to do, and your infrastructure executes it.
The separation is clean, auditable, and secure. You can log every tool call, rate-limit by user, enforce business rules before executing destructive actions (like refunds), and swap out API providers without changing Claude’s tool definitions.
The Mistakes That’ll Bite You
Before we wrap up, let me save you some pain. These are the mistakes I see most often in production Claude + API integrations:
Putting API keys in system prompts. Don’t do this. Claude doesn’t need your keys. Your execution layer does. Keep them in environment variables or a secrets manager.
Dumping raw API responses. Transform everything. Strip what Claude doesn’t need. Your token costs and response quality will both improve dramatically.
No error handling in tool execution. When an API fails and you return nothing, Claude hallucinates an answer. When you return a clean error message, Claude tells the user what happened and suggests alternatives.
Ignoring the loop. Tool use isn’t a single request-response. It’s a loop. Claude might call multiple tools, or the same tool multiple times. Your code needs to handle that.
Over-broad tool definitions. A tool called “do_anything” with a freeform input schema is useless. Claude won’t know when to use it. Be specific in your names, descriptions, and schemas.
Skipping rate limits. Your free-tier API lets you make 60 calls per minute. A chatty agent can blow through that in two conversations. Track and limit at the execution layer.
Where to Go From Here
You’ve got the patterns. Here’s how to level up from here:
Multi-tool orchestration. Define 5-10 tools and let Claude chain them. You’ll be surprised at how well it reasons about tool ordering and data dependencies.
Streaming with tool use. The Anthropic API supports streaming responses even during tool-use loops. This gives your users real-time feedback while Claude is thinking and making API calls.
Caching tool results. If your weather API data is valid for 10 minutes, cache it. If a database query returns the same result within a session, cache it. Claude doesn’t need to know you’re caching—it gets the same result either way, and you save money on both API calls and tokens.
Guardrails and business logic. Wrap destructive tools (delete, refund, cancel) with confirmation steps. Your execution layer can return “This will refund $247.99 to the customer. Should I proceed?” and let Claude ask the user for confirmation before executing.
The core pattern never changes: Claude reasons, you execute. Keep that boundary clean and you can integrate with anything—weather APIs, databases, payment processors, internal microservices, IoT devices, whatever your architecture demands. The bridge is yours to build.