All Articles Claude AI

How to Combine Claude with External APIs

Here's the thing nobody tells you when you start building with Claude: the model itself can't call your APIs. Not directly.

Here’s the thing nobody tells you when you start building with Claude: the model itself can’t call your APIs. Not directly. Not from the chat interface, not from some magic plugin system that silently fires HTTP requests in the background.

Claude doesn’t have network access. It doesn’t have a fetch function. It can’t curl your weather service or hit your Stripe endpoint. And that’s actually by design—you don’t want an AI making arbitrary network calls without your knowledge. What you do want is a clean architecture where Claude decides what to call, and your infrastructure handles the how.

That’s what this guide is about. We’re going to build the bridge between Claude’s reasoning capabilities and the outside world. By the end, you’ll have real patterns for wrapping external APIs as tools, handling auth, transforming responses, managing errors, and keeping costs under control. Let’s get into it.

The Architecture You Actually Need

Before we write a single line of code, you need to understand the two primary integration patterns. Pick the wrong one and you’ll fight your own architecture for the life of the project.

Pattern 1: Tool Use via the Anthropic API

You define tools in your API request. Claude sees the tool definitions, decides when to use them, and returns a tool_use content block telling your application what to call and with what arguments. Your application executes the actual API call, then feeds the result back to Claude as a tool_result. Claude incorporates that data and continues reasoning.

This is the pattern for custom applications—backends, CLI tools, chatbots, agents. You own the execution loop.

Pattern 2: Model Context Protocol (MCP) Servers

MCP is a standardized protocol where you build a server that exposes external APIs as tools. Claude Desktop, Claude Code, and other MCP-compatible clients discover these tools automatically. The MCP server handles the actual API calls.

This is the pattern for extending Claude’s capabilities in existing clients without building a full application.

Which One Do You Pick?

If you’re building a product, an agent, or a custom backend: Tool Use via the API. You get full control over the execution loop, error handling, retries, and auth management.

If you’re augmenting your personal Claude workflow or building internal tools for a team: MCP. You write the server once, and it works across any MCP-compatible client.

Both patterns share the same core concept: Claude never calls APIs directly. You build the bridge. Let’s see how.

Tool Use: Wrapping External APIs

This is where most production integrations live. You define a tool schema that describes what the external API does, Claude decides when to invoke it, and your code handles the actual HTTP request.

Defining a Tool That Wraps a Weather API

Here’s a complete example wrapping the OpenWeatherMap API. Pay attention to the schema—Claude uses the description fields to decide when and how to use the tool, so they need to be precise.





client = anthropic.Anthropic()

# Define the tool schema — this tells Claude what's available
tools = [
    {
        "name": "get_weather",
        "description": "Get current weather conditions for a specific city. Returns temperature, humidity, wind speed, and a text description. Use this when the user asks about current weather anywhere in the world.",
        "input_schema": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    "description": "City name, e.g. 'London' or 'San Francisco'"
                },
                "units": {
                    "type": "string",
                    "enum": ["metric", "imperial"],
                    "description": "Temperature units. Use 'imperial' for US users, 'metric' for everyone else."
                }
            },
            "required": ["city"]
        }
    }
]

def call_weather_api(city: str, units: str = "metric") -> dict:
    """Actually hit the OpenWeatherMap API."""
    api_key = os.environ["OPENWEATHER_API_KEY"]
    resp = httpx.get(
        "https://api.openweathermap.org/data/2.5/weather",
        params={"q": city, "appid": api_key, "units": units}
    )
    resp.raise_for_status()
    data = resp.json()

    # Transform the response — give Claude only what it needs
    return {
        "city": data["name"],
        "country": data["sys"]["country"],
        "temperature": data["main"]["temp"],
        "feels_like": data["main"]["feels_like"],
        "humidity": data["main"]["humidity"],
        "description": data["weather"][0]["description"],
        "wind_speed": data["wind"]["speed"],
        "units": units
    }

Notice what we’re doing in call_weather_api: we’re not dumping the entire OpenWeatherMap response into Claude’s context. We’re transforming it. The raw response includes icon codes, internal IDs, coordinate data, and a bunch of other fields Claude doesn’t need. Every token you feed Claude costs money and burns context window. Be surgical about what you pass back.

The Execution Loop

Now here’s the part that trips people up. You don’t just send the message and get a final answer. Tool use requires a conversation loop: send the message, check if Claude wants to use a tool, execute the tool, feed the result back, and let Claude generate the final response.

def chat_with_tools(user_message: str) -> str:
    """Complete tool-use loop with Claude."""
    messages = [{"role": "user", "content": user_message}]

    # First API call — Claude may request a tool
    response = client.messages.create(
        model="claude-sonnet-4-20250514",
        max_tokens=1024,
        tools=tools,
        messages=messages
    )

    # Loop until Claude gives a final text response
    while response.stop_reason == "tool_use":
        # Find the tool use block
        tool_block = next(
            b for b in response.content if b.type == "tool_use"
        )

        tool_name = tool_block.name
        tool_input = tool_block.input

        # Execute the actual API call
        if tool_name == "get_weather":
            try:
                result = call_weather_api(**tool_input)
                tool_result = {
                    "type": "tool_result",
                    "tool_use_id": tool_block.id,
                    "content": str(result)
                }
            except httpx.HTTPStatusError as e:
                tool_result = {
                    "type": "tool_result",
                    "tool_use_id": tool_block.id,
                    "content": f"Error: API returned {e.response.status_code}",
                    "is_error": True
                }

        # Feed the result back to Claude
        messages.append({"role": "assistant", "content": response.content})
        messages.append({"role": "user", "content": [tool_result]})

        response = client.messages.create(
            model="claude-sonnet-4-20250514",
            max_tokens=1024,
            tools=tools,
            messages=messages
        )

    # Extract the final text response
    return next(b.text for b in response.content if hasattr(b, "text"))

That while response.stop_reason == "tool_use" loop is the heart of it. Claude might chain multiple tool calls—check the weather in two cities, look up a database record, then call a calculation service. Your loop keeps running until Claude has everything it needs and produces a text response.

Authentication: The Part Everyone Gets Wrong

External APIs need auth. Your weather API needs an API key. Your internal services need bearer tokens. Your payment processor needs OAuth. Here’s how to handle each pattern without leaking credentials into Claude’s context.

API Keys

The simplest pattern. Keep keys in environment variables, never in tool definitions, never in system prompts.

# WRONG — don't put API keys anywhere Claude can see them
tools = [{
    "name": "get_weather",
    "description": "Get weather. API key is sk-abc123..."  # NO. STOP.
}]

# RIGHT — keys live in your execution layer
def call_weather_api(city: str, units: str = "metric") -> dict:
    api_key = os.environ["OPENWEATHER_API_KEY"]  # Claude never sees this
    resp = httpx.get(url, params={"q": city, "appid": api_key, "units": units})
    return transform_response(resp.json())

Claude doesn’t need to know how auth works. It just describes what it wants—”get the weather in Tokyo”—and your execution layer handles the credentials. This is a fundamental boundary. Claude reasons about what to do. Your code handles how.

Bearer Tokens and OAuth

For APIs that require user-specific tokens (Slack, GitHub, Google), you’ll typically have an OAuth flow in your application that stores tokens per user. Your tool execution layer retrieves the right token at runtime:

def execute_tool(tool_name: str, tool_input: dict, user_id: str) -> dict:
    """Tool executor with per-user auth."""

    # Retrieve the user's stored OAuth token
    token = token_store.get(user_id, service=tool_name_to_service(tool_name))

    if not token:
        return {"error": "User hasn't connected this service yet"}

    if token.is_expired:
        token = refresh_oauth_token(token)
        token_store.save(user_id, token)

    headers = {"Authorization": f"Bearer {token.access_token}"}
    # ... make the API call with proper auth

The key insight: auth is entirely your application’s responsibility. Claude doesn’t participate in OAuth flows, doesn’t store tokens, doesn’t refresh credentials. It just says “call this tool with these arguments” and trusts your infrastructure to handle the rest.

Error Handling: Because APIs Break

External APIs fail. They time out. They rate-limit you. They return 500s at the worst possible moment. Your tool execution layer needs to handle all of this gracefully, because if you just pass raw errors to Claude, it’ll either hallucinate a recovery or give the user a terrible experience.

The Resilient API Wrapper

Here’s a pattern I use in production for every external API integration. It handles retries, timeouts, rate limits, and gives Claude clean error information it can actually reason about.



from typing import Any

class ResilientAPIClient:
    """Wrapper for external API calls with retry logic and clean error handling."""

    def __init__(self, base_url: str, api_key: str, max_retries: int = 3):
        self.base_url = base_url
        self.max_retries = max_retries
        self.client = httpx.Client(
            base_url=base_url,
            headers={"Authorization": f"Bearer {api_key}"},
            timeout=10.0  # Hard timeout — don't let Claude's user wait forever
        )

    def call(self, method: str, path: str, **kwargs) -> dict[str, Any]:
        """Make an API call with automatic retries and clean error surfaces."""
        last_error = None

        for attempt in range(self.max_retries):
            try:
                resp = self.client.request(method, path, **kwargs)

                # Rate limited — back off and retry
                if resp.status_code == 429:
                    retry_after = int(resp.headers.get("Retry-After", 2))
                    time.sleep(min(retry_after, 10))  # Cap at 10s
                    continue

                resp.raise_for_status()
                return {"success": True, "data": resp.json()}

            except httpx.TimeoutException:
                last_error = "The service is taking too long to respond"
                time.sleep(2 ** attempt)  # Exponential backoff

            except httpx.HTTPStatusError as e:
                status = e.response.status_code
                if status >= 500:
                    last_error = "The service is experiencing issues"
                    time.sleep(2 ** attempt)
                elif status == 404:
                    return {
                        "success": False,
                        "error": "Resource not found",
                        "details": "The requested item doesn't exist"
                    }
                elif status == 403:
                    return {
                        "success": False,
                        "error": "Access denied",
                        "details": "Insufficient permissions for this operation"
                    }
                else:
                    return {
                        "success": False,
                        "error": f"Request failed with status {status}",
                        "details": e.response.text[:200]
                    }

        return {
            "success": False,
            "error": last_error or "Request failed after retries",
            "retries_exhausted": True
        }

There are a few things worth calling out here.

Clean error messages. We don’t pass raw stack traces to Claude. We give it structured error information—what happened, whether it’s retryable, what the user-facing implication is. Claude can then craft a helpful response like “I wasn’t able to check the weather right now—the service seems to be having issues. Want me to try again in a minute?” instead of dumping a traceback.

Capped retries and backoff. Three retries with exponential backoff handles transient failures without making the user wait forever. The 10-second cap on rate-limit waits prevents one slow API from blocking an entire conversation.

Structured returns. Every response from call() has a consistent shape—success boolean plus either data or error. This makes your tool execution layer predictable and easy to reason about.

Data Transformation: The Hidden Superpower

This is the part most tutorials skip, and it’s arguably the most important part of the whole integration. The data you get from external APIs is almost never in the right shape for Claude.

A Stripe payment intent response is 50+ fields. A GitHub issue has nested objects three levels deep. A database query might return 500 rows. You can’t dump all of that into Claude’s context and expect good results. You’ll burn tokens, confuse the model, and get worse answers.

Transform Before You Pass

The rule is simple: give Claude exactly the information it needs to answer the user’s question, and nothing more.

def transform_stripe_payment(raw: dict) -> dict:
    """Transform a Stripe PaymentIntent into what Claude actually needs."""
    return {
        "payment_id": raw["id"],
        "amount": raw["amount"] / 100,  # Stripe uses cents
        "currency": raw["currency"].upper(),
        "status": raw["status"],
        "customer_email": raw.get("receipt_email", "unknown"),
        "created": raw["created"],  # Unix timestamp
        "description": raw.get("description", "No description"),
        "last_four": raw.get("payment_method_details", {})
                        .get("card", {})
                        .get("last4", "N/A")
    }
    # We're dropping: metadata, charges array, shipping details,
    # payment method fingerprint, radar risk data, and 40+ other fields
    # Claude doesn't need any of that to answer "what's the status of payment X?"

Summarize Large Result Sets

When an API returns a list of results, don’t send all of them. Summarize, paginate, or filter.

def transform_database_results(rows: list[dict], query_context: str) -> dict:
    """Summarize large result sets instead of dumping everything."""
    if len(rows) <= 10:
        return {"results": rows, "total": len(rows)}

    return {
        "total_results": len(rows),
        "showing": "first 10",
        "results": rows[:10],
        "summary": {
            "date_range": f"{rows[-1]['created_at']} to {rows[0]['created_at']}",
            "unique_users": len(set(r["user_id"] for r in rows)),
        },
        "note": "More results available. Ask the user if they want to narrow the search."
    }

That note field is a nice trick. You’re giving Claude a hint about what to do next—suggest the user refine their query instead of trying to process hundreds of records.

Rate Limiting and Cost Management

Here’s where you need to think about two things at once: the rate limits of your external APIs, and the token cost of feeding responses back to Claude.

External API Rate Limits

Most APIs have rate limits. If you’re wrapping a free-tier weather API that allows 60 calls per minute, you need to track that in your execution layer—not in Claude.

from collections import defaultdict


class RateLimiter:
    def __init__(self):
        self.windows: dict[str, list[float]] = defaultdict(list)

    def check(self, api_name: str, max_calls: int, window_seconds: int) -> bool:
        now = time.time()
        calls = self.windows[api_name]
        # Remove expired entries
        self.windows[api_name] = [t for t in calls if now - t < window_seconds]

        if len(self.windows[api_name]) >= max_calls:
            return False

        self.windows[api_name].append(now)
        return True

rate_limiter = RateLimiter()

def execute_tool_with_limits(tool_name: str, tool_input: dict) -> dict:
    limits = {"get_weather": (60, 60), "search_database": (100, 60)}

    if tool_name in limits:
        max_calls, window = limits[tool_name]
        if not rate_limiter.check(tool_name, max_calls, window):
            return {
                "success": False,
                "error": "Rate limit reached for this service",
                "retry_after_seconds": window
            }

    return execute_tool(tool_name, tool_input)

Token Cost Awareness

Every character in a tool result costs tokens. A single bloated API response can cost more than the rest of the conversation combined. Here’s a rough guide:

  • Keep tool results under 2,000 tokens (~1,500 words) per call
  • If you’re returning structured data, use compact formats—skip unnecessary whitespace in JSON
  • Strip HTML, markdown formatting, and boilerplate from API responses
  • For large text content (articles, documents), summarize rather than include the full text

A Stripe webhook payload can be 4,000+ tokens raw. After transformation? Under 200 tokens. That’s a 20x cost reduction on every single tool call.

MCP: The Other Path

If you’re building an MCP server instead of a custom application, the concepts are identical—you still wrap APIs, transform data, handle errors. The difference is the transport layer and discovery mechanism.

Here’s what a minimal MCP server looks like for the same weather API:

from mcp.server.fastmcp import FastMCP

mcp = FastMCP("weather-service")

@mcp.tool()
async def get_weather(city: str, units: str = "metric") -> str:
    """Get current weather conditions for a city.

    Args:
        city: City name like 'London' or 'Tokyo'
        units: 'metric' for Celsius, 'imperial' for Fahrenheit
    """
    # Same API call, same transformation, same error handling
    try:
        result = await call_weather_api(city, units)
        return f"Weather in {result['city']}: {result['temperature']}°, " \
               f"{result['description']}, humidity {result['humidity']}%"
    except Exception as e:
        return f"Could not fetch weather: {str(e)}"

if __name__ == "__main__":
    mcp.run()

Same patterns, different packaging. The MCP framework handles the protocol details—discovery, transport, message framing. You focus on wrapping the API and transforming the data.

Real-World Integration: Putting It All Together

Let’s walk through a realistic scenario. You’re building a customer support agent that needs to:

  1. Look up customer information from your database
  2. Check their recent orders via your order management API
  3. Process refunds through Stripe

That’s three external APIs, three tool definitions, three sets of auth, and a conversation loop that might chain all three in a single interaction.

The architecture looks like this:

User asks: “I need a refund for order #4521”

Claude’s reasoning chain:

  1. “I need to look up order #4521” → calls get_order tool
  2. “The order is $47.99 for user [email protected], delivered 3 days ago” → reasons about refund eligibility
  3. “This is within the refund window, let me process it” → calls process_refund tool
  4. “Refund of $47.99 initiated, confirmation ID ref_abc123” → generates response

Your application handles all three API calls, manages the Stripe API key, validates the refund amount, and feeds clean results back to Claude. Claude never touches a network socket. It reasons about what to do, and your infrastructure executes it.

The separation is clean, auditable, and secure. You can log every tool call, rate-limit by user, enforce business rules before executing destructive actions (like refunds), and swap out API providers without changing Claude’s tool definitions.

The Mistakes That’ll Bite You

Before we wrap up, let me save you some pain. These are the mistakes I see most often in production Claude + API integrations:

Putting API keys in system prompts. Don’t do this. Claude doesn’t need your keys. Your execution layer does. Keep them in environment variables or a secrets manager.

Dumping raw API responses. Transform everything. Strip what Claude doesn’t need. Your token costs and response quality will both improve dramatically.

No error handling in tool execution. When an API fails and you return nothing, Claude hallucinates an answer. When you return a clean error message, Claude tells the user what happened and suggests alternatives.

Ignoring the loop. Tool use isn’t a single request-response. It’s a loop. Claude might call multiple tools, or the same tool multiple times. Your code needs to handle that.

Over-broad tool definitions. A tool called “do_anything” with a freeform input schema is useless. Claude won’t know when to use it. Be specific in your names, descriptions, and schemas.

Skipping rate limits. Your free-tier API lets you make 60 calls per minute. A chatty agent can blow through that in two conversations. Track and limit at the execution layer.

Where to Go From Here

You’ve got the patterns. Here’s how to level up from here:

Multi-tool orchestration. Define 5-10 tools and let Claude chain them. You’ll be surprised at how well it reasons about tool ordering and data dependencies.

Streaming with tool use. The Anthropic API supports streaming responses even during tool-use loops. This gives your users real-time feedback while Claude is thinking and making API calls.

Caching tool results. If your weather API data is valid for 10 minutes, cache it. If a database query returns the same result within a session, cache it. Claude doesn’t need to know you’re caching—it gets the same result either way, and you save money on both API calls and tokens.

Guardrails and business logic. Wrap destructive tools (delete, refund, cancel) with confirmation steps. Your execution layer can return “This will refund $247.99 to the customer. Should I proceed?” and let Claude ask the user for confirmation before executing.

The core pattern never changes: Claude reasons, you execute. Keep that boundary clean and you can integrate with anything—weather APIs, databases, payment processors, internal microservices, IoT devices, whatever your architecture demands. The bridge is yours to build.

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.