All Articles OpenClaw

OpenClaw with Ollama: Running Local Models for Zero API Cost and Maximum Privacy

Here's the problem nobody wants to say out loud: every API call you make to a cloud LLM provider leaves a footprint.

Here’s the problem nobody wants to say out loud: every API call you make to a cloud LLM provider leaves a footprint. Your data travels across networks, gets logged somewhere, and even with SOC 2 compliance and privacy policies, you’re trusting someone else with your prompts. For some teams, that’s unacceptable. For others, it’s a cost killer.

This is where Ollama enters the chat.

If you’re running OpenClaw (Anthropic’s open-source model inference framework), you can skip the cloud entirely and run quantized models locally—on your machine, completely offline, with zero API costs. No rate limits. No billing surprises. No external dependencies. Just you, your hardware, and whatever model fits in your VRAM.

Let’s walk through how to set this up, what to expect, and whether it’s right for your use case.

Why Local Models with OpenClaw? The Three-Pillar Case

Before diving into setup, let’s nail down why you’d want this:

Privacy First: Your prompts never leave your machine. Not to Anthropic, OpenAI, or anyone else. If you’re handling sensitive data—medical records, legal documents, proprietary code—local inference means nobody’s watching. This matters for compliance (HIPAA, GDPR, internal policy) and peace of mind.

Cost Zero: After your initial hardware investment, every inference is free. No per-token billing. No surprise charges because your application suddenly scaled. For high-volume use cases—batch processing, internal tools, always-on agents—this can save thousands monthly.

Control Absolute: You own the model, the context window, the inference pipeline. No provider can deprecate your endpoint, change their pricing, or sunset a feature you depend on. You set the quantization level, the context length, the sampling parameters. You’re in charge.

The tradeoff? You’re responsible for hardware. A 7B model (like Mistral or Llama 2) needs 8-16GB of VRAM. A 13B model needs 24GB+. And inference is slower than cloud APIs—typically 15-50 tokens per second on consumer GPUs, vs. 100+ from commercial providers.

But for many workflows, that tradeoff is worth it.

The Setup: Ollama + OpenClaw in Five Steps

Here’s what you’re actually doing:

  1. Install Ollama (the model runtime)
  2. Pull a quantized model (7B, 13B, or smaller)
  3. Start the Ollama server (it listens on http://127.0.0.1:11434)
  4. Configure OpenClaw to point at that server
  5. Run inference against your local model

Let’s do this.

Step 1: Install Ollama

Ollama is a lightweight wrapper around GGML (a C++ inference engine). It handles model downloading, quantization, and serving. It’s ~200MB, runs on Mac, Linux, and Windows, and it’s genuinely simple to install.

Go to ollama.ai and download the installer for your OS. Run it. It takes two minutes.

Once installed, verify it’s working:

ollama --version
# Output: ollama version 0.1.x (or whatever current version is)

Done. You’ve got the runtime.

Step 2: Pull a Model

Ollama’s model library is hosted at ollama.ai/library. This is where you find pre-quantized versions of popular open models. The naming scheme is straightforward: modelname:quantization.

Common options:

  • mistral:7b-instruct-q4_K_M — 7B Mistral, 4-bit quantization, ~4.7GB
  • llama2:7b-chat-q4_K_M — 7B Llama 2, 4-bit, ~4GB
  • neural-chat:7b-q4_K_M — Intel’s chat model, 4-bit, ~4GB
  • dolphin-mixtral:8x7b-q4_K_M — 8x7B mixture-of-experts, 4-bit, ~20GB
  • mistral:13b-instruct-q4_K_M — 13B Mistral, 4-bit, ~8GB

The quantization level (q4, q5, q8, etc.) controls quality vs. size:

  • q4 (4-bit): ~50% of float32 size, minimal quality loss, fastest
  • q5 (5-bit): ~62% size, better quality, slightly slower
  • q8 (8-bit): ~80% size, near-original quality, slowest of the quantized options
  • f16 (no quantization): Full model, massive size, best quality, requires max VRAM

For most use cases, q4_K_M is the sweet spot. You get 95% of the model’s capability at half the size.

Let’s pull a model:

ollama pull mistral:7b-instruct-q4_K_M

This downloads the model (4-5GB for 7B models, 8-10GB for 13B) and stores it locally. First time takes a few minutes depending on your connection. Subsequent runs use the cached version.

Once done:

ollama list
# Output:
# NAME                           ID              SIZE      MODIFIED
# mistral:7b-instruct-q4_K_M     1234abcd...     4.7 GB    2 minutes ago

Step 3: Start the Ollama Server

By default, Ollama runs as a background service (on Mac/Linux via systemd, on Windows via service). But you can also start it manually for testing:

ollama serve

This starts the server on http://127.0.0.1:11434. You’ll see output like:

2026-03-17 10:25:33 INFO listening on 127.0.0.1:11434

Leave this running in a terminal (or let the background service handle it).

Verify the server is alive:

curl http://127.0.0.1:11434/api/tags

You should get JSON back listing your installed models. If you do, you’re good.

Step 4: Configure OpenClaw for Ollama

OpenClaw reads its configuration from an openclaw.json file (typically at the root of your OpenClaw project or in the config directory).

Here’s the Ollama configuration block:

{
  "api": {
    "type": "openai-completions",
    "base_url": "http://127.0.0.1:11434/v1",
    "api_key": "ollama-local",
    "model": "mistral:7b-instruct-q4_K_M",
    "context_length": 32768,
    "max_tokens": 2048
  }
}

Let’s break this:

  • type: "openai-completions" means we’re using the OpenAI-compatible API format. Ollama exposes its models via this interface, so OpenClaw treats them like OpenAI models (but running locally).
  • base_url: http://127.0.0.1:11434/v1 — this is Ollama’s OpenAI-compatible endpoint. The /v1 suffix is required.
  • api_key: "ollama-local" — a dummy key. Ollama doesn’t authenticate locally, but OpenClaw still wants a value here. Any string works.
  • model: "mistral:7b-instruct-q4_K_M" — the model you just pulled. Use whatever you installed.
  • context_length: 32768 — the maximum context window you want to use. Mistral 7B supports 32k tokens. If you’re using Llama 2 (4k native), set this lower.
  • max_tokens: 2048 — the maximum output length per response. Adjust based on your needs.

Save this to openclaw.json and you’re configured.

Step 5: Test It

Run your first inference:

# Using OpenClaw CLI (if you have it)
openclaw inference --prompt "What is quantum computing?" --model mistral

# Or via curl (testing Ollama directly)
curl http://127.0.0.1:11434/api/generate \
  -d '{
    "model": "mistral:7b-instruct-q4_K_M",
    "prompt": "What is quantum computing?",
    "stream": false
  }'

You should get a response. It’ll be slower than cloud APIs (5-15 seconds for a typical response on a mid-range GPU), but it’s free, it’s local, and it’s yours.

Context Matters: The 64k+ Requirement

Here’s a detail that catches people: Ollama models in OpenClaw perform significantly better with 64k+ context windows.

Why? Because OpenClaw includes a lot of system context—chainable prompts, memory, reasoning loops, validator stages. If your context window is too small, you lose valuable prompt real estate.

When you’re configuring OpenClaw for Ollama, set context_length to the model’s maximum supported length:

  • Mistral 7B: 32k tokens (supported)
  • Llama 2 7B: 4k tokens (upgrade to Llama 2 Uncensored 32k or Code Llama for more)
  • Neural Chat 7B: 8k tokens
  • Dolphin Mixtral 8x7B: 32k tokens

If you need 64k+ specifically, you’re limited to models like:

  • Llama 2 Long (80k, but quantized variants are less common)
  • MPT-7B (84k, via Ollama)
  • Falcon-7B (2k, actually shorter—skip this)

For most setups, 32k context from Mistral is plenty. You’ll rarely hit the ceiling.

Hardware Reality Check: Do You Have Enough?

Ollama’s hardware requirements are simple but real:

  • 4GB VRAM: Gateway only. Not enough for inference. Use cloud APIs instead.
  • 8GB VRAM: Fits a 7B model quantized to q4. Expect 5-10 tokens/sec. Usable for light workloads.
  • 16GB VRAM: Comfortable for a 7B model with reasonable performance (15-30 tokens/sec). Recommended minimum.
  • 24GB VRAM: Fits a 13B model q4 (8-10GB) + system overhead. Faster, better quality.
  • 32GB+ VRAM: Run larger models (13B q5, multiple 7B models, or 70B models).

You can check your GPU VRAM:

# NVIDIA GPU
nvidia-smi

# Apple Silicon (M1/M2/M3)
system_profiler SPDisplaysDataType | grep "VRAM"

# AMD GPU
rocm-smi

If you don’t have a GPU, Ollama can run on CPU, but expect 0.5-2 tokens/sec. Not recommended for anything but testing.

Also: Ollama automatically offloads models to GPU when possible. On Apple Silicon Macs, it’s seamless. On NVIDIA, it requires CUDA (which Ollama installs automatically). On AMD, you need rocm.

Data Privacy: The Architecture

This is the core promise of local models, so let’s be precise about the privacy model:

  • Input/Output: Your prompts and responses stay on your machine. They’re never sent to Ollama’s servers (Ollama doesn’t have cloud servers—it’s pure local inference).
  • Model Weights: Downloaded once from Ollama’s model library (ollama.ai/library), stored locally. After that, air-gapped.
  • Logs: By default, Ollama doesn’t log your prompts unless you configure it. OpenClaw might log internally depending on your configuration.
  • Hardware: Your GPU/CPU sees all data. If you’re worried about supply-chain attacks or hardware vulnerabilities, that’s a different threat model.

The bottom line: Zero exfiltration of text data. If privacy is your priority, this is the setup.

Deep Dive: openclaw.json Configuration

Understanding your configuration file is critical. Let’s walk through each field and what it means for your local setup.

The api object tells OpenClaw how to talk to your Ollama server:

type: "openai-completions" is the magic here. Ollama implements the OpenAI API spec, which means any tool built for OpenAI (including OpenClaw) can talk to it transparently. You’re not using a proprietary Ollama protocol—you’re using a standardized interface. This is why Ollama integrates so cleanly with everything.

base_url: http://127.0.0.1:11434/v1 is your local Ollama server. The /v1 path is the OpenAI-compatible endpoint. Without it, requests fail. The 127.0.0.1 loopback address means Ollama is only accessible from your machine—not the network. If you wanted to expose Ollama to other machines, you’d change this to your machine’s IP, but that’s typically not recommended for local development. The port 11434 is Ollama’s default and rarely changes unless you explicitly configure it.

api_key: "ollama-local" is a dummy key. Real Ollama doesn’t authenticate. But OpenClaw expects something in this field, so we provide any string. In production setups where you might proxy Ollama through a reverse proxy with authentication, this would be a real key.

model: Must match exactly what you pulled. If you pulled mistral:7b-instruct-q4_K_M, use that exact string. If you use mistral without the version tag, it’ll try to pull the default (usually the latest), which might not be what you want.

context_length: Set this to your model’s actual max. Mistral 7B supports 32k. Llama 2 base is 4k (unless you’ve used fine-tuned variants). Code Llama goes to 100k. Setting this higher than the model supports is harmless (Ollama caps it), but setting it lower wastes potential. Match the model spec exactly.

max_tokens: This is the maximum number of tokens the model will generate in a single request. If you ask OpenClaw to write a novel, it can’t exceed this. 2048 is reasonable for most tasks. For summarization, 512 works. For creative writing, go higher—4096 or 8192.

Pro tip: You can have multiple sections in openclaw.json for different model profiles. Create api_lightweight, api_creative, api_coding with different context and max_token settings, then switch between them based on the task.

Privacy Architecture: The Complete Picture

The privacy promise of local models deserves a detailed explanation because it matters for compliance.

Data Isolation at Three Levels:

First, the network level. Your prompts never leave your machine on the wire. There’s no HTTPS connection to Anthropic’s servers, no webhooks, no telemetry calls (unless you explicitly enable Ollama logging). The data stays on your LAN, in your RAM, on your storage.

Second, the process level. Ollama runs as a local service. OpenClaw talks to it via localhost HTTP. No cloud intermediary. No third-party API gateway. It’s direct binary-to-binary communication.

Third, the persistence level. Models are stored in your ~/.ollama/models directory (or wherever Ollama is configured). That’s your machine. Offline. If your hard drive is encrypted, the model weights are encrypted. If it’s not, they’re readable only to processes running as your user (depending on permissions).

What OpenClaw Logs:

This varies by configuration. By default, OpenClaw logs:

  • Timestamps of inference requests
  • Token counts (how many tokens used)
  • Model names
  • Status codes (success/failure)

It does NOT log:

  • Prompt content
  • Response content
  • User data

If you’re paranoid (and you should be for healthcare/legal data), you can disable all logging. Set log_prompts: false in your config.

Hardware Threat Model:

Here’s the catch: your GPU and CPU see the data. If someone has physical access to your machine or can install kernel-level malware, they can snoop. But that’s a different threat model than cloud APIs. For most teams, “nobody external sees our data” is the requirement, and Ollama delivers that.

The 64k Context Requirement Explained

Earlier I mentioned that OpenClaw works better with 64k+ context. Let me explain why this matters and when it actually matters.

OpenClaw’s architecture includes several heavy context consumers:

System prompts — These describe the task, the rules, the expected output format. For complex reasoning tasks (story agents, code review), system prompts can be 5-10k tokens alone.

Conversation history — The back-and-forth between human and agent. Each exchange adds tokens. A typical multi-turn conversation easily consumes 10-20k tokens.

Memory and context injection — OpenClaw pulls relevant memory entries and injects them. For an agent that has learned a lot, this can be 5-15k tokens.

Working context — The actual task. Writing a chapter might need 10-20k tokens for the chapter being drafted.

Add it up: 10 + 15 + 10 + 15 = 50k tokens, and you haven’t written much yet. So a 32k context window (common in 7B models) gets cramped fast. You either truncate history (losing context) or cap your output (limiting what the agent can do).

With 64k or higher, you have breathing room. You can keep full conversation history, load more memories, use longer system prompts, and still have space for the actual task.

When 32k is Enough:

  • Single-turn tasks (answer this question)
  • Deterministic tasks (validate this code)
  • Shorter documents (articles under 5k words)
  • Systems that explicitly manage memory truncation

When You Need 64k+:

  • Long conversations
  • Creative writing (need to track character arcs, plot threads)
  • Code review of large files
  • Research tasks that build on each other
  • Systems that don’t want to implement aggressive truncation

The workaround if your model maxes out at 32k? Be selective. Load fewer memories. Truncate older history. Use shorter system prompts. But it’s better to have the context and not need it than to need it and not have it.

Model Selection: Beyond the Obvious Choices

I listed common models earlier, but let’s dig into selection criteria since this is where most people stumble.

Mistral 7B — Versatile, good instruction-following, decent at code. This is the safe choice. Works for 80% of use cases. ~4.7GB in q4.

Llama 2 7B — Open weights, good safety training, reasonable at most tasks. Smaller context window (4k) is a limitation. ~3.8GB. Less impressive than Mistral but fully open and battle-tested.

Neural Chat 7B — Intel’s model, optimized for dialogue. Better at conversation than Mistral. Similar size.

Code Llama — If you’re writing code, this is better than Mistral. Larger context window (100k). More expensive (~24GB even quantized to q4).

Dolphin Mixtral 8x7B — Mixture of experts architecture. Better quality than single 7B models but needs 20GB VRAM. Faster than 13B models due to sparse activation.

Selecting Based on Your Workload:

For story/creative: Use a model with good instruction-following and larger context. Mistral or Llama 2 Long if available.

For code: Use Code Llama or fine-tuned variants if possible. Mistral is acceptable fallback.

For general tasks: Mistral 7B is the default.

For constrained VRAM (8GB): Llama 2 7B q4, Neural Chat q4, or even q3 quantization.

For maximum quality: 13B models if you have 24GB+ VRAM, or q5 quantization if you’re willing to sacrifice speed.

Troubleshooting Deep Dive

The gotchas section was quick. Let me expand on each because they’re where most people get stuck.

“Connection refused” — The Most Common Error

This means OpenClaw can’t reach Ollama. Verify step-by-step:

  1. Is Ollama running? ps aux | grep ollama (on Mac/Linux) or check Windows Task Manager.
  2. Is it listening on the right port? lsof -i :11434 (Mac/Linux) or netstat -tuln | grep 11434 (Linux).
  3. Is the OpenClaw URL exactly right? Test with curl: curl -X POST http://127.0.0.1:11434/api/generate -d '{"model":"mistral:7b-instruct-q4_K_M","prompt":"hi","stream":false}'. If this works, the server is fine. If not, Ollama isn’t responding.
  4. If it’s a firewall issue (unlikely on localhost but possible), temporarily disable your firewall and test.

“Model not found”

Run ollama list and check exact spelling. The name must match perfectly. If you’re getting “model not found” from OpenClaw but ollama list shows it, check:

  1. The quantization tag must match. mistral:7b-instruct-q4_K_M is different from mistral:7b-instruct-q5_K_M.
  2. The base URL in openclaw.json might be wrong. Test the URL directly with the curl command above.
  3. Ollama might be in the process of downloading the model. Wait.

“CUDA out of memory”

Your GPU doesn’t have enough memory for the model. Solutions in order of preference:

  1. Use a smaller quantization: q3 instead of q4, q4 instead of q5.
  2. Unload other models: ollama rm modelname for models you’re not using.
  3. Reduce your context window in openclaw.json: lower context_length slightly (but this affects quality).
  4. Use a smaller model entirely (7B instead of 13B).
  5. Use CPU inference (slow, but works).

High VRAM Usage at Idle

Ollama doesn’t unload models automatically. If you ran a model and now don’t need it, manually unload: ollama rm modelname. Or restart Ollama with ollama serve (kills all loaded models).

If this is happening during a session, it means your model is larger than you thought. Check the actual size: ollama show mistral:7b-instruct-q4_K_M and look at “parameters_size”.

Slow Inference (2-3 tokens/sec)

This usually means you’re on CPU. Verify GPU is active:

ollama show mistral:7b-instruct-q4_K_M — look for gpu_layers. If it’s 0, GPU isn’t being used. If it’s >0, GPU is active but might be slow.

On NVIDIA: Check CUDA is installed (nvidia-smi shows your GPU). Ollama auto-installs CUDA requirements, but sometimes NVIDIA drivers are outdated.

On Apple Silicon: GPU should be automatic. If slow, restart Ollama and your Mac (I know, but it works).

On AMD: Requires rocm. Installation is complex. Consider using CPU if this is too much hassle.

Expected speeds: 10-50 tokens/sec on consumer GPUs, 100+ on data center GPUs, 0.5-2 on CPU.

Quantization: The Quality-Speed Tradeoff Explained

I mentioned quantization briefly, but it deserves deeper explanation because it’s the lever you control for performance.

Quantization is compression. A full-precision LLM like Mistral 7B stores each parameter as a 32-bit float. That’s 7 billion parameters × 4 bytes = 28GB. Way too big.

Quantization reduces precision. Instead of 32 bits, you use 4 bits (q4), 5 bits (q5), 8 bits (q8). Mathematically, this shouldn’t work—you lose information. But empirically, it does work. A 4-bit Mistral is nearly indistinguishable from the full model for most tasks.

How 4-bit Quantization Works:

Instead of storing each weight as 32.7234893424, you bucket them. “Is this weight in the 0-10 range?” If yes, map it to one of 16 discrete values (4 bits = 2^4 = 16 values). The model learns to reconstruct the original value from the compressed representation. It’s lossy, but the loss is tiny for inference.

The Tradeoff Spectrum:

  • f16 (float16): 2 bytes per param, 14GB for 7B model. Original precision. Perfect quality.
  • q8 (8-bit): 1 byte per param, 7GB for 7B model. ~99% quality, same speed as f16.
  • q5 (5-bit): 0.625 bytes per param, 4.4GB. ~97% quality, slightly faster inference.
  • q4 (4-bit): 0.5 bytes per param, 3.5GB. ~95% quality, 10-20% faster inference.
  • q3 (3-bit): 0.375 bytes per param, 2.6GB. ~90% quality, 30-40% faster but noticeably degraded.

Quality Loss Curve:

From q8→q4, you lose ~5% of capability but save ~50% of VRAM. That’s a good trade.

From q4→q3, you lose another ~5% but only save ~25%. Most teams skip q3.

Real-World Testing:

If you care about quality, test it. Take a task you know the full model does well (writing a poem, explaining a concept), run it on q8 and q4, compare outputs. You’ll probably not see a difference. Then try q3 and you might notice degradation.

Quantization Type Matters:

There are different quantization algorithms. GGML (what Ollama uses) has several:

  • q4_K_M (Ollama default) — “K” means it uses a key-value quantization scheme. “M” means medium. This is the sweet spot. Good quality, good speed, well-balanced.
  • q4_K_S — “S” means small. Trades 1-2% quality for smaller file size. Rarely worth it.
  • q4_0 — Older algorithm, rarely used now. Skip this.
  • q5_K_M, q8_K_M — Same idea, higher precision variants.

Stick with q4_K_M unless you have a specific reason.

Setup Optimization: Advanced Techniques

Once you’ve got the basics running, here are power moves to squeeze more performance.

GPU Memory Pinning — When Ollama loads a model, it can either stream from disk or preload to GPU memory. Preloading is faster but uses more VRAM. Configure in Ollama (you’ll need to edit the Ollama config file, usually ~/.ollama/modelfile):

PARAMETER num_gpu 99  # Use all GPU layers (default)
PARAMETER num_gpu 32  # Use only 32 layers on GPU, rest on CPU

Fewer GPU layers = slower inference but lower VRAM. More GPU layers = faster but more VRAM.

Prompt Caching — If you run the same system prompt repeatedly (which OpenClaw does), Ollama can cache it. No latency hit, just memory efficiency. This is handled automatically, but worth knowing it exists.

Batch Processing — If you have multiple inference requests, bundle them. Instead of running 10 requests serially, queue them and run in batches of 3-4. This increases GPU utilization and can give 30-50% throughput improvement.

# Serial: 10 requests × 5 seconds each = 50 seconds total
for i in {1..10}; do
  curl http://127.0.0.1:11434/api/generate -d "{...}"
done

# Batched: Still 10 requests but better GPU utilization, might be 30-35 seconds
# This requires application-level batching, not a command-line trick

Offload Tuning — On NVIDIA GPUs, you can control how many model layers run on GPU vs. CPU. More layers on GPU = faster but more VRAM. The sweet spot is usually 70-80% of layers on GPU, 20-30% on CPU. Ollama figures this out automatically, but you can tweak it.

Running Multiple Models Simultaneously

A power user question: can I run two models at once?

Yes, but with caveats. If you have 32GB VRAM, you might fit a 7B model (4.7GB q4) and a 13B model (8GB q4) simultaneously. But running inference on both at the same time will contend for GPU resources. One will slow the other down.

Strategy 1: Sequential Loading

Load one model, run inference, unload, load the next. Slower overall but minimal VRAM needed.

ollama run mistral:7b-instruct-q4_K_M
# ... use it ...
# Ctrl+C to stop
ollama run llama2:7b-chat-q4_K_M

Strategy 2: Keep Multiple Loaded

Keep both models in VRAM, switch between them. Ollama will evict unused models after a timeout (configurable).

# In your OpenClaw config, use different models for different tasks
# Story agent uses mistral, Code agent uses llama2
# They'll stay loaded if you alternate between them quickly

Strategy 3: Model Server Orchestration

For production systems, run multiple Ollama instances on different GPUs. More complex but fully parallel.

# Terminal 1: GPU 0
CUDA_VISIBLE_DEVICES=0 ollama serve --port 11434

# Terminal 2: GPU 1
CUDA_VISIBLE_DEVICES=1 ollama serve --port 11435

# In OpenClaw config, have different agents use different ports

Most people stick with Strategy 1 (sequential). It’s simple and works fine.

When to Run Ollama vs. When to Scale

This is the honest conversation: local models are great until they’re not.

As you increase workload (more requests per day, larger batch sizes, more complex tasks), you’ll hit limits:

  • VRAM exhaustion: You can’t fit a bigger model.
  • Latency: 20 tokens/sec is fine for batch processing but too slow for real-time user-facing features.
  • Uptime: Your machine crashes or needs to restart. Local Ollama goes down with it.
  • Scaling: You can’t distribute load across multiple machines.

At this point, consider hybrid or full cloud:

Hybrid Approach:

  • Use Ollama for simple, deterministic tasks (routing, classification, fact validation)
  • Use cloud APIs (Claude, GPT-4) for complex reasoning and creative tasks
  • Route requests based on complexity

Full Cloud:

  • Use OpenClaw with Claude API or OpenAI API
  • Infinite scale, no hardware management, access to best models
  • Costs money but infrastructure is someone else’s problem

Most teams start local (good for dev, good for privacy), then add cloud as workload grows.

Long-Term Maintenance: Keeping Ollama Healthy

If you’re running Ollama for months or years, a few hygiene habits help.

Weekly: Run ollama list and remove models you no longer use.

Monthly: Check storage usage (du -sh ~/.ollama/models). If it’s growing faster than expected, you might have duplicate quantizations. Clean up.

Quarterly: Update Ollama itself. New versions often include performance improvements.

# On Mac
brew upgrade ollama

# On Linux/Windows, download the latest from ollama.ai

When something breaks: Ollama logs are in ~/.ollama/logs on most systems. Check there before panicking.

The point: Ollama is stable and mature at this point. Set-and-forget mostly works. Just don’t ignore it for a year and expect it to run perfectly.

This is the pragmatic question: should your OpenClaw setup point at Ollama or Claude/OpenAI?

Choose Local (Ollama) if:

  • Privacy is non-negotiable (healthcare, legal, proprietary data)
  • You process high volume and costs matter (100+ requests/day)
  • You want 100% uptime control (no service outages, rate limits, or deprecations)
  • You have >16GB VRAM and tolerance for slower inference
  • You’re building internal tools that never talk to the outside world

Choose Cloud if:

  • You need cutting-edge models (GPT-4, Claude 3 Opus)
  • Inference speed matters (real-time user-facing features)
  • You want zero infrastructure overhead
  • Your workload is variable/bursty (no sense owning hardware)
  • You’re fine with data traversing networks

Choose Hybrid if:

  • You do simple tasks locally (classifying, routing) and complex tasks in the cloud
  • You’re building agents that spawn different models for different steps
  • You need privacy for most flows but speed for a few

With OpenClaw, all three approaches are supported. Pick what fits.

Next Steps

Once you’ve got Ollama + OpenClaw running:

  1. Test different models: Mistral, Llama 2, Neural Chat, Code Llama. See which fits your workload.
  2. Tune quantization: Try q5 or q3 to find your quality-speed tradeoff.
  3. Increase context: If 32k feels cramped, look for 64k-context models.
  4. Monitor performance: Track token/sec and VRAM usage. Optimize from there.
  5. Automate: Set up Ollama as a service so it starts on boot.

You’ve got a free, private, always-on LLM at your fingertips. The setup takes 30 minutes. The payoff is real.


Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.