# Run OpenClaw 100% Free with Ollama: No API Keys, No Monthly Bills

Canonical: https://clawdocx.com/blog/openclaw-ollama-free-local-setup
Author: Sam Okafor
Published: 2026-03-08
Updated: 2026-09-19

> Run OpenClaw on a local model with Ollama: which tag fits your RAM, the documented install and pull commands, and the openclaw.json provider block.

## Which local model should you run OpenClaw on?

Your memory picks the model, so start there. On **16GB**, pull **`gemma4:12b`**, which Ollama's library lists as a 7.6GB download with a 256K context window and tools, thinking and vision support. On **24 to 32GB**, pull **`qwen3.6:27b`**, an 18GB download with the same 256K window and the same three capability badges. On **32GB at a squeeze and 48GB comfortably**, pull **`qwen3.6:35b`**, a 23GB download that wants the extra headroom a long context needs.

One ranking and one vendor recommendation sit behind these picks. haimaker.ai's ranked test of local models for OpenClaw, updated July 23, 2026, puts `qwen3.6:27b` first and scores models on SWE-bench Verified results, tool-calling reliability and real-world agent performance. Ollama's own OpenClaw integration page recommends the `gemma4` family for local use at around 16 GB of VRAM and `qwen3.5` at around 11 GB, which is what puts gemma4 in the 16GB row.

The cost side is the easy part. OpenClaw's Ollama documentation is blunt about it: "Ollama runs locally and is free, so all model costs are `0` for both auto-discovered and manually defined models." The bill you are replacing is per-token; what you keep paying is hardware and electricity.

## Local models by hardware tier

Download sizes, context windows and capability badges below are Ollama's own figures from each library page as of September 19, 2026. Capability badges are published per model family, not per tag. The memory column leaves headroom for the operating system and the KV cache rather than matching the download size to the memory size, which is why a 23GB model is not a 24GB machine's model. On Apple silicon the model shares one pool of unified memory with the system. On a PC with a discrete GPU the number that matters is VRAM, not system RAM, and a model that does not fit in VRAM spills to the CPU; confirm the split under `PROCESSOR` in `ollama ps`.

| Memory for the model | Model | Ollama tag | Size, window and badges |
|---|---|---|---|
| 16GB | Gemma 4 12B | `gemma4:12b` | 7.6GB download, 256K window, tools, thinking and vision. Ollama's OpenClaw page recommends the gemma4 family locally at around 16 GB of VRAM; the 12b build is the one that leaves room underneath that on a 16GB machine. The 256K window is not reachable here: Ollama's VRAM-tiered default gives 4k, and Step 4 raises it to 64k. |
| 16GB | Qwen3.5 9B | `qwen3.5:9b` | 6.6GB download, 256K window, tools, thinking and vision. The smaller of the two. Ollama's OpenClaw page puts the qwen3.5 family at around 11 GB of VRAM. |
| 24 to 32GB | gpt-oss 20B | `gpt-oss:20b` | 14GB download, 128K window, tools and thinking. No vision badge, so no image input. |
| 24 to 32GB | Qwen3.6 27B | `qwen3.6:27b` | 18GB download, 256K window, tools, thinking and vision. First in haimaker.ai's ranking, updated July 23, 2026. The same applies here: the 24 to 48 GiB default is 32k. |
| 24 to 32GB | Gemma 4 26B | `gemma4:26b` | 19GB download, 256K window, tools, thinking and vision. The tag OpenClaw's own timeout example uses. A mixture-of-experts build with 4B active parameters, so it runs faster than its size suggests, but it is a gigabyte tighter than the 27b on a 24GB machine. |
| 32GB at a squeeze, 48GB and up comfortably | Qwen3.6 35B | `qwen3.6:35b` | 23GB download, 256K window, tools, thinking and vision. Fits on 32GB; 48GB leaves room for macOS and a long context. |
| 96GB and up | gpt-oss 120B | `gpt-oss:120b` | 65GB download, 128K window, tools and thinking. |

Two models worth naming and then skipping. `mistral:7b` is only a 4.4GB download, but Ollama's library lists its context window at 32K, below the 64k tokens Ollama's OpenClaw page recommends for local agent use. `llama4` starts at a 67GB download for the `16x17b` tag, so its 10M context window is not a consumer-hardware option.

> **Note**
>
> Ollama also publishes `-mlx` variants of several of these tags for Apple silicon, at slightly different download sizes. `qwen3.6:27b-mlx` is 19GB against 18GB for `qwen3.6:27b`. Pick one and pin it rather than letting a default tag decide.

## Step 1: Install Ollama

Ollama publishes a download page per platform with the command on it. On macOS and Linux it is the same script:

```bash
curl -fsSL https://ollama.com/install.sh | sh
```

On Windows, Ollama's download page gives a PowerShell one-liner instead:

```powershell
irm https://ollama.com/install.ps1 | iex
```

Each platform has a documented floor. Ollama's macOS page requires "MacOS Sonoma (v14) or newer" and Apple M series silicon for GPU support, with x86 Macs limited to CPU. The Windows page requires "Windows 10 22H2 or newer, Home or Pro", runs as a native Windows application with NVIDIA and AMD Radeon GPU support, and notes that the API is served on `http://localhost:11434`. Ollama binds `127.0.0.1` on port 11434 by default, which is the address the config block below uses.

The macOS app and the Linux systemd service start the server for you. Check that it answers:

```bash
ollama -v
curl http://127.0.0.1:11434/api/tags
```

If nothing answers, start the server yourself and leave it running in its own terminal:

```bash
ollama serve
```

> **Tip**
>
> Ollama ships a shortcut for this entire post. `ollama launch openclaw` prompts to install OpenClaw through npm if it is missing, shows a security notice about tool access on first launch, lets you pick a model, then "configures the provider, installs the gateway daemon, sets your model as the primary, and enables OpenClaw's bundled Ollama web search". Use it if you want the short path; read on if you want to know what it wrote.

If OpenClaw is not installed yet, its documented installer is one line, and the latest release listed in the OpenClaw docs as of September 19, 2026 is v2026.9.5:

```bash
curl -fsSL https://openclaw.ai/install.sh | bash
```

## Step 2: Pull a model

Pull the tag you picked, then run it once so the weights are on disk and loaded. The second command opens a prompt; type `/bye` to leave it:

```bash
ollama pull qwen3.6:27b
ollama run qwen3.6:27b
```

Pin the tag. Ollama's library lists `qwen3.6:latest` at 23GB and shows the `latest` label on the `qwen3.6:35b` row, so a bare `ollama pull qwen3.6` fetches the 23GB build rather than the 18GB `27b` one. That is a 5GB difference and, on a 24GB machine, the difference between comfortable and not.

Two commands tell you what you have and what is loaded:

```bash
ollama ls
ollama ps
```

`ollama ps` is the more useful of the two for agent work. Ollama's context length page documents its output columns and says to "Verify the split under `PROCESSOR` using `ollama ps`", because a model that spills onto the CPU is the usual cause of a local agent feeling stuck. The same output carries a `CONTEXT` column, which is how you confirm the tuning in step 4 actually took effect.

## Step 3: Point OpenClaw at Ollama

OpenClaw "reads an optional JSON5 config from `~/.openclaw/openclaw.json`", and falls back to safe defaults when the file is missing. Model refs take the form `provider/model` and live under `agents.defaults.model`, where `primary` is what the agent starts on and `fallbacks` is an ordered list it tries when the primary is unavailable.

```json5
// ~/.openclaw/openclaw.json
{
  models: {
    providers: {
      ollama: {
        baseUrl: "http://127.0.0.1:11434",
        apiKey: "ollama-local",
        api: "ollama",
        models: [
          { id: "qwen3.6:27b", name: "Qwen3.6 27B" },
        ],
      },
    },
  },
  agents: {
    defaults: {
      model: {
        primary: "ollama/qwen3.6:27b",
      },
    },
  },
}
```

Four details in that block come straight from the OpenClaw docs. The `baseUrl` carries no `/v1` suffix, because the docs warn: "Do not add `/v1`. That path selects OpenAI-compatible mode, where tool calling is not reliable." The base URL above is Ollama's default local endpoint; the docs' own examples use a remote host. `api: "ollama"` is marked in the docs as "Explicit: guarantees native tool-calling behavior", and OpenClaw talks to Ollama's native `/api/chat` rather than the OpenAI-compatible `/v1` endpoint. `apiKey: "ollama-local"` is the marker OpenClaw uses for loopback, private-network and `.local` hosts, which the docs say "do not need a real bearer token"; you can omit it when `OLLAMA_API_KEY` is set, because OpenClaw fills it in for availability checks. And listing an explicit `models` array turns discovery off: the docs annotate the same shape with "This example uses a nonempty manual model list, so it skips discovery."

If you would rather not edit the file, the documented CLI path does the same job.

```bash
export OLLAMA_API_KEY="ollama-local"
openclaw models list --provider ollama
openclaw models set ollama/qwen3.6:27b
```

Use `ollama-local` only for a loopback or LAN host; ollama.com needs the real key.

Then smoke-test the route with the command from OpenClaw's own verification recipe:

```bash
openclaw infer model run \
  --model ollama/qwen3.6:27b \
  --prompt "Reply with exactly: ok"
```

## Step 4: Tune for agent use

Three documented settings separate a local model that chats from one that runs an agent loop.

**Give it context.** Ollama's context length page sets defaults by VRAM: 4k under 24 GiB, 32k between 24 and 48 GiB, and 256k at 48 GiB and above. It also states the agent case directly: "Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens." Ollama's OpenClaw page repeats the number, recommending "a context window of at least 64k tokens if using local models". Server side, that is one environment variable:

```bash
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
```

Inside OpenClaw there are two keys and they are not interchangeable. The docs describe them plainly: "`contextTokens` caps OpenClaw's active-input budget; `params.num_ctx` sets Ollama's request context. Keep them aligned when hardware cannot run the model's full advertised context." OpenClaw's app, interactive CLI and non-interactive setup use a 32,768-token runtime context by default, or the model's native window if smaller, so raising it to the 64k Ollama recommends is a deliberate edit.

**Keep it loaded.** Ollama unloads a model after five minutes of idle by default. OpenClaw forwards `params.keep_alive` as a top-level `keep_alive` on native `/api/chat` requests, and the docs say to "set it per model when first-turn load time is the bottleneck".

**Give the first load time to finish.** `timeoutSeconds` on the provider entry is documented as covering "the model HTTP request: connection setup, headers, body streaming, and the total guarded-fetch abort", and the troubleshooting page reaches for it under the heading "Cold local model times out".

```json5
{
  models: {
    providers: {
      ollama: {
        baseUrl: "http://127.0.0.1:11434",
        apiKey: "ollama-local",
        api: "ollama",
        timeoutSeconds: 300,
        maxTokens: 8192,
        models: [
          {
            id: "qwen3.6:27b",
            name: "Qwen3.6 27B",
            contextTokens: 65536,
            params: {
              num_ctx: 65536,
              keep_alive: "15m",
            },
          },
        ],
      },
    },
  },
}
```

The 65,536 pair is the aligned version of Ollama's 64k recommendation. Ollama's own warning applies: "Setting a larger context length will increase the amount of memory required to run a model." Check the `CONTEXT` column in `ollama ps` after a turn, and drop both numbers together if the machine starts swapping.

## What you give up versus cloud models

Four trade-offs are visible in the documentation rather than in anyone's opinion.

**Context ceiling.** The largest window in the table above is 256K, on the qwen3.6, qwen3.5 and gemma4 tags. Claude Sonnet 5 lists a 1M token context window and Claude Haiku 4.5 lists 200K. The local ceiling is also softer than it looks, because context costs memory you have already spent on weights.

**Tool-calling reliability.** OpenClaw ships an escape hatch specifically for this: "If a small local model still fails on tool schemas, set `compat.supportsTools: false` on that model entry and retest." A provider that documents how to turn tools off for small local models is telling you something about small local models.

**Cold starts.** "Large local models can need a long first load" is the OpenClaw docs' own framing, and the fix is the 300 second timeout above rather than a faster model.

**Memory contention.** An 18GB model is 18GB your other applications do not get. Ollama's default context behaviour makes the same point from the other side: under 24 GiB of VRAM it hands you a 4k context, which is not enough for an agent turn.

What you get back is the part the price tables show. Claude Haiku 4.5 runs $1 per million input tokens and $5 per million output tokens, and Claude Sonnet 5 runs $2 and $10, both as listed by Anthropic on September 19, 2026. Every local token is $0 of that. Ollama's FAQ is also explicit on privacy for the local path: "Ollama runs locally. We don't see your prompts or data when you run locally."

## The hybrid setup

Run the local model as the primary and keep one cheap cloud model behind it:

```json5
{
  agents: {
    defaults: {
      model: {
        primary: "ollama/qwen3.6:27b",
        fallbacks: ["anthropic/claude-haiku-4-5"],
      },
    },
  },
}
```

Fallbacks are tried in order, and OpenClaw rotates auth profiles inside a provider before moving to the next fallback model. Note what that is and is not. The docs classify an automatic switch as "Auto fallback", a "Temporary recovery state" that OpenClaw reprobes and clears on recovery, announcing each transition once. It is failover, not a cost router, so do not expect it to hand hard turns to the cloud and easy ones to the laptop.

Invert it when availability of a specific answer matters more than the bill: put the cloud model first and the local model in `fallbacks`, and an outage or a rate limit degrades your agent to local instead of stopping it. If you want cloud capacity without a second provider account, OpenClaw's recipes cover a signed-in Ollama daemon serving both, with `primary: "ollama/gemma4"` and `fallbacks: ["ollama/kimi-k2.5:cloud"]` through one provider block.

For real cost control, split the work instead. `agents.entries.*.model` overrides `agents.defaults.model`, so a sub-agent that only summarizes an inbox can sit on the local model while your main agent stays on a cloud one. Our [cost optimization guide](/docs/cost-optimization) has the budgeting framework for deciding which is which.

## Common problems

**The `/v1` suffix.** The misconfiguration the OpenClaw docs warn about most directly. OpenClaw's Ollama page states: "Do not use the `/v1` OpenAI-compatible URL (`http://host:11434/v1`). It breaks tool calling and models can emit raw tool-call JSON as plain text. Use the native URL: `baseUrl: "http://host:11434"` (no `/v1`)." The canonical key is `baseUrl`; `baseURL` is accepted but new config should use `baseUrl`.

**Tool calls arriving as text.** Same root cause, seen from the agent side. The docs' diagnosis under "Model outputs tool JSON as text" is that "Usually the provider is in OpenAI-compatible mode, or the model cannot handle tool schemas", and the fix is to prefer native mode with `api: "ollama"`. Only if a small model still fails should you set `compat.supportsTools: false`, which disables tool use for that model entirely.

**The wrong default tag.** `ollama pull qwen3.6` gets the 23GB `latest` build, not the 18GB `27b` one. Check with `ollama ls` before blaming the machine.

**Running out of memory, or crawling.** The docs put this under "Large-context model is too slow or runs out of memory" and note that "Many models advertise contexts larger than your hardware can run comfortably." Three dials, in order: lower the model entry's `contextTokens` if OpenClaw sends too much prompt, lower `params.num_ctx` if Ollama's runtime context is too large for the machine, and lower `maxTokens` if generation runs too long.

**No models showing up.** This is the discovery rule, and it cuts both ways. "A nonempty manual model list skips discovery; an explicit self-hosted endpoint with `models: []` does not." For ambient localhost discovery, set `OLLAMA_API_KEY` or an auth profile. Discovery is also narrower than people expect: on a fresh guided setup it "considers only models already loaded in memory, as reported by `/api/ps`, with tool support and at least 16K of context confirmed by `/api/show`", and an eligible model sitting on disk but not loaded is not a candidate. If you listed models explicitly, that is the list, so add every model you want available.

## What to read next

- [The best LLM models for OpenClaw](/blog/best-llm-models-openclaw) has current provider prices and the fallback chains worth copying if you decide the hybrid setup is the right shape.
- [Running OpenClaw on a Mac mini](/blog/openclaw-mac-mini-server) covers the always-on machine most local-model users end up building, including how much memory each qwen3.6 build actually needs.
- [Cost optimization](/docs/cost-optimization) is the budgeting framework for the cloud half of a hybrid setup.

Model tags and download sizes on ollama.com move faster than this post does. Check the library page for your tag before you pull it, and check `ollama ps` after you do.

## Frequently asked questions

**Is running OpenClaw on Ollama really free?**

There is no per-token charge. The OpenClaw docs state that Ollama runs locally and is free, so all model costs are 0 for both auto-discovered and manually defined models. You still pay for the hardware and the electricity it draws, and for any cloud model you keep in the fallback chain.

**Which local model should I run on a 16GB machine?**

Pull gemma4:12b, which the Ollama library lists as a 7.6GB download with a 256K context window and tools, thinking and vision support as of September 19, 2026. Ollama's own OpenClaw integration page recommends the gemma4 family for local use at around 16 GB of VRAM. The 18GB qwen3.6:27b download does not leave room for the operating system on a 16GB machine.

**Is a local model good enough for OpenClaw agent work?**

For tool-using agent work, pick a model whose Ollama page carries the tools badge and give it a large context. Ollama's OpenClaw page recommends at least 64k tokens of context for local models, and the OpenClaw docs document a compat.supportsTools escape hatch for small models that reliably fail on tool schemas, which is a fair warning that not every local model handles them. In haimaker.ai's ranking of local models for OpenClaw, updated July 23, 2026, qwen3.6:27b placed first.

**Can I mix a local model and a cloud model in OpenClaw?**

Yes. Set agents.defaults.model.primary to your Ollama ref and agents.defaults.model.fallbacks to an ordered list that includes a cloud ref such as anthropic/claude-haiku-4-5. The OpenClaw docs describe fallbacks as tried in order, with auth-profile rotation happening inside a provider before OpenClaw moves to the next fallback model.

**Why does my local model print tool calls as plain text?**

The provider is almost certainly in OpenAI-compatible mode. The OpenClaw docs warn not to add /v1 to the Ollama baseUrl because that path selects OpenAI-compatible mode, where tool calling is not reliable and models can emit raw tool-call JSON as plain text. Use the native URL with no /v1 and set api to ollama.

## Sources

- [Download Ollama on macOS](https://ollama.com/download/mac)
- [Download Ollama on Linux](https://ollama.com/download/linux)
- [Download Ollama on Windows](https://ollama.com/download/windows)
- [Ollama docs: macOS](https://docs.ollama.com/macos)
- [Ollama docs: Linux](https://docs.ollama.com/linux)
- [Ollama docs: Windows](https://docs.ollama.com/windows)
- [Ollama docs: CLI reference](https://docs.ollama.com/cli)
- [Ollama docs: quickstart](https://docs.ollama.com/quickstart)
- [Ollama docs: context length](https://docs.ollama.com/context-length)
- [Ollama docs: FAQ](https://docs.ollama.com/faq)
- [Ollama docs: OpenClaw integration](https://docs.ollama.com/integrations/openclaw)
- [Qwen3.6 in the Ollama library](https://ollama.com/library/qwen3.6)
- [Qwen3.5 in the Ollama library](https://ollama.com/library/qwen3.5)
- [Gemma 4 in the Ollama library](https://ollama.com/library/gemma4)
- [gpt-oss in the Ollama library](https://ollama.com/library/gpt-oss)
- [Mistral in the Ollama library](https://ollama.com/library/mistral)
- [Llama 4 in the Ollama library](https://ollama.com/library/llama4)
- [OpenClaw: Install](https://docs.openclaw.ai/install)
- [OpenClaw: Configuration](https://docs.openclaw.ai/gateway/configuration)
- [OpenClaw: Models CLI](https://docs.openclaw.ai/concepts/models)
- [OpenClaw: Ollama](https://docs.openclaw.ai/providers/ollama)
- [OpenClaw: Ollama setup](https://docs.openclaw.ai/providers/ollama/setup)
- [OpenClaw: Ollama configuration](https://docs.openclaw.ai/providers/ollama/configuration)
- [OpenClaw: Ollama config recipes](https://docs.openclaw.ai/providers/ollama/recipes)
- [OpenClaw: Ollama advanced configuration](https://docs.openclaw.ai/providers/ollama/advanced)
- [OpenClaw: Ollama troubleshooting](https://docs.openclaw.ai/providers/ollama/troubleshooting)
- [OpenClaw: v2026.9.5](https://docs.openclaw.ai/releases/2026.9.5)
- [haimaker.ai: best local LLMs for OpenClaw, ranked](https://haimaker.ai/blog/best-local-models-for-openclaw/)
- [Claude API pricing](https://platform.claude.com/docs/en/about-claude/pricing)
- [Claude models overview](https://platform.claude.com/docs/en/about-claude/models/overview)