Back to Blog

Run OpenClaw 100% Free with Ollama: No API Keys, No Monthly Bills

Run OpenClaw on a local model with Ollama: which tag fits your RAM, the documented install and pull commands, and the openclaw.json provider block.

Sam Okafor14 min read

#Which local model should you run OpenClaw on?

Your memory picks the model, so start there. On 16GB, pull gemma4:12b, which Ollama's library lists as a 7.6GB download with a 256K context window and tools, thinking and vision support. On 24 to 32GB, pull qwen3.6:27b, an 18GB download with the same 256K window and the same three capability badges. On 32GB at a squeeze and 48GB comfortably, pull qwen3.6:35b, a 23GB download that wants the extra headroom a long context needs.

One ranking and one vendor recommendation sit behind these picks. haimaker.ai's ranked test of local models for OpenClaw, updated July 23, 2026, puts qwen3.6:27b first and scores models on SWE-bench Verified results, tool-calling reliability and real-world agent performance. Ollama's own OpenClaw integration page recommends the gemma4 family for local use at around 16 GB of VRAM and qwen3.5 at around 11 GB, which is what puts gemma4 in the 16GB row.

The cost side is the easy part. OpenClaw's Ollama documentation is blunt about it: "Ollama runs locally and is free, so all model costs are 0 for both auto-discovered and manually defined models." The bill you are replacing is per-token; what you keep paying is hardware and electricity.

#Local models by hardware tier

Download sizes, context windows and capability badges below are Ollama's own figures from each library page as of September 19, 2026. Capability badges are published per model family, not per tag. The memory column leaves headroom for the operating system and the KV cache rather than matching the download size to the memory size, which is why a 23GB model is not a 24GB machine's model. On Apple silicon the model shares one pool of unified memory with the system. On a PC with a discrete GPU the number that matters is VRAM, not system RAM, and a model that does not fit in VRAM spills to the CPU; confirm the split under PROCESSOR in ollama ps.

Memory for the modelModelOllama tagSize, window and badges
16GBGemma 4 12Bgemma4:12b7.6GB download, 256K window, tools, thinking and vision. Ollama's OpenClaw page recommends the gemma4 family locally at around 16 GB of VRAM; the 12b build is the one that leaves room underneath that on a 16GB machine. The 256K window is not reachable here: Ollama's VRAM-tiered default gives 4k, and Step 4 raises it to 64k.
16GBQwen3.5 9Bqwen3.5:9b6.6GB download, 256K window, tools, thinking and vision. The smaller of the two. Ollama's OpenClaw page puts the qwen3.5 family at around 11 GB of VRAM.
24 to 32GBgpt-oss 20Bgpt-oss:20b14GB download, 128K window, tools and thinking. No vision badge, so no image input.
24 to 32GBQwen3.6 27Bqwen3.6:27b18GB download, 256K window, tools, thinking and vision. First in haimaker.ai's ranking, updated July 23, 2026. The same applies here: the 24 to 48 GiB default is 32k.
24 to 32GBGemma 4 26Bgemma4:26b19GB download, 256K window, tools, thinking and vision. The tag OpenClaw's own timeout example uses. A mixture-of-experts build with 4B active parameters, so it runs faster than its size suggests, but it is a gigabyte tighter than the 27b on a 24GB machine.
32GB at a squeeze, 48GB and up comfortablyQwen3.6 35Bqwen3.6:35b23GB download, 256K window, tools, thinking and vision. Fits on 32GB; 48GB leaves room for macOS and a long context.
96GB and upgpt-oss 120Bgpt-oss:120b65GB download, 128K window, tools and thinking.

Two models worth naming and then skipping. mistral:7b is only a 4.4GB download, but Ollama's library lists its context window at 32K, below the 64k tokens Ollama's OpenClaw page recommends for local agent use. llama4 starts at a 67GB download for the 16x17b tag, so its 10M context window is not a consumer-hardware option.

Ollama also publishes -mlx variants of several of these tags for Apple silicon, at slightly different download sizes. qwen3.6:27b-mlx is 19GB against 18GB for qwen3.6:27b. Pick one and pin it rather than letting a default tag decide.

Want step-by-step guides for this and more?

ClawDocx Pro includes 500+ curated prompts, setup guides, SKILL.md files, and templates — everything to make your AI agent unstoppable.

See plans & pricing

#Step 1: Install Ollama

Ollama publishes a download page per platform with the command on it. On macOS and Linux it is the same script:

bash
curl -fsSL https://ollama.com/install.sh | sh

On Windows, Ollama's download page gives a PowerShell one-liner instead:

powershell
irm https://ollama.com/install.ps1 | iex

Each platform has a documented floor. Ollama's macOS page requires "MacOS Sonoma (v14) or newer" and Apple M series silicon for GPU support, with x86 Macs limited to CPU. The Windows page requires "Windows 10 22H2 or newer, Home or Pro", runs as a native Windows application with NVIDIA and AMD Radeon GPU support, and notes that the API is served on http://localhost:11434. Ollama binds 127.0.0.1 on port 11434 by default, which is the address the config block below uses.

The macOS app and the Linux systemd service start the server for you. Check that it answers:

bash
ollama -v
curl http://127.0.0.1:11434/api/tags

If nothing answers, start the server yourself and leave it running in its own terminal:

bash
ollama serve

Ollama ships a shortcut for this entire post. ollama launch openclaw prompts to install OpenClaw through npm if it is missing, shows a security notice about tool access on first launch, lets you pick a model, then "configures the provider, installs the gateway daemon, sets your model as the primary, and enables OpenClaw's bundled Ollama web search". Use it if you want the short path; read on if you want to know what it wrote.

If OpenClaw is not installed yet, its documented installer is one line, and the latest release listed in the OpenClaw docs as of September 19, 2026 is v2026.9.5:

bash
curl -fsSL https://openclaw.ai/install.sh | bash

#Step 2: Pull a model

Pull the tag you picked, then run it once so the weights are on disk and loaded. The second command opens a prompt; type /bye to leave it:

bash
ollama pull qwen3.6:27b
ollama run qwen3.6:27b

Pin the tag. Ollama's library lists qwen3.6:latest at 23GB and shows the latest label on the qwen3.6:35b row, so a bare ollama pull qwen3.6 fetches the 23GB build rather than the 18GB 27b one. That is a 5GB difference and, on a 24GB machine, the difference between comfortable and not.

Two commands tell you what you have and what is loaded:

bash
ollama ls
ollama ps

ollama ps is the more useful of the two for agent work. Ollama's context length page documents its output columns and says to "Verify the split under PROCESSOR using ollama ps", because a model that spills onto the CPU is the usual cause of a local agent feeling stuck. The same output carries a CONTEXT column, which is how you confirm the tuning in step 4 actually took effect.

#Step 3: Point OpenClaw at Ollama

OpenClaw "reads an optional JSON5 config from ~/.openclaw/openclaw.json", and falls back to safe defaults when the file is missing. Model refs take the form provider/model and live under agents.defaults.model, where primary is what the agent starts on and fallbacks is an ordered list it tries when the primary is unavailable.

json5
// ~/.openclaw/openclaw.json
{
models: {
providers: {
ollama: {
baseUrl: "http://127.0.0.1:11434",
apiKey: "ollama-local",
api: "ollama",
models: [
{ id: "qwen3.6:27b", name: "Qwen3.6 27B" },
],
},
},
},
agents: {
defaults: {
model: {
primary: "ollama/qwen3.6:27b",
},
},
},
}

Four details in that block come straight from the OpenClaw docs. The baseUrl carries no /v1 suffix, because the docs warn: "Do not add /v1. That path selects OpenAI-compatible mode, where tool calling is not reliable." The base URL above is Ollama's default local endpoint; the docs' own examples use a remote host. api: "ollama" is marked in the docs as "Explicit: guarantees native tool-calling behavior", and OpenClaw talks to Ollama's native /api/chat rather than the OpenAI-compatible /v1 endpoint. apiKey: "ollama-local" is the marker OpenClaw uses for loopback, private-network and .local hosts, which the docs say "do not need a real bearer token"; you can omit it when OLLAMA_API_KEY is set, because OpenClaw fills it in for availability checks. And listing an explicit models array turns discovery off: the docs annotate the same shape with "This example uses a nonempty manual model list, so it skips discovery."

If you would rather not edit the file, the documented CLI path does the same job.

bash
export OLLAMA_API_KEY="ollama-local"
openclaw models list --provider ollama
openclaw models set ollama/qwen3.6:27b

Use ollama-local only for a loopback or LAN host; ollama.com needs the real key.

Then smoke-test the route with the command from OpenClaw's own verification recipe:

bash
openclaw infer model run \
--model ollama/qwen3.6:27b \
--prompt "Reply with exactly: ok"

#Step 4: Tune for agent use

Three documented settings separate a local model that chats from one that runs an agent loop.

Give it context. Ollama's context length page sets defaults by VRAM: 4k under 24 GiB, 32k between 24 and 48 GiB, and 256k at 48 GiB and above. It also states the agent case directly: "Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens." Ollama's OpenClaw page repeats the number, recommending "a context window of at least 64k tokens if using local models". Server side, that is one environment variable:

bash
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

Inside OpenClaw there are two keys and they are not interchangeable. The docs describe them plainly: "contextTokens caps OpenClaw's active-input budget; params.num_ctx sets Ollama's request context. Keep them aligned when hardware cannot run the model's full advertised context." OpenClaw's app, interactive CLI and non-interactive setup use a 32,768-token runtime context by default, or the model's native window if smaller, so raising it to the 64k Ollama recommends is a deliberate edit.

Keep it loaded. Ollama unloads a model after five minutes of idle by default. OpenClaw forwards params.keep_alive as a top-level keep_alive on native /api/chat requests, and the docs say to "set it per model when first-turn load time is the bottleneck".

Give the first load time to finish. timeoutSeconds on the provider entry is documented as covering "the model HTTP request: connection setup, headers, body streaming, and the total guarded-fetch abort", and the troubleshooting page reaches for it under the heading "Cold local model times out".

json5
{
models: {
providers: {
ollama: {
baseUrl: "http://127.0.0.1:11434",
apiKey: "ollama-local",
api: "ollama",
timeoutSeconds: 300,
maxTokens: 8192,
models: [
{
id: "qwen3.6:27b",
name: "Qwen3.6 27B",
contextTokens: 65536,
params: {
num_ctx: 65536,
keep_alive: "15m",
},
},
],
},
},
},
}

The 65,536 pair is the aligned version of Ollama's 64k recommendation. Ollama's own warning applies: "Setting a larger context length will increase the amount of memory required to run a model." Check the CONTEXT column in ollama ps after a turn, and drop both numbers together if the machine starts swapping.

#What you give up versus cloud models

Four trade-offs are visible in the documentation rather than in anyone's opinion.

Context ceiling. The largest window in the table above is 256K, on the qwen3.6, qwen3.5 and gemma4 tags. Claude Sonnet 5 lists a 1M token context window and Claude Haiku 4.5 lists 200K. The local ceiling is also softer than it looks, because context costs memory you have already spent on weights.

Tool-calling reliability. OpenClaw ships an escape hatch specifically for this: "If a small local model still fails on tool schemas, set compat.supportsTools: false on that model entry and retest." A provider that documents how to turn tools off for small local models is telling you something about small local models.

Cold starts. "Large local models can need a long first load" is the OpenClaw docs' own framing, and the fix is the 300 second timeout above rather than a faster model.

Memory contention. An 18GB model is 18GB your other applications do not get. Ollama's default context behaviour makes the same point from the other side: under 24 GiB of VRAM it hands you a 4k context, which is not enough for an agent turn.

What you get back is the part the price tables show. Claude Haiku 4.5 runs $1 per million input tokens and $5 per million output tokens, and Claude Sonnet 5 runs $2 and $10, both as listed by Anthropic on September 19, 2026. Every local token is $0 of that. Ollama's FAQ is also explicit on privacy for the local path: "Ollama runs locally. We don't see your prompts or data when you run locally."

#The hybrid setup

Run the local model as the primary and keep one cheap cloud model behind it:

json5
{
agents: {
defaults: {
model: {
primary: "ollama/qwen3.6:27b",
fallbacks: ["anthropic/claude-haiku-4-5"],
},
},
},
}

Fallbacks are tried in order, and OpenClaw rotates auth profiles inside a provider before moving to the next fallback model. Note what that is and is not. The docs classify an automatic switch as "Auto fallback", a "Temporary recovery state" that OpenClaw reprobes and clears on recovery, announcing each transition once. It is failover, not a cost router, so do not expect it to hand hard turns to the cloud and easy ones to the laptop.

Invert it when availability of a specific answer matters more than the bill: put the cloud model first and the local model in fallbacks, and an outage or a rate limit degrades your agent to local instead of stopping it. If you want cloud capacity without a second provider account, OpenClaw's recipes cover a signed-in Ollama daemon serving both, with primary: "ollama/gemma4" and fallbacks: ["ollama/kimi-k2.5:cloud"] through one provider block.

For real cost control, split the work instead. agents.entries.*.model overrides agents.defaults.model, so a sub-agent that only summarizes an inbox can sit on the local model while your main agent stays on a cloud one. Our cost optimization guide has the budgeting framework for deciding which is which.

#Common problems

The /v1 suffix. The misconfiguration the OpenClaw docs warn about most directly. OpenClaw's Ollama page states: "Do not use the /v1 OpenAI-compatible URL (http://host:11434/v1). It breaks tool calling and models can emit raw tool-call JSON as plain text. Use the native URL: baseUrl: "http://host:11434" (no /v1)." The canonical key is baseUrl; baseURL is accepted but new config should use baseUrl.

Tool calls arriving as text. Same root cause, seen from the agent side. The docs' diagnosis under "Model outputs tool JSON as text" is that "Usually the provider is in OpenAI-compatible mode, or the model cannot handle tool schemas", and the fix is to prefer native mode with api: "ollama". Only if a small model still fails should you set compat.supportsTools: false, which disables tool use for that model entirely.

The wrong default tag. ollama pull qwen3.6 gets the 23GB latest build, not the 18GB 27b one. Check with ollama ls before blaming the machine.

Running out of memory, or crawling. The docs put this under "Large-context model is too slow or runs out of memory" and note that "Many models advertise contexts larger than your hardware can run comfortably." Three dials, in order: lower the model entry's contextTokens if OpenClaw sends too much prompt, lower params.num_ctx if Ollama's runtime context is too large for the machine, and lower maxTokens if generation runs too long.

No models showing up. This is the discovery rule, and it cuts both ways. "A nonempty manual model list skips discovery; an explicit self-hosted endpoint with models: [] does not." For ambient localhost discovery, set OLLAMA_API_KEY or an auth profile. Discovery is also narrower than people expect: on a fresh guided setup it "considers only models already loaded in memory, as reported by /api/ps, with tool support and at least 16K of context confirmed by /api/show", and an eligible model sitting on disk but not loaded is not a candidate. If you listed models explicitly, that is the list, so add every model you want available.

  • The best LLM models for OpenClaw has current provider prices and the fallback chains worth copying if you decide the hybrid setup is the right shape.
  • Running OpenClaw on a Mac mini covers the always-on machine most local-model users end up building, including how much memory each qwen3.6 build actually needs.
  • Cost optimization is the budgeting framework for the cloud half of a hybrid setup.

Model tags and download sizes on ollama.com move faster than this post does. Check the library page for your tag before you pull it, and check ollama ps after you do.

Frequently asked questions

Is running OpenClaw on Ollama really free?
There is no per-token charge. The OpenClaw docs state that Ollama runs locally and is free, so all model costs are 0 for both auto-discovered and manually defined models. You still pay for the hardware and the electricity it draws, and for any cloud model you keep in the fallback chain.
Which local model should I run on a 16GB machine?
Pull gemma4:12b, which the Ollama library lists as a 7.6GB download with a 256K context window and tools, thinking and vision support as of September 19, 2026. Ollama's own OpenClaw integration page recommends the gemma4 family for local use at around 16 GB of VRAM. The 18GB qwen3.6:27b download does not leave room for the operating system on a 16GB machine.
Is a local model good enough for OpenClaw agent work?
For tool-using agent work, pick a model whose Ollama page carries the tools badge and give it a large context. Ollama's OpenClaw page recommends at least 64k tokens of context for local models, and the OpenClaw docs document a compat.supportsTools escape hatch for small models that reliably fail on tool schemas, which is a fair warning that not every local model handles them. In haimaker.ai's ranking of local models for OpenClaw, updated July 23, 2026, qwen3.6:27b placed first.
Can I mix a local model and a cloud model in OpenClaw?
Yes. Set agents.defaults.model.primary to your Ollama ref and agents.defaults.model.fallbacks to an ordered list that includes a cloud ref such as anthropic/claude-haiku-4-5. The OpenClaw docs describe fallbacks as tried in order, with auth-profile rotation happening inside a provider before OpenClaw moves to the next fallback model.
Why does my local model print tool calls as plain text?
The provider is almost certainly in OpenAI-compatible mode. The OpenClaw docs warn not to add /v1 to the Ollama baseUrl because that path selects OpenAI-compatible mode, where tool calling is not reliable and models can emit raw tool-call JSON as plain text. Use the native URL with no /v1 and set api to ollama.

Sources

  1. Download Ollama on macOS
  2. Download Ollama on Linux
  3. Download Ollama on Windows
  4. Ollama docs: macOS
  5. Ollama docs: Linux
  6. Ollama docs: Windows
  7. Ollama docs: CLI reference
  8. Ollama docs: quickstart
  9. Ollama docs: context length
  10. Ollama docs: FAQ
  11. Ollama docs: OpenClaw integration
  12. Qwen3.6 in the Ollama library
  13. Qwen3.5 in the Ollama library
  14. Gemma 4 in the Ollama library
  15. gpt-oss in the Ollama library
  16. Mistral in the Ollama library
  17. Llama 4 in the Ollama library
  18. OpenClaw: Install
  19. OpenClaw: Configuration
  20. OpenClaw: Models CLI
  21. OpenClaw: Ollama
  22. OpenClaw: Ollama setup
  23. OpenClaw: Ollama configuration
  24. OpenClaw: Ollama config recipes
  25. OpenClaw: Ollama advanced configuration
  26. OpenClaw: Ollama troubleshooting
  27. OpenClaw: v2026.9.5
  28. haimaker.ai: best local LLMs for OpenClaw, ranked
  29. Claude API pricing
  30. Claude models overview

Get the full experience with ClawDocx Pro

Access 500+ prompts, step-by-step guides, SKILL.md files, and more. Everything you need to master OpenClaw.

Start Free Trial

Related Posts