Papr logoPapr

Guide · Updated September 2026

How to run Claude Code free with local models (Ollama)

Claude Code is a free download, but normally every request goes to Anthropic and needs a Claude plan or an API key. Ollama can stand in for Anthropic's API, so Claude Code talks to a model running on your own Mac or PC instead. No subscription, no per-token bill, and your code stays on your machine.

The trade-off is quality. This guide covers the setup on Mac and Windows, which model to pick for how much memory you have, and what to expect once it is running.

What you need

  • Ollama installed and running.
  • At least 16 GB of memory. 32 GB or more is where it starts to feel useful.
  • On Windows, a GPU with 8 GB+ of video memory helps a lot. On a Mac with Apple silicon, system memory is shared, so total RAM is what counts.

The quick way

Recent versions of Ollama can set Claude Code up for you. Run this, pick a model, and it installs what is missing and starts Claude Code pointed at Ollama:

ollama launch claude

If that works, skip to choosing a model.

Manual setup on Mac

If you want to control it yourself, or the quick way is not available in your Ollama version, set three environment variables so Claude Code sends requests to Ollama:

# 1. Install Claude Code
curl -fsSL https://claude.ai/install.sh | bash

# 2. Download a model (pick one from the table below)
ollama pull qwen3.5

# 3. Point Claude Code at Ollama instead of Anthropic
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434

# 4. Start it
claude --model qwen3.5

Add the three export lines to ~/.zshrc if you want them every time. Leave them out if you also use Claude Code with your normal Claude account.

Manual setup on Windows

Same idea in PowerShell:

# 1. Install Claude Code (PowerShell)
irm https://claude.ai/install.ps1 | iex

# 2. Download a model
ollama pull qwen3.5

# 3. Point Claude Code at Ollama
$env:ANTHROPIC_AUTH_TOKEN = "ollama"
$env:ANTHROPIC_API_KEY = ""
$env:ANTHROPIC_BASE_URL = "http://localhost:11434"

# 4. Start it
claude --model qwen3.5

Raise the context window (don't skip this)

This is the most common reason local Claude Code feels broken. Ollama gives models only 4k tokens of context on machines with under 24 GB of GPU memory. Claude Code needs far more to hold your files and its own instructions. Set it to at least 64k, either with the slider in the Ollama app settings or when you start the server:

OLLAMA_CONTEXT_LENGTH=64000 ollama serve

Run ollama ps to check. The CONTEXT column should show 64000 or more, and PROCESSOR should say 100% GPU. If part of the model spills onto the CPU, drop to a smaller model.

Which model for how much memory

MemoryModelWhat to expect
8 GBqwen3.5:2bRuns, but too small for real coding. Fine for questions about a file.
16 GBqwen3.5 or gemma4:e4bSmall edits and explanations. Expect it to lose track on multi-file changes.
32 GBqwen3.5:27bThe first size that feels usable for everyday coding tasks.
48 GB+qwen3.5:27b with 64k+ contextRoom for a long context, which matters more than raw model size for Claude Code.

Use a model that supports tool calling. Without it, Claude Code can chat but cannot edit files or run commands.

What to expect

Local models handle explaining code, writing a single function, small fixes and questions about a repo reasonably well. They struggle with the things Claude Code is best known for: long sessions that touch many files, keeping a plan straight over dozens of steps, and recovering from their own mistakes.

A setup many people land on: a local model for quick, private or offline work, and a Claude plan or API key for the big tasks. Switching is just a matter of whether the three environment variables are set.

If you don't want the terminal

Plenty of people searching for a free local Claude Code are not trying to edit a codebase. They want an AI that builds things for them: a dashboard, a scraper, a weekly report, a tool for their team.

That is what Papr Work is for. It is a desktop app for Mac and Windows that installs Ollama and a Qwen model for you when you pick a local model, then builds apps and scheduled jobs that run on your computer. Running locally is free and unlimited. It is not a code editor, so for working inside an existing repository, stick with Claude Code.

Download Papr Work

Frequently asked questions

Can I use Claude Code for free?

Yes, if you run it against a local model. Claude Code itself is a free download. Pointed at Ollama, it sends requests to a model on your own computer instead of Anthropic, so you do not need a Claude subscription or an API key. You still need a computer with enough memory to run the model.

Is Claude Code with a local model as good as with Claude?

No. Local open models are much better than they were a year ago, but on long, multi-step coding tasks they still make more mistakes than Claude Sonnet or Opus. They are good for small edits, explaining code and working offline.

What is the best local model for Claude Code?

Pick the largest model your memory allows that supports tool calling, and give it at least 64k of context. On a 32 GB machine, qwen3.5:27b is a good default. On 16 GB, qwen3.5 or gemma4:e4b will run but struggle with bigger tasks.

Does it work on Windows?

Yes. Install Ollama and Claude Code for Windows, then set the same three environment variables in PowerShell. A dedicated GPU helps a lot; on CPU only, responses are slow.

Why does Claude Code keep forgetting what it was doing?

Usually the context window is too small. Ollama defaults to 4k tokens on machines with under 24 GB of GPU memory, which is far too little for Claude Code. Raise it to 64k in the Ollama app settings or with OLLAMA_CONTEXT_LENGTH.

Is there a local Claude Code alternative that does not need the terminal?

If you want to build apps, automations and reports rather than edit a codebase, Papr Work is a desktop app that installs Ollama and a model for you and runs everything on your computer for free. For editing an existing codebase, Claude Code or another coding agent is the better tool.

Commands from Ollama's Claude Code integration docs, September 2026.