Ollama & local models
Run real AI models at home, free and private. Installation, your first model, and how to wire it up to your coding agent.
In this guide
Guide checked 3 months ago: some commands may have changed. Let us know if so.
In short
Ollama runs open AI models on your own machine: one command to install it, one to start a model, and a local API on port 11434 that speaks both the OpenAI and Anthropic formats. To start, qwen3.5:9b fits in 16 GB and qwen3.8:27b needs a 24 GB graphics card or 32 GB of unified memory. OpenCode plugs into it natively, so does Claude Code since Ollama 0.15 with ollama launch claude, and Codex with ollama launch codex or codex --oss. Set the context size yourself, because the default depends on memory and climbs to 256,000 tokens on big machines.
Do this first: Choosing and sizing your model
Up to now, your coding agent was talking to models in the cloud. Now we run real AI models directly on your mini-PC. And the thing that makes that painless is Ollama.
Ollama is the simplest way to run open-weight LLMs locally. One command to install it, one to download a model, and you end up with a local API on port 11434 that speaks the OpenAI format and, since Ollama 0.14 (January 2026), Anthropic’s too: just about every tool out there can plug into it. Three advantages that need no comment: it’s private (nothing leaves the machine), it’s free, and it works offline.
Installing Ollama
One line. The script installs the binary and launches it as a system service, it runs in the background, ready to answer.
curl -fsSL https://ollama.com/install.sh | sh
That’s it. Ollama is now listening on http://localhost:11434.
Your first model
Let’s start with a model that fits almost anywhere: qwen3.5:9b, 9 billion parameters, a 6.6 GB download. In the Quelle IA ranking (in French, 28 September 2026 edition), it’s the top-rated model that fits in 8 to 14 GB of usable memory, so the right pick for a 16 GB machine.
# the starting point, for a 16 GB machine
ollama run qwen3.5:9b
# with a 24 GB graphics card, or 32 GB of unified memory and up
ollama run qwen3.8:27b
qwen3.8:27b (18 GB) is, as of the same date, the top-rated model that fits anywhere from a 24 GB graphics card up to a 128 GB mini-PC. These picks move fast: for yours, use Find my model or the local models ranking on Quelle IA (in French). The details are in Choosing your local model.
The first launch downloads the model (a few GB, be patient). Then you land in a chat right in the terminal: ask it a question, ask it for a snippet of code, check that it answers. Type /bye to exit.
0 of 3 steps done Your ticks stay in this browser.
-
See what you have
ollama list # every downloaded model, with its size -
See what's running
ollama ps # the models loaded in memory right now -
Clean house
ollama rm qwen3.5:9b # delete a model to free up space
Wiring it up to your coding agent
All three agents plug into Ollama, each by its own route. They don’t all offer the same online models: Quelle IA’s coding tools page (in French) shows what each one accepts, personal keys and local models included. Whichever you pick, keep in mind that a local model still sits below the best online models for code: check the gap on the Quelle IA code ranking (in French) before expecting miracles.
Claude Code plugs into Ollama. Since version 0.14, Ollama speaks Anthropic’s API, and since 0.15 a single command does all the wiring:
ollama launch claude # picks a model and starts Claude Code on it
ollama launch claude --model qwen3.8:27b # or straight away with the model you want
Without ollama launch, three variables are enough, for the length of a session:
export ANTHROPIC_AUTH_TOKEN=ollama # required, but ignored by Ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
claude --model qwen3.8:27b
Pick a model that can call tools, and raise the context to at least 64,000 tokens: Ollama’s docs recommend it for Claude Code. Also expect the gap in level: same harness, smaller brain. For the big jobs, long reasoning and top-tier reliability, Claude Code on Anthropic’s models stays ahead.
Ollama’s local API also serves side tasks, without switching models in Claude Code: embeddings, quick classification, summaries, small scripts that hit http://localhost:11434 without costing a cent or sending your data elsewhere. The cloud for the brain, local for the plumbing.
This is the royal road to a 100% local coding agent. OpenCode talks natively to Ollama. The shortest path, since Ollama 0.15:
ollama launch opencode # picks the model and starts OpenCode, without touching your config
By hand, you declare Ollama as a provider in opencode.json, pointing it at the local API:
http://localhost:11434/v1
then you pick the local model with /models, for example qwen3.8:27b. OpenCode needs a context of at least 64,000 tokens (see the settings below). From there, your agent thinks, reads your code, and writes files without a single byte leaving the machine. Free, private, offline. This is exactly the scenario OpenCode exists for.
Codex knows Ollama out of the box: it’s one of its built-in “open source” providers, switched on with --oss. The shortest route goes through Ollama itself:
ollama launch codex # Ollama sets up a dedicated Codex profile and starts the session
By hand, one flag is enough:
codex --oss -m qwen3.8:27b # interactive session on the local model
codex exec --oss -m qwen3.8:27b "Summarize the README" # same thing, non-interactive
To avoid repeating the provider choice, set oss_provider = "ollama" in ~/.codex/config.toml: without it, the interactive session asks you to choose between Ollama and LM Studio, and codex exec --oss stops with an error. Here too, Ollama’s docs ask for a context of at least 64,000 tokens. Same caveat as for Claude Code: same harness, a smaller model than OpenAI’s.
The three settings that matter
Ollama works right out of the box, but three environment variables make all the difference when you push it. You set them in the service’s environment (systemctl edit ollama then Environment="...").
The trap to burn into your memory
Nine times out of ten, when someone says “Ollama is slow on my machine,” the culprit is the context set too high. And on a well-equipped machine, it’s now the default that sets it too high.
Here’s why. The bigger the context window, the more the KV cache (the model’s working memory) eats RAM, and it grows fast. If you ask for a huge context on a tight machine, you spill over into swap (the disk used as backup RAM), and then everything collapses: the model crawls, every token takes forever, you think your hardware is junk when it’s actually suffocating.
The rule: set the context to what the task needs, not to the maximum. A small script? 4096 is plenty. A coding agent on a real repo? 64,000, but watch your RAM: ollama ps shows the size actually loaded and, in its CONTEXT column, the context it settled on. The right setting is the smallest one that does the job.
All the commands in this guide
Frequently asked questions
Why is Ollama so slow on my machine?
Nine times out of ten, the context window is set too high. The bigger it is, the more RAM the KV cache (the model's working memory) eats, and on a tight machine everything spills over into swap: each token then takes forever. Set the context to what the task needs, and check with ollama ps how much is really loaded and which context was kept.
How do I stop Ollama from reloading the model on every request?
By default, Ollama keeps a model loaded for 5 minutes after the last request. With OLLAMA_KEEP_ALIVE=-1 it stays resident permanently, which becomes essential for an agent that fires off many calls. You set the variable in the service's environment, with systemctl edit ollama.
Can I use Ollama from another device?
Yes, as long as it stays private: leave the OLLAMA_HOST listening address on localhost, or go through Tailscale to reach it from another device. Never expose port 11434 to the public internet, that would hand your GPU to anyone who comes along.
How do I see and delete downloaded models?
The ollama list command shows every downloaded model with its size, and ollama ps shows the ones loaded in memory right now. To free up space, ollama rm followed by the model name deletes it.
Does a local model code as well as Claude or GPT?
No, a local model stays below the best online models for code: with Claude Code or Codex plugged into it, the harness is the same and the model is more modest. It has other strengths: it is private, free and works offline. It also shines on side tasks such as embeddings, quick classification or summaries.
Terms in this guide: AgentOllamaOpen-weight modelAPIParametersGPU (graphics card)Context windowRAMTokenRepository (repo)
Spotted a mistake?
A command stopped working, a price changed?
Tools change every month. Tell me what is wrong in this chapter and I will fix it and update its date.