Skip to content
The 31 guidesFREN中文
Guides
Part 4 · guide 7 of 8 Level: Easy Reading time: 12 min Platforms: Linux

Ollama & local models

Run real AI models at home, free and private. Installation, your first model, and how to wire it up to your coding agent.

In this guide
  1. 01Installing Ollama
  2. 02Your first model
  3. 03Wiring it up to your coding agent
  4. 04The three settings that matter
  5. 05The trap to burn into your memory
  6. 06Frequently asked questions

In short

Ollama runs open AI models on your own machine: one command to install it, one to start a model, and a local API on port 11434 that speaks both the OpenAI and Anthropic formats. To start, qwen3.5:9b fits in 16 GB and qwen3.8:27b needs a 24 GB graphics card or 32 GB of unified memory. OpenCode plugs into it natively, so does Claude Code since Ollama 0.15 with ollama launch claude, and Codex with ollama launch codex or codex --oss. Set the context size yourself, because the default depends on memory and climbs to 256,000 tokens on big machines.

Do this first: Choosing and sizing your model

Up to now, your coding agent was talking to models in the cloud. Now we run real AI models directly on your mini-PC. And the thing that makes that painless is Ollama.

Ollama is the simplest way to run open-weight LLMs locally. One command to install it, one to download a model, and you end up with a local API on port 11434 that speaks the OpenAI format and, since Ollama 0.14 (January 2026), Anthropic’s too: just about every tool out there can plug into it. Three advantages that need no comment: it’s private (nothing leaves the machine), it’s free, and it works offline.

Installing Ollama

One line. The script installs the binary and launches it as a system service, it runs in the background, ready to answer.

curl -fsSL https://ollama.com/install.sh | sh

That’s it. Ollama is now listening on http://localhost:11434.

Your first model

Let’s start with a model that fits almost anywhere: qwen3.5:9b, 9 billion parameters, a 6.6 GB download. In the Quelle IA ranking (in French, 28 September 2026 edition), it’s the top-rated model that fits in 8 to 14 GB of usable memory, so the right pick for a 16 GB machine.

# the starting point, for a 16 GB machine
ollama run qwen3.5:9b
# with a 24 GB graphics card, or 32 GB of unified memory and up
ollama run qwen3.8:27b

qwen3.8:27b (18 GB) is, as of the same date, the top-rated model that fits anywhere from a 24 GB graphics card up to a 128 GB mini-PC. These picks move fast: for yours, use Find my model or the local models ranking on Quelle IA (in French). The details are in Choosing your local model.

The first launch downloads the model (a few GB, be patient). Then you land in a chat right in the terminal: ask it a question, ask it for a snippet of code, check that it answers. Type /bye to exit.

0 of 3 steps done Your ticks stay in this browser.

  1. See what you have

    ollama list   # every downloaded model, with its size
    
  2. See what's running

    ollama ps     # the models loaded in memory right now
    
  3. Clean house

    ollama rm qwen3.5:9b   # delete a model to free up space
    

Wiring it up to your coding agent

All three agents plug into Ollama, each by its own route. They don’t all offer the same online models: Quelle IA’s coding tools page (in French) shows what each one accepts, personal keys and local models included. Whichever you pick, keep in mind that a local model still sits below the best online models for code: check the gap on the Quelle IA code ranking (in French) before expecting miracles.

Claude Code plugs into Ollama. Since version 0.14, Ollama speaks Anthropic’s API, and since 0.15 a single command does all the wiring:

ollama launch claude                       # picks a model and starts Claude Code on it
ollama launch claude --model qwen3.8:27b   # or straight away with the model you want

Without ollama launch, three variables are enough, for the length of a session:

export ANTHROPIC_AUTH_TOKEN=ollama        # required, but ignored by Ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
claude --model qwen3.8:27b

Pick a model that can call tools, and raise the context to at least 64,000 tokens: Ollama’s docs recommend it for Claude Code. Also expect the gap in level: same harness, smaller brain. For the big jobs, long reasoning and top-tier reliability, Claude Code on Anthropic’s models stays ahead.

Ollama’s local API also serves side tasks, without switching models in Claude Code: embeddings, quick classification, summaries, small scripts that hit http://localhost:11434 without costing a cent or sending your data elsewhere. The cloud for the brain, local for the plumbing.

The three settings that matter

Ollama works right out of the box, but three environment variables make all the difference when you push it. You set them in the service’s environment (systemctl edit ollama then Environment="...").

The trap to burn into your memory

Nine times out of ten, when someone says “Ollama is slow on my machine,” the culprit is the context set too high. And on a well-equipped machine, it’s now the default that sets it too high.

Here’s why. The bigger the context window, the more the KV cache (the model’s working memory) eats RAM, and it grows fast. If you ask for a huge context on a tight machine, you spill over into swap (the disk used as backup RAM), and then everything collapses: the model crawls, every token takes forever, you think your hardware is junk when it’s actually suffocating.

The rule: set the context to what the task needs, not to the maximum. A small script? 4096 is plenty. A coding agent on a real repo? 64,000, but watch your RAM: ollama ps shows the size actually loaded and, in its CONTEXT column, the context it settled on. The right setting is the smallest one that does the job.

Frequently asked questions

Why is Ollama so slow on my machine?

Nine times out of ten, the context window is set too high. The bigger it is, the more RAM the KV cache (the model's working memory) eats, and on a tight machine everything spills over into swap: each token then takes forever. Set the context to what the task needs, and check with ollama ps how much is really loaded and which context was kept.

How do I stop Ollama from reloading the model on every request?

By default, Ollama keeps a model loaded for 5 minutes after the last request. With OLLAMA_KEEP_ALIVE=-1 it stays resident permanently, which becomes essential for an agent that fires off many calls. You set the variable in the service's environment, with systemctl edit ollama.

Can I use Ollama from another device?

Yes, as long as it stays private: leave the OLLAMA_HOST listening address on localhost, or go through Tailscale to reach it from another device. Never expose port 11434 to the public internet, that would hand your GPU to anyone who comes along.

How do I see and delete downloaded models?

The ollama list command shows every downloaded model with its size, and ollama ps shows the ones loaded in memory right now. To free up space, ollama rm followed by the model name deletes it.

Does a local model code as well as Claude or GPT?

No, a local model stays below the best online models for code: with Claude Code or Codex plugged into it, the harness is the same and the model is more modest. It has other strengths: it is private, free and works offline. It also shines on side tasks such as embeddings, quick classification or summaries.

Terms in this guide: AgentOllamaOpen-weight modelAPIParametersGPU (graphics card)Context windowRAMTokenRepository (repo)

Spotted a mistake?

A command stopped working, a price changed?

Tools change every month. Tell me what is wrong in this chapter and I will fix it and update its date.

Only the page, your message and the optional contact are kept. Nothing else.

Guide 25 of 31 · part 4 no guides read yet Open the list of guides