Skip to content
The 31 guidesFREN中文
Guides
Part 4 · guide 6 of 8 Level: Intermediate Reading time: 16 min Platforms: Linux, Mac and Windows

Choosing and sizing your model

Which local model should you run on YOUR machine? The 2026 open-weight landscape (mostly Chinese), and the simple rule for fitting them to your RAM.

In this guide
  1. 01The sizing rule (the one that decides everything)
  2. 02The shortcut: let a tool size it for you
  3. 03The 2026 model landscape
  4. 04Quantization, plainly
  5. 05Wiring the right quant into Ollama
  6. 06Honestly: local or Claude?
  7. 07Frequently asked questions

In short

Your machine's memory decides: count about 0.6 GB per billion parameters at 4 bits, plus the context. According to Quelle IA (28 September 2026 edition), Qwen3.5-9B is the best pick up to 16 GB, and Qwen 3.8 27B the best from a 24 GB graphics card up to a 128 GB mini-PC, with Gemma 4 26B A4B as the faster option and gpt-oss-120b for agents. The largest open models, such as GLM-5.3-Flash or DeepSeek V4.1 Flash, need more than 200 GB. For code, the best local model still sits well below Claude: local for bounded, private tasks, the cloud for long jobs.

Do this first: Choosing the hardware

Here’s the truth nobody likes to hear: you don’t choose your model, your machine does. You can dream of running the biggest open-weight monster on the planet, but if you have 32 GB of RAM, it’ll never launch. The good news is that in 2026 the local landscape has become excellent, as long as you aim right. We’ll first lay down the rule that decides everything (sizing), then look at the catalog, and finally settle the real question: when your mini-PC is enough, and when Claude stays in front.

The sizing rule (the one that decides everything)

Before model names, physics. A model is billions of parameters (the “weights”), and each one takes up space in memory. The mental formula fits on one line:

RAM needed ≈ (number of parameters × bytes per weight) + the context cache (the KV cache)

The bytes per weight depend on precision (we’ll come back to that with quantization). In practice, in Q4_K_M, the standard format, count ~0.6 GB per billion parameters (the rule Quelle IA follows, explained in its “Comment on compte” section, in French). Add 0.5 to 2 GB for the system, and +30 to 50% if you work with a long context or several parallel requests.

A few reference points in Q4_K_M, to train your eye: a Llama 8B fits in ~6 GB, a dense 32B model in ~19 GB, a Llama 70B demands 40-43 GB. And here’s the table to keep handy, by memory tier. These are Quelle IA’s picks in the 28 September 2026 edition (in French); they change often, so check your machine’s page before downloading 20 GB.

Machine memoryTop-rated model that fitsIn Ollama
8 GB graphics cardQwen3.5-4Bqwen3.5:4b
16 GB (graphics card, Mac, laptop)Qwen3.5-9Bqwen3.5:9b
24 GB card, 32 GB Mac or Ryzen AI MaxQwen 3.8 27B at 4 bits; Gemma 4 26B A4B, fasterqwen3.8:27b, gemma4:26b
64 to 128 GB (Ryzen AI Max, Mac)Qwen 3.8 27B at 8 bits or full; for agents on 128 GB, gpt-oss-120bqwen3.8:27b-q8_0, gpt-oss:120b
Over 200 GB (DGX Spark cluster, 512 GB Mac)GLM-5.3-Flash, DeepSeek V4.1 Flashbeyond a mini-PC

For your exact setup: each machine’s page (for example the 128 GB Ryzen AI Max+ 395 or the 32 GB Mac mini M6), or Find my model in four questions (in French).

# what the integrated GPU took from the system, in bytes
cat /sys/class/drm/card*/device/mem_info_gtt_used

# reserve 28 GB of RAM for the rest of the machine (value in bytes)
# set this in the Ollama service environment
OLLAMA_GPU_OVERHEAD=30064771072

Without that setting, Ollama sizes its layers as if it were alone on the machine. It sees a 128 GB GPU, so it accepts a 65 GB model on a 94 GB machine that hosts thirty other services.

The shortcut: let a tool size it for you

The whole rule we just laid down (counting GBs, gauging the quant, spotting MoEs), a small open-source tool does it for you in two seconds: llmfit (MIT license, repo AlexsJones/llmfit). It scans your machine (RAM, CPU cores, VRAM, multi-GPU) and the installed runtimes (Ollama, llama.cpp, LM Studio, MLX, Docker Model Runner), cross-references that with a catalog of models and quants, and gives you a scored ranking: what fits, what will run fast, and the quant to aim for. Enough to avoid pulling 20 GB of model for nothing.

0 of 2 steps done Your ticks stay in this browser.

  1. Install it

    Through your usual package manager, depending on your system:

    # macOS / Linux (Homebrew)
    brew install AlexsJones/llmfit/llmfit
    # macOS / Linux (quick script)
    curl -fsSL https://llmfit.axjns.dev/install.sh | sh
    # Windows (Scoop)
    scoop install llmfit
    # Without installing anything permanently: Docker
    docker run ghcr.io/alexsjones/llmfit
    
  2. Run the analysis

    With no argument, you get an interactive terminal UI (and a web dashboard at http://<machine-ip>:8787). Prefer a raw table or JSON to script with? The options are there.

    llmfit              # interactive UI + web dashboard
    llmfit --cli        # classic table in the terminal
    llmfit recommend --json --use-case coding --limit 3   # top 3 for coding, as JSON
    

Both tools tell you what fits. To know what is best rated among what fits, task by task (code, agents, French…), use Quelle IA, machine by machine (in French): 64 real machines, with dated prices.

The 2026 model landscape

The headline: Chinese labs dominate open-weight (Alibaba, DeepSeek, Z.ai, Moonshot, MiniMax, Xiaomi), with Google (Gemma) and OpenAI (gpt-oss) as the main American options. Two silhouettes stand out. The giants (from a few hundred billion up to 2.8 T total parameters) for servers or clusters of machines. And the compact models (around 25-30B, dense or MoE) which actually run on a mini-PC. It’s this second family that interests us.

Here are the safe picks as of 28 September 2026, according to Quelle IA (in French):

ModelType / sizeWhat forCan you run it?
Qwen 3.8 27B (Alibaba, August 2026)Dense, 27B, 256K contextTop-rated under 128 GB, code included✅ 24 GB card or 32 GB unified: qwen3.8:27b
Gemma 4 26B A4B (Google, April 2026)MoE, 26B / ~4B active, 256K contextFast, non-Chinese, the test machine’s reference model✅ From 32 GB: gemma4:26b
Qwen3.5-9B / 4B (Alibaba, February 2026)DenseMachines with 8 to 16 GB✅ qwen3.5:9b, qwen3.5:4b
gpt-oss-120b (OpenAI, August 2025)MoE, 117BTop-rated for agents under 128 GB⚠️ 128 GB machine, 65 GB download: gpt-oss:120b
Qwen3.6 27B (Alibaba, April 2026)Dense, 27.8B, 256K contextThe predecessor, still solid✅ From 24 GB: qwen3.6:27b
GLM-4.7-Flash (Z.ai)~19 GB at Q4, 198K contextThe GLM that genuinely fits locally✅ From 32 GB: glm-4.7-flash
GLM-5.3-Flash (Z.ai, August 2026)MoE, 320BThe best local model for code and agents❌ ~220 GB at 4 bits; online via glm-5.3-flash:cloud
DeepSeek V4 Flash / V4.1 FlashMoE; V4 Flash: 284B, 13B active, 1M contextLong context; V4.1 Flash tops the local ranking❌ DGX Spark cluster
Kimi K3 (Moonshot, July 2026)2.8 T total parametersLong-haul agentic work❌ Server only

The first choice on a mini-PC is qwen3.8:27b, as soon as it fits. If speed matters more than the last few quality points, gemma4:26b answers much faster thanks to its 4B active parameters. And on a 128 GB machine running agents, gpt-oss:120b is worth a try. The ranking moves every week: local models ranking, all open models (including those that won’t fit at home), and Compare to put two models side by side (all in French).

Quantization, plainly

You’ve seen names like Q4_K_M go by. Let’s decode them. GGUF is the file format for local models, and the suffix indicates the precision: how many bits per weight you keep. Fewer bits = lighter file = it fits in less RAM. The trade-off is a small loss of quality.

LevelQualityWhen to use it
Q8_0Near perfectRarely justified for code (too heavy)
Q6_K / Q5_K_MExcellentCritical reasoning, if you have the memory headroom
Q4_K_MVery goodThe de facto standard: your default choice
< Q4Visible degradationOnly if memory forces you

Q4_K_M weighs about 30% of the original FP16 model for only ~2-3% loss. For code specifically, it holds up very well, moving to Q5 or Q6 doesn’t perceptibly improve code generation. You only drop below Q4 when forced, and you feel it fast (code and reasoning are the first to suffer).

Wiring the right quant into Ollama

In practice, it’s simple: with Ollama, the tag already encodes the quant. You choose your precision by writing it into the name.

0 of 2 steps done Your ticks stay in this browser.

  1. Pull the model at the right quant

    The suffix after the tag is the quantization. Without a suffix, Ollama takes a default (often Q4_K_M).

    # The default (usually Q4_K_M), perfect to start
    ollama pull qwen3.8:27b
    # Or the explicit quant, if you want to be sure
    ollama pull qwen3.8:27b-q4_K_M
    # More faithful, if you have the memory (64 GB and up)
    ollama pull qwen3.8:27b-q8_0
    
  2. Fit the context to the real need

    Context costs RAM (the famous KV cache). The default for OLLAMA_CONTEXT_LENGTH depends on graphics memory: 4,000 tokens under 24 GiB, 32,000 from 24 to 48 GiB, 256,000 above. Too low for an agent on a small card, very high on a Ryzen AI Max or a big Mac, where the cache swells silently (in GTT memory, on Strix Halo). Too much context can also push you into swap: that’s the #1 cause of “Ollama is slow.” Set OLLAMA_CONTEXT_LENGTH or num_ctx to what the task really needs: 64,000 for a coding agent according to Ollama’s docs, far less for a script.

Honestly: local or Claude?

We won’t sell you a dream. Here’s where local stands in Quelle IA’s 28 September 2026 edition, unfiltered.

For bounded and private tasks, short-window generation, refactoring, autocomplete, targeted debugging, local is genuinely competitive. It’s private, it’s free to use, and the benchmark gap has narrowed sharply (GLM-5.3 and DeepSeek V4 are closing in on the best, Kimi K3 is strong on agentics). And a point too often forgotten: an open-weight model performs much better in a good harness (a well-configured OpenCode) than in a raw chat.

But keep two limits in mind:

  • The summit is still Claude. On the Quelle IA code ranking (in French), Claude Opus 5.5 (99.9) and Claude Fable 5.1 (98.6) lead. The best model that fits on a 128 GB mini-PC, Qwen 3.8 27B, scores 81.1 there (still a provisional score).
  • The best open models don’t run on a mini-PC. Kimi K3 (2.8 T parameters), GLM-5.3 (744B), DeepSeek V4 Pro (1.6 T): that’s API or server. Even GLM-5.3-Flash, the best local model for code (88.9), needs two DGX Sparks. What fits (Qwen 3.8 27B, Gemma 4 26B A4B, gpt-oss-120b) is a notch below on long-haul agentic reliability and tool use.

Frequently asked questions

What is an MoE model and how much memory does it need?

An MoE (Mixture of Experts) model has a large total number of parameters but only activates a small share of them for each token. You size the memory on the total, since everything has to be loaded, while speed depends on the active parameters. Gemma 4 26B A4B, for example, takes 16 to 19 GB yet activates only about 4 billion parameters per token, which makes it much faster than a dense model of its size.

Which quantization should I choose for a local model?

Q4_K_M is the de facto standard: it weighs about 30% of the original FP16 model for only 2 to 3% loss, and holds up very well for code. Move up to Q5 or Q6 only for critical reasoning, if you have memory to spare. Only go below Q4 when forced to, because code and reasoning suffer first.

Why is Ollama slow on my machine?

The number one cause is a context that's too large, which makes the system spill onto disk, into swap. Context costs memory, and Ollama's default rises to 256,000 tokens on machines with more than 48 GiB of graphics memory. Set OLLAMA_CONTEXT_LENGTH or num_ctx to what the task really needs: 64,000 for a coding agent according to Ollama's docs, far less for a script.

How can I find out which models run on my PC without doing the maths?

The open-source tool llmfit scans your machine (RAM, CPU cores, VRAM) and the installed runtimes such as Ollama or LM Studio, then ranks what fits, what will run fast and which quant to aim for. It can also simulate a configuration you don't have yet, which is handy before a purchase. Without installing anything, the CanIRun.ai website gives an estimate right in your browser.

Why doesn't the free command show the memory my models are using?

On a shared-memory APU such as AMD Strix Halo, model weights take system RAM as GTT memory, which is invisible in free and in a Docker container's memory limit. The machine can fill up without any warning. Before keeping several models loaded, check what the integrated GPU has actually taken, and reserve RAM for the rest of the machine with OLLAMA_GPU_OVERHEAD.

Terms in this guide: Open-weight modelRAMParametersQuantizationOllamaGPU (graphics card)Open source (vs open-weight)Repository (repo)VRAMDockerOpenCodeAPITools (tool use)

Spotted a mistake?

A command stopped working, a price changed?

Tools change every month. Tell me what is wrong in this chapter and I will fix it and update its date.

Only the page, your message and the optional contact are kept. Nothing else.

Guide 24 of 31 · part 4 no guides read yet Open the list of guides