Going hybrid: cloud + local
The best of both worlds. A frontier model on a subscription (Claude, GPT, Grok or Gemini) plans, reasons and reviews, and delegates the repetitive heavy lifting to a local model. Private, frugal, and devastatingly effective.
In this guide
Guide checked 3 months ago: some commands may have changed. Let us know if so.
In short
The hybrid setup gives hard reasoning, architecture and review to a frontier model on a subscription (Claude with Claude Code, GPT with Codex, Grok with Grok Build, Gemini with Antigravity CLI), and repetitive, sensitive or offline work to a local model served by Ollama. Claude Code orchestrates: it writes the scripts that call the local API on port 11434. Since Ollama 0.15, Claude Code can also run on the local model itself with ollama launch claude, for code that must not leave the machine or when the connection drops. The golden rule: only delegate well-bounded tasks to local, and have the result reviewed.
Do this first: Ollama & local models
You now have two worlds at hand. On one side, the big vendors’ frontier models, taken on a subscription with their coding agent: Claude with Claude Code, GPT with Codex, Grok with Grok Build, Gemini with Antigravity CLI. On the other, models that run for free at home (Ollama). The wrong question is “which one to choose?” The right answer is: both, at the same time, each in its place.
That’s the hybrid setup, and it’s the most effective way to work with a mini-PC. The machine orchestrates: it runs the agents and keeps your projects and scripts. The frontier plan, at a fixed price every month ($20 for Claude Pro or ChatGPT Plus, $30 for SuperGrok, €21.99 for Google AI Pro, prices checked on October 1, 2026), does the bulk of the reasoning. The local model takes the confidential, the repetitive and the offline. Tiers, limits and what each vendor allows are compared in Frontier models on a subscription.
The examples in this guide put Claude Code in the conductor’s seat, because it’s the most battle-tested setup; it carries over as is to Codex or OpenCode. The cloud agent decides who does what, and delegates the bulk work to the local model.
- Claude· Claude Code
- ChatGPT· Codex
- Grok· Grok Build
- Gemini· Antigravity CLI
- gemma4:26b
- qwen3.8:27b
- nomic-embed-text
Why mix rather than choose
The two worlds have opposite strengths, and that’s exactly what makes them complementary:
- The cloud (Claude, GPT, Grok, Gemini) is unbeatable on hard reasoning, long tasks, reliable tool use, architecture. But plans have caps (per 5-hour window and per week), the API is billed per token, and your data goes to a third party.
- Local (Ollama) is free to use, private, available offline, and largely good enough for bounded and repetitive tasks. But it tires on long agentics and cutting-edge reasoning (we talked about it bluntly in Choosing your model, and the Quelle IA local models ranking, in French, puts numbers on the gap).
The hybrid setup takes the best of each: you keep cutting-edge intelligence where it really matters, and you knock down cost and data leaks on everything else, that is, 80% of the volume.
Claude Code as orchestrator: how it works
The thing that makes this possible: Claude Code can run commands. And Ollama exposes a dead-simple local API on http://localhost:11434. So Claude Code can call your local model, via a curl, a script, or a small tool, exactly as it would call any other command.
Concretely, you tell Claude: “for this bulk task, don’t do it yourself, delegate it to the local model via the Ollama API.” It writes the script that loops over your files, hits the local model for each one, and brings back the consolidated result. It keeps the big picture; local does the grunt work.
# The basic move Claude Code orchestrates: call the local model (here gemma4:26b, fast)
curl -s http://localhost:11434/api/generate -d '{
"model": "gemma4:26b",
"prompt": "Summarize this file in 3 bullets: '"$(cat report.md)"'",
"stream": false
}' | jq -r .response
Four concrete ways to route the work
0 of 4 steps done Your ticks stay in this browser.
-
Route by cost: volume to local, the cutting edge to cloud
Need to rewrite 300 product descriptions, classify 2,000 comments, or generate boilerplate by the truckload? It’s repetitive and bounded: local model. Need to design the module’s architecture or debug a nasty race condition? It’s rare and hard: Claude. Claude writes the pipeline, local runs the 300 calls for free.
-
Route by confidentiality: the sensitive stuff stays home
Proprietary code, customer data, things that must not leave? You have them processed by the local model: nothing leaves the machine. Claude keeps a coordination role on the non-sensitive part. And if the sensitive work needs a real agent, start a separate session with
ollama launch claude: same Claude Code, but the model requests stay at home. It’s a strong argument in a professional context (cf. Securing access). -
Homemade RAG: local embeddings, cloud reasoning
Want the agent to know your corpus (your docs, your articles, your code)? Generate the embeddings locally with Ollama (
nomic-embed-textor equivalent), store them, and let Claude reason over the most relevant passages you serve it. The linking and indexing, free and private and local; the final intelligence, on Claude’s side. -
The offline net: OpenCode + local takes over
No connection? Train, plane, outage? You switch to OpenCode wired to your local model, or to Claude Code itself with
ollama launch claude, and keep coding. The cloud is no longer a single point of failure: your machine stays a self-sufficient workshop.
Wiring both up, in practice
Claude Code is the default orchestrator. You have nothing special to install: it already knows how to run commands, so it already knows how to call Ollama. Just give it the instruction in your CLAUDE.md:
# Hybrid strategy
- For repetitive, bulk tasks (rewriting, classification,
summaries, boilerplate generation), delegate to the local model via the
Ollama API (http://localhost:11434), don't do them yourself.
- Keep complex reasoning, architecture, and review for yourself.
- Code and data marked "sensitive": local model only.
From now on, when you hand it a big batch, it writes the script that hits local and brings back the result. You steer, it distributes.
The second setup, since Ollama 0.15: the whole of Claude Code on the local model. Ollama speaks Anthropic’s API, and one command does the wiring:
# a Claude Code session with a local brain
ollama launch claude --model qwen3.8:27b
Keep two terminals: your usual session on Anthropic’s models for architecture and review, and this local session for sensitive or offline code. Plan for a context of at least 64,000 tokens and a model that can call tools (details in Ollama & local models). Model requests don’t leave the machine; Claude Code still sends operational telemetry, with no code or prompts according to its documentation, which CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 turns off.
OpenCode makes routing explicit: you switch models depending on the task. Wire up both a cloud provider (Anthropic/OpenRouter) and Ollama locally, then switch via /models:
- hard task, reasoning, architecture → cloud model;
- bulk task, sensitive, or offline → your local model.
It’s the ideal approach if you want to see and control which model handles what, rather than letting an orchestrator decide. And to see which other coding tools accept a local model (and which refuse it), Quelle IA keeps an up-to-date list (in French). And if Claude Code stays your conductor, OpenCode + local makes an excellent “arm” that Claude can invoke from the command line.
All the commands in this guide
Frequently asked questions
How much does a frontier model subscription cost?
As of October 1, 2026, expect $20 a month for Claude Pro or ChatGPT Plus, $30 for SuperGrok and €21.99 for Google AI Pro. The price is fixed, but the plan has usage caps, per 5-hour window and per week. That is one more reason to hand high-volume work to a local model, which costs nothing to use.
How do I tell Claude Code to delegate to the local model?
Write the instruction once in the project's CLAUDE.md file: repetitive, high-volume tasks go to the local model through the Ollama API at http://localhost:11434, complex reasoning, architecture and review stay with Claude, and data marked sensitive goes to the local model only. When you then hand it a big batch, it writes the script that calls the local model itself and brings you back the result.
How do I build a RAG with a local model?
Generate the embeddings for your corpus locally with Ollama, for example with nomic-embed-text, and store them. You then feed Claude the most relevant passages, and it does the reasoning on them. Indexing stays free and private on your machine, and only the final thinking goes through the cloud.
If Claude Code runs on a local model, does it still send data out?
Requests to the model stay on your machine. Claude Code still sends operational telemetry, with no code or prompts according to its documentation. The CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 variable turns it off.
Can I choose which model handles each task myself?
Yes, with OpenCode: you plug in both a cloud provider and Ollama locally, then switch between them with /models. The cloud takes the hard tasks, the local model takes high-volume, sensitive or offline work. It is the right approach if you want to see and control who handles what instead of letting an orchestrator decide.
Terms in this guide: Frontier modelAgentCodexGrok BuildAntigravity CLIClaude CodeOpenCodeTools (tool use)APITokenOllama
Spotted a mistake?
A command stopped working, a price changed?
Tools change every month. Tell me what is wrong in this chapter and I will fix it and update its date.