Chinese LLMs: MiMo, Qwen, DeepSeek, GLM, Kimi
Chinese models close behind the best, for a fraction of the price. Coding plan, pay-as-you-go API or local on the mini-PC: how to plug them into Claude Code, OpenCode and Codex, and what you should know about your data.
In this guide
Guide checked 3 months ago: some commands may have changed. Let us know if so.
In short
Chinese models (Xiaomi's MiMo, Moonshot's Kimi, Z.ai's GLM, DeepSeek, Alibaba's Qwen, MiniMax) fill the top sixteen rows of the Quelle IA open-weights ranking (28 September 2026 edition), about ten points below the best American models, often for much less: MiMo-V2.6-Pro costs nearly fifteen times less than Claude Opus 5.5 per million tokens. There are three ways to use them: a coding plan from the vendor ($6 to $199 a month), the pay-as-you-go API or an aggregator such as OpenCode Go, or locally with Ollama for the versions that fit in memory, such as Qwen 3.8 27B. Claude Code connects through ANTHROPIC_BASE_URL, OpenCode through /connect, Codex through a provider in ~/.codex/config.toml. Depending on the offer, your requests go through China, Singapore, Europe or the United States: never send them secrets or client code.
Do this first: Frontier models on a subscription: Claude, ChatGPT, Grok, Gemini
Two years ago, a Chinese model was mostly a small Qwen you ran locally out of curiosity. In October 2026, Xiaomi’s MiMo, Moonshot’s Kimi or Z.ai’s GLM code almost as well as the leading models from Anthropic and OpenAI, and often cost much less. Many of them plug straight into Claude Code, OpenCode or Codex, with a monthly plan designed for exactly that.
This guide covers it all: why they are worth a look, the three ways to use them, the exact configuration for each tool, and the awkward question, the one about your data.
Why they are worth a look
Three reasons, with numbers (rankings from Quelle IA, in French, 28 September 2026 edition).
- They are close behind the best. In the overall ranking, the top Chinese model, Xiaomi’s MiMo-V2.6-Pro, is 13th out of 246 with a score of 90.5, followed by Kimi K3 (14th), Qwen 3.8 Max (21st), GLM-5.3 (24th) and DeepSeek V4 Pro 0813 (27th). At the top, Claude Opus 5.5 scores 99.9. In coding, GLM-5.3 ranks 9th with 91.5 (Code ranking).
- They cost much less. MiMo-V2.6-Pro costs €0.44 per million tokens read and €0.87 per million written, against €4 and €20 for Claude Opus 5.5. In the value-for-money ranking, four of the top five are Chinese: three DeepSeek models and MiMo-V2.6-Pro.
- Their weights are open. The top sixteen rows of the open-weights ranking are Chinese. Open means downloadable: any host can serve them, and the smaller ones run at home. On a machine with between 22 and 189 GB of usable memory, the best-rated model that fits is Alibaba’s Qwen 3.8 27B (82.4).
Three ways to use them
1. The vendor’s coding plan
Most Chinese vendors sell a monthly subscription reserved for coding, to plug into your agent. You pay a fixed amount, the vendor counts your usage its own way, and you are cut off (or slowed down) at the cap.
| Plan | Vendor and models | Price per month | What it counts | Tools named by the vendor |
|---|---|---|---|---|
| Token Plan | Xiaomi: MiMo-V2.6-Pro, MiMo-V2.6-Flash | $6, $16, $50 or $100 | credits per month (4.1 to 82 billion) | Claude Code, OpenCode, Codex, MiMo Code, Cline |
| GLM Coding Plan | Z.ai: GLM-5.3, GLM-5.3-Flash | $18, $80 or $168 | credits per 5-hour window and per week | Claude Code, OpenCode, Codex, Cursor, Cline |
| Kimi Code | Moonshot: K3, K2.7 Code | $19, $39, $99 or $199 | credits over 7 days plus a 5-hour window (roughly 300 to 1,200 requests) | Claude Code, OpenCode, Codex, Roo Code |
| M Plan | MiniMax: M3.1 Flash Preview | $22, $55 or $132 | 5-hour and 7-day windows | Claude Code, OpenCode, Codex, Cursor |
| Token Plan, personal edition | Alibaba Cloud: Qwen 3.8, Qwen 3.7, DeepSeek V4, GLM-5.3 | $8, $16, $25 or $80 ($6, $10, $18 or $68 on promotion) | credits per month | Claude Code, Codex, Cursor, Qwen Code |
| Coding Plan Pro | Alibaba Cloud: Qwen 3.7 Plus, Kimi K2.5, GLM-5, MiniMax M2.5 | $50 (limited places) | 6,000 requests per 5 hours, 90,000 per month | Claude Code, OpenCode, Codex, Cline |
Public prices in dollars, before tax, recorded on 30 September 2026 by Quelle IA from each offer’s page. One vendor’s credits cannot be compared with another’s: each has its own scale. The full table, with yearly discounts and exact models, is on the Quelle IA Tools page (in French).
2. Pay as you go: the vendor’s API or an aggregator
Every vendor also sells its API per million tokens, with no subscription: you top up a balance and pay for what you use. That is the normal route for a script, a bot or an application, and the only one at DeepSeek, which sells no coding plan. Each model’s prices are in the Quelle IA comparison tool (in French).
If you want to try several models without opening five accounts, go through an aggregator (we introduce them in Installing the agent):
- OpenRouter: one key for hundreds of models, including every model in this guide, paid as you go. Perfect in OpenCode. In Claude Code, OpenRouter warns that it is only guaranteed to work with Anthropic’s models.
- OpenCode Zen: the OpenCode team’s pay-as-you-go gateway, with models tested inside an agent. Its page says it plainly: “All our models are hosted in the US”.
- OpenCode Go: the same team’s plan, $10 a month ($40 for Go Plus), with GLM-5.3, Kimi K3, MiMo-V2.6-Pro, MiniMax M3, Qwen 3.8 and DeepSeek V4.1 Flash, among others. Each model has its own monthly cap in dollars of usage (recorded on 30 September 2026). A good way to try them all for the price of a lunch.
3. Locally on the mini-PC
This is the route this site prefers: open models run at home with Ollama, for free, without sending anything to anyone. The big Chinese models (Kimi K3, GLM-5.3, MiMo-V2.6-Pro) need several hundred gigabytes of memory: out of reach for a mini-PC. Their smaller cousins fit very well.
# 16 GB of memory: Qwen3.5-9B, a 6.6 GB download
ollama pull qwen3.5:9b
# 32 GB of unified memory or a 24 GB graphics card: Qwen 3.8 27B, 18 GB
# the best-rated model that fits, up to 189 GB of usable memory
ollama pull qwen3.8:27b
# 64 GB and up: Qwen3.5-35B-A3B, 24 GB, faster because it only activates 3 billion parameters at a time
ollama pull qwen3.5:35b-a3b
Even on a 128 GB Ryzen AI Max+ 395, Qwen 3.8 27B remains the best-rated model that fits (that machine’s page on Quelle IA, in French). A bigger model is not necessarily a better one: choose with Choosing and sizing your model and Quelle IA, machine by machine.
Plugging them into your tools
Most of these vendors expose two doors: one compatible with Anthropic’s API, which Claude Code knows how to use, and one compatible with OpenAI’s, for OpenCode and Codex. You just tell the tool which address to call and with which key.
0 of 4 steps done Your ticks stay in this browser.
-
Get the key
Create your key in the vendor’s console: a plan key (it starts with
tp-at Xiaomi, for instance) or a pay-as-you-go API key. The two are not interchangeable, and each has its own address. -
Pick the right address
Several vendors have one address per region. Xiaomi’s Token Plan has three:
token-plan-cn(China),token-plan-sgp(Singapore) andtoken-plan-ams(Europe, Amsterdam); the console gives you yours. Kimi Code separatesapi.kimi.com(China) andapi.kimi.ai(outside China). -
Configure the tool
Follow your agent’s tab below.
-
Check
In Claude Code,
/statusshows the address and model in use. In OpenCode and Codex,/modelsand/modellist what is connected. Run a small task and watch the usage in the vendor’s console.
Claude Code reads three settings: ANTHROPIC_BASE_URL (the address), a key, and the model names. The trick to keep your usual Claude Code intact: put these settings in a separate file, loaded only when you ask for it with --settings. Example with the GLM Coding Plan, following Z.ai’s documentation:
# a settings file just for GLM, next to your normal configuration
cat > ~/.claude/glm.json <<'EOF'
{
"env": {
"ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
"ANTHROPIC_AUTH_TOKEN": "your_z_ai_key",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3[1m]",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3[1m]",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
"API_TIMEOUT_MS": "3000000",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"
}
}
EOF
chmod 600 ~/.claude/glm.json # the key is inside: readable by you only
# a Claude Code session on GLM; plain "claude" stays on Anthropic
claude --settings ~/.claude/glm.json
For the other vendors, change the address, the key variable and the models, following their documentation:
| Offer | ANTHROPIC_BASE_URL | Key in | Model |
|---|---|---|---|
| Xiaomi, Token Plan | https://token-plan-ams.xiaomimimo.com/anthropic (or -sgp, -cn) | ANTHROPIC_AUTH_TOKEN | mimo-v2.6-pro |
| Xiaomi, pay as you go | https://api.xiaomimimo.com/anthropic | ANTHROPIC_AUTH_TOKEN | mimo-v2.6-pro |
| Kimi Code | https://api.kimi.ai/coding/ | ANTHROPIC_API_KEY | k3-256k |
| MiniMax, M Plan | https://api.minimax.io/anthropic | ANTHROPIC_AUTH_TOKEN | MiniMax-M3.1-Flash-Preview |
| DeepSeek, pay as you go | https://api.deepseek.com/anthropic | ANTHROPIC_AUTH_TOKEN | deepseek-flash or deepseek-v4-pro |
If ANTHROPIC_BASE_URL or ANTHROPIC_AUTH_TOKEN are already lying around in your ~/.bashrc, remove them: Xiaomi and MiniMax explicitly ask for this to avoid conflicts.
OpenCode already knows most of these providers: no file to write. Start opencode, type /connect, search for the provider, paste the key, then pick the model with /models.
opencode
# inside OpenCode:
/connect # search for "Z.AI Coding Plan", for instance, then paste the key
/models # pick GLM-5.3
The names to search for, as OpenCode’s catalogue shows them: Xiaomi Token Plan (Europe), (Singapore) or (China) depending on the address in your console, Xiaomi for the pay-as-you-go API, Z.AI Coding Plan or Z.AI, Kimi For Coding, MiniMax Token Plan (minimax.io) for the M Plan, DeepSeek, Alibaba Coding Plan and OpenCode Go. Keys are stored in ~/.local/share/opencode/auth.json.
MiMo Code, Xiaomi’s agent, is derived from OpenCode: same interface, same commands, with Xiaomi connected out of the box.
curl -fsSL https://mimo.xiaomi.com/install | bash # installs the mimo command
mimo auth login # pick MiMo, sign in through the browser
cd ~/projects/my-project && mimo # then /models to switch models
Codex speaks a single protocol, OpenAI’s Responses API: wire_api = "responses" is the only accepted value. Xiaomi, Z.ai, Kimi and DeepSeek expose it. You declare the provider in ~/.codex/config.toml, then select it in a profile, a small file layered over the configuration for one session. Example with Kimi Code, following its documentation:
# ~/.codex/config.toml: declares the provider without making it the default
[model_providers.kimi]
name = "Kimi"
base_url = "https://api.kimi.ai/coding/v1" # outside China; api.kimi.com in China
env_key = "KIMI_API_KEY" # Codex reads the key from this variable
wire_api = "responses"
# ~/.codex/kimi.config.toml: the "kimi" profile
model_provider = "kimi"
model = "k3-256k"
model_catalog_json = "~/.codex/models.json" # model sheet supplied by Kimi
export KIMI_API_KEY="your_kimi_key" # put it in ~/.bashrc to keep it
codex --profile kimi # plain "codex" stays on OpenAI
The models.json file describes the models to Codex (context size, reasoning levels): copy the one from the vendor’s documentation. The other Responses-compatible addresses: https://api.z.ai/api/v1 at Z.ai, https://api.xiaomimimo.com/v1 (pay as you go) or https://token-plan-ams.xiaomimimo.com/v1 (plan, Europe region) at Xiaomi, https://api.deepseek.com/ at DeepSeek.
Your data: where it goes, what is done with it
An open model and an online service are not the same thing. The weights of GLM or Qwen can be downloaded and run anywhere. When you go through the vendor’s API or plan, your code goes to them, and their terms apply. Here is what those terms say, read on 1 October 2026 on the official pages:
| Offer | Where your requests go | Training on your content |
|---|---|---|
| Xiaomi (API, Token Plan) | China, Singapore or Europe depending on the cluster; outside mainland China, the agreement provides for storage in Europe and Singapore | Not without your prior consent, according to the platform’s policy, which however covers use in China; the version for other countries could not be found |
| Z.ai (GLM) | Singapore, “generally” | Not without explicit consent for API customers; possible for individual accounts. The terms do not say which box the GLM Coding Plan falls into |
| Moonshot (Kimi) | Singapore for the API platform; for Kimi Code, only the Beijing entity’s policy (storage in China) is online | Allowed by the API terms unless otherwise agreed in writing, while another page of the same documentation says the opposite for enterprises |
| MiniMax | A data center in the United States, according to its privacy policy | Inputs and outputs may be used to improve the services; no opt-out found |
| Alibaba Cloud (Token Plan) | Singapore region, “Global” inference: requests cross borders | Not without your consent, according to the product terms |
| DeepSeek | China, for both storage and processing | Yes, with a right to opt out provided by the policy |
For these international offers, the contract is signed with a Singapore company (Alibaba goes through its Dutch subsidiary for a European billing address), except DeepSeek, which contracts from Hangzhou under Chinese law.
None of this is unique to China: with any online provider, American ones included, your code leaves the machine and depends on terms you do not control. What weighs here: Chinese parent companies, terms that sometimes contradict each other from one page to the next, and little leverage over them from Europe. The right attitude fits in one rule: only send these services what you would publish without regret.
All the commands in this guide
Frequently asked questions
Where do you feel the gap between a Chinese model and Claude or GPT?
On long, hard tasks: a big refactor, a nasty bug, an architecture to design. For everyday code, the ten-point gap in the rankings shows far less than it does on the bill. Each score also has its margin of error, which Quelle IA publishes for every model.
Can you use a Chinese coding plan in a script or an application?
Generally not: these plans are for coding inside a tool. Xiaomi and Alibaba forbid using the plan key in an automated script or an application server, Z.ai forbids direct calls from your apps or bots, and MiniMax steers production use to pay-as-you-go. For a bot that runs overnight, use the pay-as-you-go API.
Which Chinese model can you run locally on a mini-PC?
Flagship models such as Kimi K3, GLM-5.3 or MiMo-V2.6-Pro need several hundred gigabytes of memory, out of reach for a mini-PC. Their smaller siblings run very well with Ollama: Qwen3.5-9B with 16 GB of memory, Qwen 3.8 27B with 32 GB of unified memory or a 24 GB graphics card, Qwen3.5-35B-A3B from 64 GB. Even on a 128 GB machine, Qwen 3.8 27B remains the best-rated model that fits.
Does an Ollama model tagged :cloud run locally?
No. In the Ollama library, kimi-k3 and minimax-m3 only exist as :cloud versions: the command ollama run kimi-k3:cloud sends the request to Ollama Cloud's servers, hosted mostly in the United States according to Ollama. It is convenient, but your data leaves the machine.
Does DeepSeek use your data to train its models?
Yes, according to its policy read on 1 October 2026, with a right to opt out. DeepSeek stores and processes requests in China and contracts from Hangzhou under Chinese law. For sensitive code, only an open model running on your own machine with Ollama guarantees that nothing leaves.
Terms in this guide: Claude CodeOpenCodeCodexTokenAgentAPICoding planUnified APIInference
Spotted a mistake?
A command stopped working, a price changed?
Tools change every month. Tell me what is wrong in this chapter and I will fix it and update its date.