Skip to content
The 31 guidesFREN中文
Guides
Part 1 · guide 8 of 8 Level: Intermediate Reading time: 16 min

Chinese LLMs: MiMo, Qwen, DeepSeek, GLM, Kimi

Chinese models close behind the best, for a fraction of the price. Coding plan, pay-as-you-go API or local on the mini-PC: how to plug them into Claude Code, OpenCode and Codex, and what you should know about your data.

In this guide
  1. 01Why they are worth a look
  2. 02Three ways to use them
  3. 03Plugging them into your tools
  4. 04Your data: where it goes, what is done with it
  5. 05Frequently asked questions

In short

Chinese models (Xiaomi's MiMo, Moonshot's Kimi, Z.ai's GLM, DeepSeek, Alibaba's Qwen, MiniMax) fill the top sixteen rows of the Quelle IA open-weights ranking (28 September 2026 edition), about ten points below the best American models, often for much less: MiMo-V2.6-Pro costs nearly fifteen times less than Claude Opus 5.5 per million tokens. There are three ways to use them: a coding plan from the vendor ($6 to $199 a month), the pay-as-you-go API or an aggregator such as OpenCode Go, or locally with Ollama for the versions that fit in memory, such as Qwen 3.8 27B. Claude Code connects through ANTHROPIC_BASE_URL, OpenCode through /connect, Codex through a provider in ~/.codex/config.toml. Depending on the offer, your requests go through China, Singapore, Europe or the United States: never send them secrets or client code.

Do this first: Frontier models on a subscription: Claude, ChatGPT, Grok, Gemini

Two years ago, a Chinese model was mostly a small Qwen you ran locally out of curiosity. In October 2026, Xiaomi’s MiMo, Moonshot’s Kimi or Z.ai’s GLM code almost as well as the leading models from Anthropic and OpenAI, and often cost much less. Many of them plug straight into Claude Code, OpenCode or Codex, with a monthly plan designed for exactly that.

This guide covers it all: why they are worth a look, the three ways to use them, the exact configuration for each tool, and the awkward question, the one about your data.

Why they are worth a look

Three reasons, with numbers (rankings from Quelle IA, in French, 28 September 2026 edition).

  • They are close behind the best. In the overall ranking, the top Chinese model, Xiaomi’s MiMo-V2.6-Pro, is 13th out of 246 with a score of 90.5, followed by Kimi K3 (14th), Qwen 3.8 Max (21st), GLM-5.3 (24th) and DeepSeek V4 Pro 0813 (27th). At the top, Claude Opus 5.5 scores 99.9. In coding, GLM-5.3 ranks 9th with 91.5 (Code ranking).
  • They cost much less. MiMo-V2.6-Pro costs €0.44 per million tokens read and €0.87 per million written, against €4 and €20 for Claude Opus 5.5. In the value-for-money ranking, four of the top five are Chinese: three DeepSeek models and MiMo-V2.6-Pro.
  • Their weights are open. The top sixteen rows of the open-weights ranking are Chinese. Open means downloadable: any host can serve them, and the smaller ones run at home. On a machine with between 22 and 189 GB of usable memory, the best-rated model that fits is Alibaba’s Qwen 3.8 27B (82.4).

Three ways to use them

1. The vendor’s coding plan

Most Chinese vendors sell a monthly subscription reserved for coding, to plug into your agent. You pay a fixed amount, the vendor counts your usage its own way, and you are cut off (or slowed down) at the cap.

PlanVendor and modelsPrice per monthWhat it countsTools named by the vendor
Token PlanXiaomi: MiMo-V2.6-Pro, MiMo-V2.6-Flash$6, $16, $50 or $100credits per month (4.1 to 82 billion)Claude Code, OpenCode, Codex, MiMo Code, Cline
GLM Coding PlanZ.ai: GLM-5.3, GLM-5.3-Flash$18, $80 or $168credits per 5-hour window and per weekClaude Code, OpenCode, Codex, Cursor, Cline
Kimi CodeMoonshot: K3, K2.7 Code$19, $39, $99 or $199credits over 7 days plus a 5-hour window (roughly 300 to 1,200 requests)Claude Code, OpenCode, Codex, Roo Code
M PlanMiniMax: M3.1 Flash Preview$22, $55 or $1325-hour and 7-day windowsClaude Code, OpenCode, Codex, Cursor
Token Plan, personal editionAlibaba Cloud: Qwen 3.8, Qwen 3.7, DeepSeek V4, GLM-5.3$8, $16, $25 or $80 ($6, $10, $18 or $68 on promotion)credits per monthClaude Code, Codex, Cursor, Qwen Code
Coding Plan ProAlibaba Cloud: Qwen 3.7 Plus, Kimi K2.5, GLM-5, MiniMax M2.5$50 (limited places)6,000 requests per 5 hours, 90,000 per monthClaude Code, OpenCode, Codex, Cline

Public prices in dollars, before tax, recorded on 30 September 2026 by Quelle IA from each offer’s page. One vendor’s credits cannot be compared with another’s: each has its own scale. The full table, with yearly discounts and exact models, is on the Quelle IA Tools page (in French).

2. Pay as you go: the vendor’s API or an aggregator

Every vendor also sells its API per million tokens, with no subscription: you top up a balance and pay for what you use. That is the normal route for a script, a bot or an application, and the only one at DeepSeek, which sells no coding plan. Each model’s prices are in the Quelle IA comparison tool (in French).

If you want to try several models without opening five accounts, go through an aggregator (we introduce them in Installing the agent):

  • OpenRouter: one key for hundreds of models, including every model in this guide, paid as you go. Perfect in OpenCode. In Claude Code, OpenRouter warns that it is only guaranteed to work with Anthropic’s models.
  • OpenCode Zen: the OpenCode team’s pay-as-you-go gateway, with models tested inside an agent. Its page says it plainly: “All our models are hosted in the US”.
  • OpenCode Go: the same team’s plan, $10 a month ($40 for Go Plus), with GLM-5.3, Kimi K3, MiMo-V2.6-Pro, MiniMax M3, Qwen 3.8 and DeepSeek V4.1 Flash, among others. Each model has its own monthly cap in dollars of usage (recorded on 30 September 2026). A good way to try them all for the price of a lunch.

3. Locally on the mini-PC

This is the route this site prefers: open models run at home with Ollama, for free, without sending anything to anyone. The big Chinese models (Kimi K3, GLM-5.3, MiMo-V2.6-Pro) need several hundred gigabytes of memory: out of reach for a mini-PC. Their smaller cousins fit very well.

# 16 GB of memory: Qwen3.5-9B, a 6.6 GB download
ollama pull qwen3.5:9b

# 32 GB of unified memory or a 24 GB graphics card: Qwen 3.8 27B, 18 GB
# the best-rated model that fits, up to 189 GB of usable memory
ollama pull qwen3.8:27b

# 64 GB and up: Qwen3.5-35B-A3B, 24 GB, faster because it only activates 3 billion parameters at a time
ollama pull qwen3.5:35b-a3b

Even on a 128 GB Ryzen AI Max+ 395, Qwen 3.8 27B remains the best-rated model that fits (that machine’s page on Quelle IA, in French). A bigger model is not necessarily a better one: choose with Choosing and sizing your model and Quelle IA, machine by machine.

Plugging them into your tools

Most of these vendors expose two doors: one compatible with Anthropic’s API, which Claude Code knows how to use, and one compatible with OpenAI’s, for OpenCode and Codex. You just tell the tool which address to call and with which key.

0 of 4 steps done Your ticks stay in this browser.

  1. Get the key

    Create your key in the vendor’s console: a plan key (it starts with tp- at Xiaomi, for instance) or a pay-as-you-go API key. The two are not interchangeable, and each has its own address.

  2. Pick the right address

    Several vendors have one address per region. Xiaomi’s Token Plan has three: token-plan-cn (China), token-plan-sgp (Singapore) and token-plan-ams (Europe, Amsterdam); the console gives you yours. Kimi Code separates api.kimi.com (China) and api.kimi.ai (outside China).

  3. Configure the tool

    Follow your agent’s tab below.

  4. Check

    In Claude Code, /status shows the address and model in use. In OpenCode and Codex, /models and /model list what is connected. Run a small task and watch the usage in the vendor’s console.

Claude Code reads three settings: ANTHROPIC_BASE_URL (the address), a key, and the model names. The trick to keep your usual Claude Code intact: put these settings in a separate file, loaded only when you ask for it with --settings. Example with the GLM Coding Plan, following Z.ai’s documentation:

# a settings file just for GLM, next to your normal configuration
cat > ~/.claude/glm.json <<'EOF'
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
    "ANTHROPIC_AUTH_TOKEN": "your_z_ai_key",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3[1m]",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3[1m]",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
    "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
    "API_TIMEOUT_MS": "3000000",
    "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"
  }
}
EOF
chmod 600 ~/.claude/glm.json        # the key is inside: readable by you only

# a Claude Code session on GLM; plain "claude" stays on Anthropic
claude --settings ~/.claude/glm.json

For the other vendors, change the address, the key variable and the models, following their documentation:

OfferANTHROPIC_BASE_URLKey inModel
Xiaomi, Token Planhttps://token-plan-ams.xiaomimimo.com/anthropic (or -sgp, -cn)ANTHROPIC_AUTH_TOKENmimo-v2.6-pro
Xiaomi, pay as you gohttps://api.xiaomimimo.com/anthropicANTHROPIC_AUTH_TOKENmimo-v2.6-pro
Kimi Codehttps://api.kimi.ai/coding/ANTHROPIC_API_KEYk3-256k
MiniMax, M Planhttps://api.minimax.io/anthropicANTHROPIC_AUTH_TOKENMiniMax-M3.1-Flash-Preview
DeepSeek, pay as you gohttps://api.deepseek.com/anthropicANTHROPIC_AUTH_TOKENdeepseek-flash or deepseek-v4-pro

If ANTHROPIC_BASE_URL or ANTHROPIC_AUTH_TOKEN are already lying around in your ~/.bashrc, remove them: Xiaomi and MiniMax explicitly ask for this to avoid conflicts.

Your data: where it goes, what is done with it

An open model and an online service are not the same thing. The weights of GLM or Qwen can be downloaded and run anywhere. When you go through the vendor’s API or plan, your code goes to them, and their terms apply. Here is what those terms say, read on 1 October 2026 on the official pages:

OfferWhere your requests goTraining on your content
Xiaomi (API, Token Plan)China, Singapore or Europe depending on the cluster; outside mainland China, the agreement provides for storage in Europe and SingaporeNot without your prior consent, according to the platform’s policy, which however covers use in China; the version for other countries could not be found
Z.ai (GLM)Singapore, “generally”Not without explicit consent for API customers; possible for individual accounts. The terms do not say which box the GLM Coding Plan falls into
Moonshot (Kimi)Singapore for the API platform; for Kimi Code, only the Beijing entity’s policy (storage in China) is onlineAllowed by the API terms unless otherwise agreed in writing, while another page of the same documentation says the opposite for enterprises
MiniMaxA data center in the United States, according to its privacy policyInputs and outputs may be used to improve the services; no opt-out found
Alibaba Cloud (Token Plan)Singapore region, “Global” inference: requests cross bordersNot without your consent, according to the product terms
DeepSeekChina, for both storage and processingYes, with a right to opt out provided by the policy

For these international offers, the contract is signed with a Singapore company (Alibaba goes through its Dutch subsidiary for a European billing address), except DeepSeek, which contracts from Hangzhou under Chinese law.

None of this is unique to China: with any online provider, American ones included, your code leaves the machine and depends on terms you do not control. What weighs here: Chinese parent companies, terms that sometimes contradict each other from one page to the next, and little leverage over them from Europe. The right attitude fits in one rule: only send these services what you would publish without regret.

Frequently asked questions

Where do you feel the gap between a Chinese model and Claude or GPT?

On long, hard tasks: a big refactor, a nasty bug, an architecture to design. For everyday code, the ten-point gap in the rankings shows far less than it does on the bill. Each score also has its margin of error, which Quelle IA publishes for every model.

Can you use a Chinese coding plan in a script or an application?

Generally not: these plans are for coding inside a tool. Xiaomi and Alibaba forbid using the plan key in an automated script or an application server, Z.ai forbids direct calls from your apps or bots, and MiniMax steers production use to pay-as-you-go. For a bot that runs overnight, use the pay-as-you-go API.

Which Chinese model can you run locally on a mini-PC?

Flagship models such as Kimi K3, GLM-5.3 or MiMo-V2.6-Pro need several hundred gigabytes of memory, out of reach for a mini-PC. Their smaller siblings run very well with Ollama: Qwen3.5-9B with 16 GB of memory, Qwen 3.8 27B with 32 GB of unified memory or a 24 GB graphics card, Qwen3.5-35B-A3B from 64 GB. Even on a 128 GB machine, Qwen 3.8 27B remains the best-rated model that fits.

Does an Ollama model tagged :cloud run locally?

No. In the Ollama library, kimi-k3 and minimax-m3 only exist as :cloud versions: the command ollama run kimi-k3:cloud sends the request to Ollama Cloud's servers, hosted mostly in the United States according to Ollama. It is convenient, but your data leaves the machine.

Does DeepSeek use your data to train its models?

Yes, according to its policy read on 1 October 2026, with a right to opt out. DeepSeek stores and processes requests in China and contracts from Hangzhou under Chinese law. For sensitive code, only an open model running on your own machine with Ollama guarantees that nothing leaves.

Terms in this guide: Claude CodeOpenCodeCodexTokenAgentAPICoding planUnified APIInference

Spotted a mistake?

A command stopped working, a price changed?

Tools change every month. Tell me what is wrong in this chapter and I will fix it and update its date.

Only the page, your message and the optional contact are kept. Nothing else.

Guide 8 of 31 · part 1 no guides read yet Open the list of guides