China’s open models went from “cheap curiosity” to “four of the top five open-weight slots” faster than almost anyone predicted. But “there are good Chinese models now” isn’t a buying decision. This is the map — who’s who, what each one is actually best at, and how to choose without the tribal nonsense.
Why this map matters now
Quick context before the picks. In 2026, Chinese labs hold most of the top open-weight positions, and on some coding benchmarks their best models trade blows with the closed frontier. The old line — “Chinese models are only for cheap experiments” — is aging badly.
You don’t have to make that political. You do have to understand the tradeoffs, because “open and cheaper” is only a bargain if the model actually wins your jobs. So here’s the field, one lab at a time.
Kimi (Moonshot) — the frontier overachiever
Kimi is the one making the loudest noise for good reason. Its latest flagship is a massive open-weight model with a million-token context window and clever attention engineering that makes that long context genuinely cheap to run. At launch it topped a frontend-coding arena, ahead of leading closed models on that test.
Reach for Kimi when: you’re doing serious coding, or you need to reason across huge documents and codebases at once. Just know the biggest versions need server-class hardware, so most people will use it via API rather than self-hosting.
DeepSeek — the efficiency champion
DeepSeek is the lab that kicked this whole global conversation off, and its reputation is still the most durable one on the list: exceptional reasoning and capability per dollar. It’s the “how is it this good and this cheap?” model.
Reach for DeepSeek when: you want strong reasoning at a genuinely low cost, especially for high-volume work where price per token actually moves your margins.
Qwen (Alibaba) — the versatile giant
Qwen is the broadest family of the four. There’s an open Qwen line you can self-host and a closed, API-only “Max” flagship for top-end performance. Its standout strength is language: it’s arguably the best multilingual model going, especially across Chinese, Japanese, and Korean, and it comes with real commerce integrations behind it.
Reach for Qwen when: you need multilingual work, or you want one flexible family that spans open self-hosting and a hosted top tier.
GLM (Zhipu) — the quiet all-rounder
GLM doesn’t grab as many headlines, but it has repeatedly ranked at or near the top of the Chinese open-weight pack, with strong coding and solid general-purpose performance. It’s the dependable all-rounder that belongs in any serious comparison.
Reach for GLM when: you want a well-balanced open model for general and coding tasks and you’re shortlisting beyond the two or three names everyone already knows.
How to actually choose
Don’t pick by vibes or by which lab trended this week. Run the same simple test across two or three of them.
- Pick five real prompts from your own work — not just benchmark puzzles.
- Run them on one US default plus two Chinese options.
- Score quality, speed, cost, and trust.
- Read the license and data-residency terms in plain language.
- Decide by task, not by nationality.
The one rule that keeps you safe
A quick reality check before you rewire your stack: data residency matters for client work, content filters differ by region, self-hosting carries an ops burden, and export-control noise can change access overnight. So don’t bet the whole business on a single foreign dependency you can’t swap.
For most freelancers, the smart play is a hosted multi-model habit — a trusted default for sensitive work, and a cheap, capable Chinese option for drafts, experiments, and long-context muscle. Test widely, deploy carefully, keep a backup. That’s how you turn this whole competitive wave into your advantage instead of your headache.
Which to try first — and how to access it
If you want to test one this weekend without touching a server, here’s the shortcut. You don’t self-host these; you reach them through hosted providers that serve open models via a simple API or chat interface, often with a free tier. For a first taste, pick by your job: try DeepSeek for cheap reasoning, Qwen if you work across languages, or Kimi for coding or feeding in huge documents. Run the exact same five prompts you’d give your current default, and compare. One weekend of real testing beats a month of reading benchmark threads.
Which of these four have you actually tried, and for what? Tell me in the comments — I’m always curious how they hold up on real work versus benchmarks.