The China frontier in ten labs

AIAcademy · AIAcademy · 2026-05-16

Read CSIS on DeepSeek, Huawei and export controls

Western coverage of Chinese AI still over-indexes on DeepSeek. DeepSeek V4-Pro is now within 0.2 points of Claude Opus 4.6 on SWE-bench Verified — but it is one of ten labs that matter, and arguably not the most strategically interesting story of the year.

The current cast: DeepSeek (V4-Pro, R2 pending), Alibaba's Qwen (3.6-Plus, native 1M context — see the Qwen team's blog), Zhipu (GLM-5.1), Moonshot (Kimi K2.6 with a 300-agent "Swarm 2.0"), ByteDance (Doubao 2.0 + Seedream 5.0), MiniMax, StepFun, Xiaomi, Baichuan, and Yi. As of early 2026, six of those — Xiaomi, Alibaba, MiniMax, Zhipu, DeepSeek, StepFun — account for ~45% of weekly OpenRouter volume. This is no longer a long tail.

The single sharpest signal is GLM-5.1, which was trained entirely on Huawei Ascend chips and currently tops SWE-Bench Pro, beating GPT-5.4 and Opus 4.6. CSIS's analysis frames why this matters: U.S. export controls were designed around a bottleneck that domestic Chinese silicon is now demonstrably capable of routing around at frontier scale. The chip-sovereignty story and the model-capability story are the same story.

Two structural points the West tends to miss. First, the Chinese frontier is permissively open: GLM-5 ships under MIT license, Qwen variants under Apache 2.0 — the open-weights center of gravity has quietly moved east as Meta pivots Llama into a frozen tier behind closed Muse Spark. Second, the labs are tightly differentiated rather than redundant: DeepSeek on cost-per-capability, Qwen on agentic coding, Kimi on multi-agent swarms, Doubao on coordinated multimodal, GLM on sovereign silicon.

The contested claim worth naming: how much of the apparent parity is real frontier capability versus targeted benchmark optimization? Both camps have evidence. What is no longer defensible is the 2024 framing that there is a single Chinese flagship called DeepSeek and the rest is noise.