OpenRouter, the world's largest neutral LLM router, revealed a striking shift: as of May 2026, Chinese open-weight models consumed roughly 61% of the platform's total tokens. DeepSeek alone accounted for 17.6% in a single week. Two years ago, Meta's Llama owned the open-weight throne; today it has completely fallen off the leaderboard.
Behind this reversal lies an underappreciated truth: the intelligence of open-weight models consistently trails US frontier labs by only three to six months, and that gap is not expanding. For any organization scrutinizing its cloud bills, migrating workloads from frontier models to open-weight alternatives delivers real savings.
DeepSeek Drives Prices to the Floor
DeepSeek V4 Flash is the first open-weight model used directly in real agentic pipelines as a replacement for Anthropic- or OpenAI-class alternatives. The larger V4 Pro scored 80.6% on SWE-bench Verified, the highest among open-weight models. Pricing is aggressive: V4 Flash cache-hit input costs $0.0028 per million tokens, output $0.28; V4 Pro cache-miss input $0.30, output $0.50; deep reasoning model R1 output $2.19.
GLM Takes the Quality Crown
GLM 5.2, released by z-ai, scored 51 points on Artificial Analysis Intelligence Index v4.1, ranking first among open-weight models, ahead of Nemotron 3 Ultra (48), MiniMax M3 and DeepSeek V4 Pro (44). It trails only closed-source Claude Fable 5 by about five points. On the agentic benchmark GDPval-AA, it performs roughly on par with GPT-5.5. GLM 5.2 excels at planning and long-horizon tasks, at a higher inference cost: ~$0.447 per million input tokens and $3.31 output.
Timing adds intrigue: days before GLM 5.2 launched, a US export control order forced Anthropic to widely disable Fable 5 and Mythos 5 to prevent access by foreign nationals. One side sees closed models cut off due to geopolitics; the other offers MIT-licensed, near-frontier, self-hostable open weights.
Team USA: Nvidia's Nemotron 3 Ultra
China doesn't own the open-weight table. Nvidia recently released Nemotron 3 Ultra, scoring 48 points on the same index and becoming the strongest US open-weight model. With 550B parameters, 55B active, and a hybrid Mamba-2/Transformer architecture under OpenMDW license—weights, training data, recipes, and evaluation tools all open—Nvidia's play is blunt: the more open models are used, the more Blackwell chips, CUDA, and enterprise services it sells.

