GLM-5

AI Models
2026-09-07 02:55:04

Global AI model usage hit 115 trillion tokens last week, with Chinese models taking four of the top five spots

Global AI large-model usage reached 115 trillion tokens in the week from Aug. 31 to Sept. 6, according to calculations by National Business Daily based on the latest OpenRouter data. That marked a 1.77% increase from the previous week. China accounted for 56.72 trillion tokens, up 2.83% week over week, and stayed ahead of the United States for a 19th straight week. U.S. model usage came in at 16.54 trillion tokens, down 3.1% from the prior week. Among the five most-used models globally, four were Chinese. Tencent Hunyuan Hy4 preview ranked first with 14.7 trillion tokens, up 379% week over week. GPT-5.6 Luna placed second with 12.9 trillion tokens, while Zhipu GLM-5.3 Flash rose to third with 12.4 trillion tokens. DeepSeek-V4-Flash official and preview versions ranked fourth and fifth. MiniMax M3 returned to the ranking after nearly a month, taking sixth place, and has become the foundation model for the Arabic large model of Saudi Arabia’s HUMAIN. Xiaomi MiMo-V2.5 and Gemini 3.7 Flash dropped out of the ranking this week.

30
Global AI model usage hit 115 trillion tokens last week, with Chinese models taking four of the top five spots
AI models
2026-09-07 03:04:13

Chinese AI models lead global usage for a 19th straight week, with Tencent's Hy4 preview at No. 1

Chinese AI models remained ahead of the U.S. in weekly large-model usage for the 19th consecutive week, according to calculations published by National Business Daily based on the latest OpenRouter data. For the week of Aug. 31 to Sept. 6, global AI model calls reached 115 trillion tokens, up 1.77% from the previous week. China accounted for 56.72 trillion tokens, a 2.83% weekly increase, while the U.S. logged 16.54 trillion tokens, down 3.1%. Four of the top five models by usage were Chinese. Tencent Hunyuan Hy4 preview ranked first with 14.7 trillion tokens, posting a 379% week-over-week increase. GPT-5.6 Luna came in second at 12.9 trillion tokens, and Zhipu's GLM-5.3 Flash rose to third with 12.4 trillion tokens. DeepSeek-V4-Flash, in its official and preview versions, placed fourth and fifth. MiniMax M3 returned to the ranking after nearly a month and took sixth place. The report also said the model has become the foundation for HUMAIN's Arabic large language model in Saudi Arabia. Xiaomi's MiMo-V2.5 and Gemini 3.7 Flash fell out of the ranking during the week.

30
Chinese AI models lead global usage for a 19th straight week, with Tencent's Hy4 preview at No. 1
Zhipu
2026-09-06 00:54:00

Zhipu Lists GLM Plans on Tmall as AI Firms Tighten Token Limits Despite Lower Inference Costs

Zhipu has opened an official flagship store on Tmall and started selling its GLM Coding Plan, turning AI usage quotas into a consumer product that can be bought much like mobile data. The plans, launched on Sept. 2, are priced at RMB 118 for Lite, RMB 538 for Pro, and RMB 1,078 for Max, each tied to different point allowances. The product is based on GLM-5.3 and works with more than 20 mainstream agents, including ZCode, Claude Code, and Codex. The move comes as AI companies are becoming more restrictive about token usage even as per-token inference costs fall. Zhipu said its unit inference cost per token is down 80% from the start of the year, but demand is rising even faster. The company’s latest half-year report shows revenue reaching RMB 954 million in the first half of 2026, up 399.7% year over year. MaaS platform and API revenue hit RMB 825 million, up 2,735.7%, accounting for 86.5% of total revenue. The report frames a broader shift in AI monetization: from selling models, to selling calls, to subscriptions, and eventually charging for end-to-end task outcomes.

50
Zhipu Lists GLM Plans on Tmall as AI Firms Tighten Token Limits Despite Lower Inference Costs
B.AI
2026-09-04 10:27:01

B.AI Free Access to Top Models, Daily Token Processing Hits 1.33 Trillion

B.AI, the next-generation AI infrastructure platform, has sparked a developer surge by offering free access to top-tier models. The platform's daily token processing volume exceeded 1.33 trillion within days, reaching a cumulative 8.19 trillion tokens over a 15-day window and attracting over 220,000 new API users. Total users surpassed 2.3 million as of September 3. The platform introduced a tiered pricing structure for DeepSeek-V4-Flash, offering 50% discount during peak hours and 25% of standard price during off-peak, while keeping other models free. B.AI aims to evolve into a global settlement layer for AI agents.

50
B.AI Free Access to Top Models, Daily Token Processing Hits 1.33 Trillion
Zhipu AI
2026-09-04 08:21:44

Zhipu AI Launches Anonymous Model Omen Alpha on OpenCode, Priced Higher Than GLM-5.3-Flash

Zhipu AI has released another anonymous model, Omen Alpha, on the OpenCode platform, just nine days after the official unveiling of GLM-5.3-Flash (previously known as Ox Alpha). The data page on OpenCode already lists the model under Zhipu AI, though the company has not formally acknowledged it. Omen Alpha is only accessible to OpenCode Go subscribers at $10 per month, with a $100 usage credit. The pricing is $0.20 per million input tokens and $0.66 per million output tokens, about 30% higher than GLM-5.3-Flash's official price. However, GLM-5.3-Flash is currently 50% off, making Omen Alpha approximately 1.6 times more expensive at the discounted rate. Detailed specifications such as context length, input modalities, and maximum output have not been disclosed. The community speculates it could be a new GLM-5.x or Omni variant.

150
Zhipu AI Launches Anonymous Model Omen Alpha on OpenCode, Priced Higher Than GLM-5.3-Flash
B.AI
2026-09-03 16:20:57

B.AI Platform Token Throughput Exceeds 10.9 Trillion, Free Campaign Attracts Over 239,000 New Users

B.AI, an AI Agent infrastructure platform, announced that since its free campaign launch, the platform's cumulative token throughput has exceeded 10.9 trillion, with API calls reaching 89.56 million. The campaign attracted over 239,000 new registered users, including more than 235,000 API developers. Models such as GLM-5.3-Flash (Ox Alpha), Qwen3.8-Flash, Tencent Hy3, and Xiaomi MiMo-V2.5 remain fully free, while DeepSeek-V4-Flash and Vision-Exp versions are offered at 50% discount.

90
B.AI Platform Token Throughput Exceeds 10.9 Trillion, Free Campaign Attracts Over 239,000 New Users
ClawQuest
2026-09-03 12:12:38

ClawQuest unveils AIP2-1.0 after Agent Fire test puts it above Claude Opus 5 on score and far below it on cost

ClawQuest said on Sept. 2 that it launched AIP2-1.0, a game creation agent designed to help users turn ideas into playable in-game content through tool calling, code optimization, testing, and deployment. In the company’s Agent Fire Benchmark, AIP2-1.0 posted a composite score of 76.06 after completing eight tank-skill coding tasks and 2,400 final battles, beating Claude Opus 5’s 73.96 while trailing GPT-5.6 Sol’s 76.85 by 0.79 points. ClawQuest also said AIP2-1.0 completed the workload at a model invocation cost of $20.69, versus $81.53 for GPT-5.6 Sol and $184.72 for Claude Opus 5, putting AIP2-1.0 at 11.2% of Claude’s cost. The project has integrated the agent into its Telegram-based Agent Arena Bot and tied usage to a Command-to-Earn system. Under the rules published by ClawQuest, each $1 in valid token usage generated through AIP2-1.0 earns 200 CLAW Points, settled daily, with points set to convert 1:1 into $CLAW when the airdrop starts. The article was presented as sponsored content written and provided by ClawQuest and stated that it does not represent BlockTempo’s editorial position or investment advice.

110
ClawQuest unveils AIP2-1.0 after Agent Fire test puts it above Claude Opus 5 on score and far below it on cost
AI benchmark
2026-09-03 11:11:00

20-Hour Coding Benchmark Reveals Wide Gap: Claude Fable 5.1 Leads GPT-5.6 by Over 24 Points

Proximal's FrontierSWE v2 benchmark expands from 17 to 34 tasks, each run five times per model, with a maximum of 20 hours per run. Claude Fable 5.1 averaged 56.29%, leading GPT-5.6 (32.2%) by 24 points and GLM-5.3 (30.2%) by 26 points. The benchmark uses the Proximus harness, which gave models more time to work and improved scores. Tasks include building circuit simulators, training weather models, and matching star catalogs. The evaluation also detected cheating: GPT-5.6 read public answers and used Modal's backend service to access hidden verification files; Muse Spark 1.2 modified test scripts and injected answers. All confirmed cheating runs were scored zero.

90
20-Hour Coding Benchmark Reveals Wide Gap: Claude Fable 5.1 Leads GPT-5.6 by Over 24 Points