GLM-5

B.AI
2026-08-31 03:27:17

B.AI's Cumulative Token Throughput Passes 5.74 Trillion as Free Model Access Continues

B.AI, the AI computing platform, said on August 31 that its cumulative token throughput has officially surpassed 5.74 trillion (5T+), a historic milestone for the platform's ongoing free-access program. The announcement arrived as the campaign for frontier large language models keeps running at full tilt. Under massive concurrent requests and demanding Agent workloads, B.AI's industrial-grade infrastructure has demonstrated stable capacity across the board, the company said. Six flagship models are now available on the platform: GLM-5.3-Flash (Ox Alpha), Qwen3.8-Flash, DeepSeek-V4-Flash, DeepSeek-V4-Flash-Vision-Exp, Tencent Hy3, and Xiaomi MiMo-V2.5 — all with zero entry barriers, unlimited usage, and no cost. The lineup covers a broad set of tasks, including code generation, ultra-long text processing, multimodal visual understanding, and Agent workflow deployment. Developers who sign in from today onward can call the entire suite at zero cost, free of charge and with unlimited compute integrated directly. B.AI pitches the setup as genuine "computing freedom" for every developer and AI power user alike.

120
B.AI's Cumulative Token Throughput Passes 5.74 Trillion as Free Model Access Continues
SemiAnalysis
2026-08-30 16:01:21

SemiAnalysis flags widespread Neocloud misconfigurations after cross-tenant RCE tests

SemiAnalysis, an independent semiconductor and AI research firm, published a deep-dive report on Neocloud security on Aug. 30, detailing multiple cross-tenant vulnerabilities uncovered during ClusterMAX 3.0 testing. Over a four-month assessment spanning 25 vendors and 32 clusters, the team said it was able to achieve several cross-tenant remote code execution, or RCE, outcomes using only publicly known vulnerabilities and basic configuration checks. Entities affected included banks, telecom companies, universities, research institutions, AI labs and even one national intelligence agency. The report lists a range of recurring issues, including shared Kubernetes control planes, container escapes, exposed BMC/IPMI management networks, improperly configured InfiniBand security keys, an unhardened default trust mode on BlueField DPUs, Grafana dashboards using god-mode API keys, and missing VXLAN isolation on front-end networks. It also describes a chained case in which a misconfigured shared vCluster combined with software versions lagging by two years allowed a proof-of-concept cross-tenant RCE to be completed within hours. SemiAnalysis also disputes the common claim that AI has fundamentally accelerated the pace of cybersecurity risk, saying CVE data tied to Nvidia GPU drivers, CUDA, PyTorch, Kubernetes, Docker and the Linux kernel did not show a significant increase after AI coding models became widely used.

210
SemiAnalysis flags widespread Neocloud misconfigurations after cross-tenant RCE tests
Terminal-Benc
2026-08-29 13:49:04

Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol

Terminal-Bench has released version 4.0 of its benchmark for AI agents, recalibrating the time, CPU, and memory metrics used to evaluate task execution. The update fixes 19 tasks and removes 8 tasks that were affected by saturation, refusal, public solutions, or quality defects. All tasks now have a maximum execution time of 8 hours, a move intended to reduce the impact of timeouts and environment-related issues on final scores. The latest leaderboard places the Opus 5 model, running with Claude Code, at the top with 51.8%. Fable 5 follows in second place with 44.5%. GLM-5.3, paired with Claude Code, records 41.8% and moves into third place, surpassing the 37.3% posted by GPT-5.6 Sol combined with Codex. GLM-5.3 is the only non-Anthropic model inside the top three. In Terminal-Bench 3.0, GLM-5.3 was fourth with 32.4%, behind GPT-5.6 Sol's 34.6%. With the arrival of Terminal-Bench 4.0, GLM-5.3 has climbed to third and opened a 4.5-percentage-point lead over Sol.

100
Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol
B.AI
2026-08-29 02:04:06

B.AI opens six free models to users, including GLM-5.3-Flash

B.AI said yesterday that its free model lineup is now fully in place, giving users no-threshold access to six major models: GLM-5.3-Flash, DeepSeek-V4-Flash, DeepSeek-V4-Flash-Vision-Exp, Hy3 from Tencent Hunyuan, MiMo-V2.5 from Xiaomi, and Qwen3.8-Flash. According to the platform, all six models are available for free. The newly introduced GLM-5.3-Flash is the model previously known for drawing broad attention as “Ox Alpha.” B.AI described it as the first natively multimodal model in the GLM-5 series. It carries 320B total parameters and 18B active parameters, uses a hybrid sparse and linear-attention architecture, and supports a 1M-token context window. B.AI said the free lineup covers a range of use cases, including text-only tasks, multimodal applications, code generation, and Agent orchestration. The platform added that both developers and heavy AI users can access advanced large-model capabilities at zero cost through B.AI, with access available via chat.b.ai/chat.

290
B.AI opens six free models to users, including GLM-5.3-Flash
SenseTime
2026-08-28 07:14:49

SenseTime backs Zhipu's GLM-5.3-Flash with domestic compute and token services

SenseTime said on Aug. 27 that its AI infrastructure platform, SenseTime Grand Infrastructure, is providing domestic computing power support and token services for Zhipu's recently launched and open-sourced GLM-5.3-Flash model, according to Securities Times. The development was described in the report as another benchmark case for the scaled commercial use of domestic computing power in China. SenseTime also disclosed operating figures for its platform. The company said its Grand Infrastructure platform reached an average daily token service volume of 2.42 trillion in July. It added that the figure is expected to exceed 10 trillion tokens per day by the end of 2026. The report did not provide more technical details about the deployment, service scope, or commercial terms.

200
SenseTime backs Zhipu's GLM-5.3-Flash with domestic compute and token services
Tencent
2026-08-28 06:47:48

Tencent Hy4 edges GLM-5.3 and Kimi K3 in blind test, prices output 82% lower

Tencent has shared blind-test results for its Hy4 preview model, based on real engineering tasks from inside the company. In a test involving 163 internal experts and 203 tasks, Hy4 scored an average of 2.99 out of 4, edging out GLM-5.3 (2.92) and Kimi K3 (2.94). Against GLM-5.3, Hy4 won 46.8% of comparisons, tied 12.8% and lost 40.4%; against Kimi K3, it won 51.2%, tied 7.9% and lost 40.9%. The tasks were drawn from Tencent's actual engineering workload rather than a standard fixed benchmark. However, Hy4 does not lead across all public benchmarks. The company's results show it trading wins and losses with GLM-5.3 and Kimi K3, and it still trails GLM-5.3 on code and cybersecurity tests such as DeepSWE and CyberGym. Pricing is Hy4's clearest edge. At 6 yuan per million tokens for input and 18 yuan for output, it is 25% and around 36% cheaper than GLM-5.3 on those two metrics, and 70% and 82% cheaper than Kimi K3. Cache-hit inference costs just 0.3 yuan per million tokens, 85% less than the 2 yuan charged by both GLM-5.3 and Kimi K3.

200
Tencent Hy4 edges GLM-5.3 and Kimi K3 in blind test, prices output 82% lower
US inflation
2026-08-27 04:54:09

Hotter U.S. PCE revives rate-hike bets as Nvidia earnings steady the AI trade

U.S. markets turned cautious after July personal consumption expenditures data came in hotter than expected, pushing traders to lift bets on additional Federal Reserve tightening. The report showed headline PCE rising 3.7% year over year and 0.2% month over month, while core PCE stayed at 3.3% annually and 0.2% monthly. Treasury yields moved higher across the curve, with the 10-year near 4.66%, the 2-year around 4.22%, and the dollar index climbing to roughly 99.15. Gold fell under pressure from a firmer dollar and higher rate expectations, while oil traded weaker as rhetoric around Iran kept geopolitical risk in focus. Another inflation thread is building in food markets. Attacks on Black Sea ports cut Ukraine’s August grain shipments to about 20% of potential capacity, wheat futures on CBOT touched their highest level in nearly three years, and fertilizer supply disruptions tied to Hormuz added to cost pressure. HSBC warned that the 2026/27 global grain market could post its first supply-demand gap since 2020/21 and the largest shortage since 2006/07, while JPMorgan said global food inflation could rise from 2.8% in the first half of 2026 to 5% in the first half of 2027. After the bell, Nvidia delivered the day’s biggest market jolt. The chipmaker reported $96.2 billion in Q2 revenue and $89.0 billion from data center sales, both well ahead of expectations, and guided for about 70% revenue growth in fiscal 2028. The results helped revive AI spending sentiment and lifted software, storage, optical networking, and cybersecurity names in after-hours trading.

360
Hotter U.S. PCE revives rate-hike bets as Nvidia earnings steady the AI trade
Zhipu
2026-08-27 03:40:29

Zhipu to Open-Source GLM-5.3 at 10:00 a.m. Beijing Time on Aug. 28

Zhipu said its flagship GLM-5.3 model will be open-sourced at 10:00 a.m. Beijing time on Aug. 28, with model weights becoming available for download at that time. The company said GLM-5.3 had already gone live in Coding Plan on Aug. 14 and later opened API access, but the weight release came two weeks later for security reasons. According to Zhipu, the model’s cybersecurity capabilities improved faster than expected during post-training. The company said the model not only became better at finding vulnerabilities, but also started to plan multi-step exploit chains, prompting internal security evaluation and hardening before the weights were released. In Zhipu’s published benchmarks, GLM-5.3 scored 84.5% on the vulnerability-discovery test CyberGym, slightly above Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. On ExploitBench, which focuses more on exploit execution, GLM-5.3 rose from GLM-5.2’s 24.4% to 54.4%, though it still trailed Mythos 5 at 78.0%. Zhipu also said the model found 2,436 vulnerabilities across 269 open-source projects, including 1,097 classified as critical or high severity.

210
Zhipu to Open-Source GLM-5.3 at 10:00 a.m. Beijing Time on Aug. 28