GLM

Terminal-Benc
2026-08-29 13:49:04

Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol

Terminal-Bench has released version 4.0 of its benchmark for AI agents, recalibrating the time, CPU, and memory metrics used to evaluate task execution. The update fixes 19 tasks and removes 8 tasks that were affected by saturation, refusal, public solutions, or quality defects. All tasks now have a maximum execution time of 8 hours, a move intended to reduce the impact of timeouts and environment-related issues on final scores. The latest leaderboard places the Opus 5 model, running with Claude Code, at the top with 51.8%. Fable 5 follows in second place with 44.5%. GLM-5.3, paired with Claude Code, records 41.8% and moves into third place, surpassing the 37.3% posted by GPT-5.6 Sol combined with Codex. GLM-5.3 is the only non-Anthropic model inside the top three. In Terminal-Bench 3.0, GLM-5.3 was fourth with 32.4%, behind GPT-5.6 Sol's 34.6%. With the arrival of Terminal-Bench 4.0, GLM-5.3 has climbed to third and opened a 4.5-percentage-point lead over Sol.

20
Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol
B.AI
2026-08-29 02:04:06

B.AI opens six free models to users, including GLM-5.3-Flash

B.AI said yesterday that its free model lineup is now fully in place, giving users no-threshold access to six major models: GLM-5.3-Flash, DeepSeek-V4-Flash, DeepSeek-V4-Flash-Vision-Exp, Hy3 from Tencent Hunyuan, MiMo-V2.5 from Xiaomi, and Qwen3.8-Flash. According to the platform, all six models are available for free. The newly introduced GLM-5.3-Flash is the model previously known for drawing broad attention as “Ox Alpha.” B.AI described it as the first natively multimodal model in the GLM-5 series. It carries 320B total parameters and 18B active parameters, uses a hybrid sparse and linear-attention architecture, and supports a 1M-token context window. B.AI said the free lineup covers a range of use cases, including text-only tasks, multimodal applications, code generation, and Agent orchestration. The platform added that both developers and heavy AI users can access advanced large-model capabilities at zero cost through B.AI, with access available via chat.b.ai/chat.

150
B.AI opens six free models to users, including GLM-5.3-Flash
Tencent
2026-08-28 07:32:42

Tencent Hunyuan Hy4 Preview Overhauls Architecture with DeepSeek Sparse Attention and Zhipu IndexCache

In a flash update on Aug 28, BlockBeats detailed the architectural changes behind Tencent's Hunyuan Hy4 preview. The model is bigger, but the bigger story is under the hood. Hy3 used full attention; Hy4 moves to Gated DSA, a gated form of DeepSeek Sparse Attention. On long inputs, instead of recomputing the entire context, the model selects the most relevant parts and focuses compute there. In Hy3 every layer had to decide what mattered; Hy4 layers Zhipu's IndexCache on top of DSA, so some layers can reuse earlier selections and cut repeated work. Tencent credits DeepSeek and GLM as inspirations. Residual structure also gets a rethink: a standard Transformer is often described as having a single information trunk, while Hy4's iHC widens that to four trunks so the network can retain more information at depth. MoE configuration moves from Hy3's 192 experts to 256 routing experts plus one shared expert, with each token choosing 8 routing experts while always passing through the shared expert.

90
Tencent Hunyuan Hy4 Preview Overhauls Architecture with DeepSeek Sparse Attention and Zhipu IndexCache
SenseTime
2026-08-28 07:14:49

SenseTime backs Zhipu's GLM-5.3-Flash with domestic compute and token services

SenseTime said on Aug. 27 that its AI infrastructure platform, SenseTime Grand Infrastructure, is providing domestic computing power support and token services for Zhipu's recently launched and open-sourced GLM-5.3-Flash model, according to Securities Times. The development was described in the report as another benchmark case for the scaled commercial use of domestic computing power in China. SenseTime also disclosed operating figures for its platform. The company said its Grand Infrastructure platform reached an average daily token service volume of 2.42 trillion in July. It added that the figure is expected to exceed 10 trillion tokens per day by the end of 2026. The report did not provide more technical details about the deployment, service scope, or commercial terms.

60
SenseTime backs Zhipu's GLM-5.3-Flash with domestic compute and token services
Tencent
2026-08-28 06:47:48

Tencent Hy4 edges GLM-5.3 and Kimi K3 in blind test, prices output 82% lower

Tencent has shared blind-test results for its Hy4 preview model, based on real engineering tasks from inside the company. In a test involving 163 internal experts and 203 tasks, Hy4 scored an average of 2.99 out of 4, edging out GLM-5.3 (2.92) and Kimi K3 (2.94). Against GLM-5.3, Hy4 won 46.8% of comparisons, tied 12.8% and lost 40.4%; against Kimi K3, it won 51.2%, tied 7.9% and lost 40.9%. The tasks were drawn from Tencent's actual engineering workload rather than a standard fixed benchmark. However, Hy4 does not lead across all public benchmarks. The company's results show it trading wins and losses with GLM-5.3 and Kimi K3, and it still trails GLM-5.3 on code and cybersecurity tests such as DeepSWE and CyberGym. Pricing is Hy4's clearest edge. At 6 yuan per million tokens for input and 18 yuan for output, it is 25% and around 36% cheaper than GLM-5.3 on those two metrics, and 70% and 82% cheaper than Kimi K3. Cache-hit inference costs just 0.3 yuan per million tokens, 85% less than the 2 yuan charged by both GLM-5.3 and Kimi K3.

100
Tencent Hy4 edges GLM-5.3 and Kimi K3 in blind test, prices output 82% lower
NVIDIA
2026-08-27 13:03:09

NVIDIA posts 106% revenue growth, eyes $13 billion Hugging Face deal as Apple foldable iPhone is tipped for Sept. 10

TechFlow’s Aug. 27 roundup pulled together a wide set of signals across AI, crypto, chips, public equities, and consumer hardware. The biggest headline was NVIDIA’s quarterly result: Q2 revenue reached $96.2 billion, up 106% year over year, with data center revenue at $89 billion and next-quarter guidance at $108 billion. The company was also reported to be acquiring Hugging Face for about $13 billion, a move that would put one of the largest open-model communities under the control of the leading AI chip vendor. On the model side, Alibaba’s Qwen3.8-Flash was described as cutting training cost by 90% while activating only 6B parameters and outperforming Claude Opus 4.6. GLM 5.3 Flash, meanwhile, was said to be running entirely on domestic Chinese AI chip clusters, processing 70 trillion tokens in a week without relying on NVIDIA’s high-end GPUs. In crypto, Glassnode said Bitcoin faces a real-demand test above $83,000, where thicker liquidity means any move higher would need to absorb genuine sell pressure rather than depend on leverage. Polymarket also launched wagers tied to which countries may send warships through the Strait of Hormuz. Elsewhere, Apple’s first foldable iPhone was reported to be set for a Sept. 10 launch.

130
NVIDIA posts 106% revenue growth, eyes $13 billion Hugging Face deal as Apple foldable iPhone is tipped for Sept. 10
US inflation
2026-08-27 04:54:09

Hotter U.S. PCE revives rate-hike bets as Nvidia earnings steady the AI trade

U.S. markets turned cautious after July personal consumption expenditures data came in hotter than expected, pushing traders to lift bets on additional Federal Reserve tightening. The report showed headline PCE rising 3.7% year over year and 0.2% month over month, while core PCE stayed at 3.3% annually and 0.2% monthly. Treasury yields moved higher across the curve, with the 10-year near 4.66%, the 2-year around 4.22%, and the dollar index climbing to roughly 99.15. Gold fell under pressure from a firmer dollar and higher rate expectations, while oil traded weaker as rhetoric around Iran kept geopolitical risk in focus. Another inflation thread is building in food markets. Attacks on Black Sea ports cut Ukraine’s August grain shipments to about 20% of potential capacity, wheat futures on CBOT touched their highest level in nearly three years, and fertilizer supply disruptions tied to Hormuz added to cost pressure. HSBC warned that the 2026/27 global grain market could post its first supply-demand gap since 2020/21 and the largest shortage since 2006/07, while JPMorgan said global food inflation could rise from 2.8% in the first half of 2026 to 5% in the first half of 2027. After the bell, Nvidia delivered the day’s biggest market jolt. The chipmaker reported $96.2 billion in Q2 revenue and $89.0 billion from data center sales, both well ahead of expectations, and guided for about 70% revenue growth in fiscal 2028. The results helped revive AI spending sentiment and lifted software, storage, optical networking, and cybersecurity names in after-hours trading.

180
Hotter U.S. PCE revives rate-hike bets as Nvidia earnings steady the AI trade
Zhipu
2026-08-27 03:40:29

Zhipu to Open-Source GLM-5.3 at 10:00 a.m. Beijing Time on Aug. 28

Zhipu said its flagship GLM-5.3 model will be open-sourced at 10:00 a.m. Beijing time on Aug. 28, with model weights becoming available for download at that time. The company said GLM-5.3 had already gone live in Coding Plan on Aug. 14 and later opened API access, but the weight release came two weeks later for security reasons. According to Zhipu, the model’s cybersecurity capabilities improved faster than expected during post-training. The company said the model not only became better at finding vulnerabilities, but also started to plan multi-step exploit chains, prompting internal security evaluation and hardening before the weights were released. In Zhipu’s published benchmarks, GLM-5.3 scored 84.5% on the vulnerability-discovery test CyberGym, slightly above Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. On ExploitBench, which focuses more on exploit execution, GLM-5.3 rose from GLM-5.2’s 24.4% to 54.4%, though it still trailed Mythos 5 at 78.0%. Zhipu also said the model found 2,436 vulnerabilities across 269 open-source projects, including 1,097 classified as critical or high severity.

100
Zhipu to Open-Source GLM-5.3 at 10:00 a.m. Beijing Time on Aug. 28