Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol

Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol

N
News Editor
2026-08-29 13:49:04
Terminal-Bench has released version 4.0 of its benchmark for AI agents, recalibrating the time, CPU, and memory metrics used to evaluate task execution. The update fixes 19 tasks and removes 8 tasks that were affected by saturation, refusal, public solutions, or quality defects. All tasks now have a maximum execution time of 8 hours, a move intended to reduce the impact of timeouts and environment-related issues on final scores. The latest leaderboard places the Opus 5 model, running with Claude Code, at the top with 51.8%. Fable 5 follows in second place with 44.5%. GLM-5.3, paired with Claude Code, records 41.8% and moves into third place, surpassing the 37.3% posted by GPT-5.6 Sol combined with Codex. GLM-5.3 is the only non-Anthropic model inside the top three. In Terminal-Bench 3.0, GLM-5.3 was fourth with 32.4%, behind GPT-5.6 Sol's 34.6%. With the arrival of Terminal-Bench 4.0, GLM-5.3 has climbed to third and opened a 4.5-percentage-point lead over Sol.

Terminal-Bench has shipped version 4.0 of its agent evaluation suite. The update recalibrates how task execution is measured across three dimensions: time, CPU usage, and memory consumption. It also revises the task set, fixing 19 tasks and removing 8 that exhibited saturation, refusal, public solutions, or quality defects.

Every task in the benchmark now runs under a single 8-hour execution cap. That change is designed to reduce the influence of timeouts and environment-related issues on final scores.

Leaderboard changes

The latest rankings put Opus 5 combined with Claude Code in first place at 51.8%. Fable 5 follows at 44.5%. GLM-5.3, paired with Claude Code, reaches third at 41.8%, edging out the 37.3% posted by GPT-5.6 Sol plus Codex. Among the top three, GLM-5.3 is the only entry that does not come from Anthropic.

A shift from the 3.0 ranking

Compared with Terminal-Bench 3.0, the swing is clear. GLM-5.3 previously held fourth place with 32.4%, below GPT-5.6 Sol's 34.6%. In the new 4.0 results, GLM-5.3 has moved up to third and now sits 4.5 percentage points ahead of Sol.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
40

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.