Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol
Terminal-Bench has released version 4.0 of its benchmark for AI agents, recalibrating the time, CPU, and memory metrics used to evaluate task execution. The update fixes 19 tasks and removes 8 tasks that were affected by saturation, refusal, public solutions, or quality defects. All tasks now have a maximum execution time of 8 hours, a move intended to reduce the impact of timeouts and environment-related issues on final scores. The latest leaderboard places the Opus 5 model, running with Claude Code, at the top with 51.8%. Fable 5 follows in second place with 44.5%. GLM-5.3, paired with Claude Code, records 41.8% and moves into third place, surpassing the 37.3% posted by GPT-5.6 Sol combined with Codex. GLM-5.3 is the only non-Anthropic model inside the top three. In Terminal-Bench 3.0, GLM-5.3 was fourth with 32.4%, behind GPT-5.6 Sol's 34.6%. With the arrival of Terminal-Bench 4.0, GLM-5.3 has climbed to third and opened a 4.5-percentage-point lead over Sol.








