GLM-5.3 Takes Third in Terminal-Bench 4.0
GLM-5.3 scored 41.82% with Claude Code at maximum reasoning, placing third on Terminal-Bench 4.0—ahead of GPT-5.6 Sol but behind Claude Opus 5 and Claude Fable 5. The result follows Z.ai’s release of the model’s full weights for downloadable, local deployment.
GLM-5.3’s performance makes open-weight models credible contenders for real terminal-based coding work, but the leaderboard is measuring the model-agent combination—not the model in isolation.
- –GLM-5.3 reached 41.82%, versus 37.27% for GPT-5.6 Sol.
- –Terminal-Bench 4.0 covers 66 tasks with five trials and up to eight hours per task.
- –The benchmark changed substantially, fixing 19 tasks and removing eight, so scores are not directly comparable with earlier versions.
- –Claude Code’s harness and maximum reasoning setting materially influence the result.
- –Full weights make the result more actionable for developers, though practical local deployment remains hardware-intensive.
DISCOVERED
1h ago
2026-08-29
PUBLISHED
2h ago
2026-08-29
RELEVANCE
AUTHOR
AGTPinsights