Terminal-Bench 4.0 Cuts Noise, Retires Saturated Tasks
Terminal-Bench 4.0 calibrates CPU, memory, and timeout resources, fixes 19 tasks, and removes eight saturated or unreliable tasks. The resulting 66-task benchmark reduces infrastructure noise, but its scores are not directly comparable with version 3.0. [Announcement](https://www.tbench.ai/news/terminal-bench-4-0)
This is the unglamorous maintenance agent benchmarks need to remain credible. The catch is that a cleaner exam is still a new exam, requiring fresh baselines and careful reporting.
- –A flat eight-hour timeout reduces failures caused by arbitrary per-task limits.
- –Anthropic found resource configuration alone could shift Terminal-Bench scores by six percentage points, underscoring the value of calibration. [Research](https://www.anthropic.com/engineering/infrastructure-noise)
- –Removing tasks solved consistently by frontier systems keeps the leaderboard discriminative.
- –Developers should record the benchmark version, agent harness, effort setting, and resource limits when comparing results.
- –Because Terminal-Bench measures complete agent systems—not just base models—it is especially useful for evaluating real coding workflows. [Homepage](https://www.tbench.ai/)
DISCOVERED
1d ago
2026-09-03
PUBLISHED
1d ago
2026-09-03
RELEVANCE
AUTHOR
WorldofAI