YOU ARE VIEWING ONE ITEM FROM THE AICRIER FEED

Terminal-Bench 4.0 Cuts Noise, Retires Saturated Tasks

AICrier tracks AI developer news across Product Hunt, GitHub, Hacker News, YouTube, X, arXiv, and more. This page keeps the article you opened front and center while giving you a path into the live feed.

// WHAT AICRIER DOES

7+

TRACKED FEEDS

24/7

SCRAPED FEED

Short summaries, external links, screenshots, relevance scoring, tags, and featured picks for AI builders.

Terminal-Bench 4.0 Cuts Noise, Retires Saturated Tasks
OPEN LINK ↗
// 1d agoPRODUCT UPDATE

Terminal-Bench 4.0 Cuts Noise, Retires Saturated Tasks

Terminal-Bench 4.0 calibrates CPU, memory, and timeout resources, fixes 19 tasks, and removes eight saturated or unreliable tasks. The resulting 66-task benchmark reduces infrastructure noise, but its scores are not directly comparable with version 3.0. [Announcement](https://www.tbench.ai/news/terminal-bench-4-0)

// ANALYSIS

This is the unglamorous maintenance agent benchmarks need to remain credible. The catch is that a cleaner exam is still a new exam, requiring fresh baselines and careful reporting.

  • A flat eight-hour timeout reduces failures caused by arbitrary per-task limits.
  • Anthropic found resource configuration alone could shift Terminal-Bench scores by six percentage points, underscoring the value of calibration. [Research](https://www.anthropic.com/engineering/infrastructure-noise)
  • Removing tasks solved consistently by frontier systems keeps the leaderboard discriminative.
  • Developers should record the benchmark version, agent harness, effort setting, and resource limits when comparing results.
  • Because Terminal-Bench measures complete agent systems—not just base models—it is especially useful for evaluating real coding workflows. [Homepage](https://www.tbench.ai/)
// TAGS
terminal-bench-4.0evaluationbenchmarkcoding-agentai-codingtestingopen-source

DISCOVERED

1d ago

2026-09-03

PUBLISHED

1d ago

2026-09-03

RELEVANCE

9/ 10

AUTHOR

WorldofAI