YOU ARE VIEWING ONE ITEM FROM THE AICRIER FEED

Terminal-Bench-Science finds agents stuck at 30%

AICrier tracks AI developer news across Product Hunt, GitHub, Hacker News, YouTube, X, arXiv, and more. This page keeps the article you opened front and center while giving you a path into the live feed.

// WHAT AICRIER DOES

7+

TRACKED FEEDS

24/7

SCRAPED FEED

Short summaries, external links, screenshots, relevance scoring, tags, and featured picks for AI builders.

Terminal-Bench-Science finds agents stuck at 30%
OPEN LINK ↗
// 1h agoBENCHMARK RESULT

Terminal-Bench-Science finds agents stuck at 30%

Terminal-Bench-Science 0.1 evaluates agents on 70 expert-curated workflows spanning five scientific domains, grading concrete outputs such as analyses, simulations, proofs, code, and data products. Claude Opus 5 with Claude Code leads at a 30.0% resolution rate, showing how far scientific agents remain from reliable autonomy.

// ANALYSIS

The important result is not the leaderboard winner but the low ceiling: real scientific work still breaks frontier agents routinely.

  • Tasks reflect research workflows rather than textbook questions, covering inference, simulation, optimization, theorem proving, image reconstruction, and scientific machine learning.
  • Three independent trials per task and reproducible, task-specific grading make the results more operationally meaningful than generic capability claims.
  • Claude Opus 5 leads GPT-5.6 Sol at 30.0% versus 22.4%; GPT-5.6 Sol matches Claude Fable 5's performance at substantially lower evaluation cost.
  • The open, continuously updated benchmark gives scientists a practical way to define what “AI for science” should actually accomplish.
  • With 70 tasks across five domains, future versions will need broader coverage to reduce variance and better represent specialized research.
// TAGS
terminal-bench-sciencebenchmarkevaluationagentresearchopen-sourcecli

DISCOVERED

1h ago

2026-09-02

PUBLISHED

1h ago

2026-09-02

RELEVANCE

8/ 10

AUTHOR

DIY Smart Code