Terminal-Bench-Science finds agents stuck at 30%
Terminal-Bench-Science 0.1 evaluates agents on 70 expert-curated workflows spanning five scientific domains, grading concrete outputs such as analyses, simulations, proofs, code, and data products. Claude Opus 5 with Claude Code leads at a 30.0% resolution rate, showing how far scientific agents remain from reliable autonomy.
The important result is not the leaderboard winner but the low ceiling: real scientific work still breaks frontier agents routinely.
- –Tasks reflect research workflows rather than textbook questions, covering inference, simulation, optimization, theorem proving, image reconstruction, and scientific machine learning.
- –Three independent trials per task and reproducible, task-specific grading make the results more operationally meaningful than generic capability claims.
- –Claude Opus 5 leads GPT-5.6 Sol at 30.0% versus 22.4%; GPT-5.6 Sol matches Claude Fable 5's performance at substantially lower evaluation cost.
- –The open, continuously updated benchmark gives scientists a practical way to define what “AI for science” should actually accomplish.
- –With 70 tasks across five domains, future versions will need broader coverage to reduce variance and better represent specialized research.
DISCOVERED
1h ago
2026-09-02
PUBLISHED
1h ago
2026-09-02
RELEVANCE
AUTHOR
DIY Smart Code