GPT-5.5 tops Agents' Last Exam benchmark
UC Berkeley's new Agents' Last Exam benchmark evaluates AI agents on complex, long-horizon professional workflows rather than academic knowledge. Initial results place OpenAI's GPT-5.5 at the top of the leaderboard with a 24.0% score, narrowly defeating Anthropic's Claude Fable 5 at 22.0%.
While GPT-5.5's narrow lead over Claude Fable 5 marks a milestone for OpenAI, the overall low scores (under 25%) highlight how far frontier AI models still have to go before they can reliably execute complex, multi-step professional work autonomously.
- –Hard Reality Check: Top-tier models scoring ~24% overall and even lower on advanced tiers prove that academic benchmark performance does not translate directly to professional-grade productivity.
- –Living Evaluation: The collaboration of 300+ experts and code-graded, reproducible testing establishes a new standard for agent evaluation, reducing data contamination and evaluation bias.
- –Narrow Frontier Gap: The 2.0% performance gap between GPT-5.5 and Claude Fable 5 indicates that frontier AI labs remain in a tight neck-and-neck race for agentic supremacy.
DISCOVERED
51d ago
2026-06-11
PUBLISHED
51d ago
2026-06-11
RELEVANCE
AUTHOR
GenAISpotlight