Epoch AI audit boosts FrontierMath scores
Epoch AI conducted an audit of its FrontierMath benchmark (Tiers 1–4), identifying evaluation errors in roughly 42% of the problems. The corrected version 2 shows that frontier models, including GPT-5.5, are performing significantly better than previously reported, achieving scores as high as 85% on Tiers 1–3 and bringing the benchmark closer to saturation.
Benchmark quality control is the new bottleneck in AI evaluation, as shown by models quickly saturating even research-level benchmarks once evaluation errors are ironed out.
* The audit, partially aided by GPT-5.5, corrected flaws in 42% of the original benchmark problems.
* Corrected scores show GPT-5.5 scoring up to 85% on Tiers 1–3, while other top models scored 76% on Tier 4.
* Rapid performance gains raise concerns about how quickly frontier benchmarks are being saturated by the latest generation of models.
DISCOVERED
51d ago
2026-06-12
PUBLISHED
51d ago
2026-06-12
RELEVANCE
AUTHOR
mark_k
