Opus 5 hits 30% on ARC-AGI-3 benchmark
François Chollet announced that Opus 5 achieved a state-of-the-art score of 30% on the ARC-AGI-3 benchmark, which measures an AI system's ability to solve novel problems without prior exposure. The result sparked community discussion regarding whether exposure to earlier ARC benchmark iterations contributes to performance gains.
Opus 5 reaching 30% on ARC-AGI-3 marks a notable leap in novel problem-solving capability, though questions persist around how benchmark evolution impacts out-of-distribution evaluation.
- –Opus 5 achieves a landmark 30% score on ARC-AGI-3, demonstrating measurable progress in zero-shot pattern abstraction.
- –ARC-AGI benchmarks target core reasoning performance where standard pre-training scaling typically yields diminishing returns.
- –Researchers debate whether performance on newer ARC iterations benefits from models being optimized against the paradigms established in ARC 1 and 2.
DISCOVERED
2h ago
2026-07-27
PUBLISHED
2d ago
2026-07-24
RELEVANCE
AUTHOR
fchollet