ARC-AGI-4 targets autonomous open-ended invention
François Chollet announced that the upcoming ARC-AGI-4 and ARC 5 benchmarks will evaluate autonomous open-ended invention rather than closed reasoning puzzles. Scheduled for an open-source release in Q1 next year, ARC-AGI-4 aims to measure scientific innovation where humans still vastly outperform AI.
Benchmarking open-ended invention is the essential next hurdle for measuring genuine AGI as test-time compute and search techniques increasingly saturate closed-world reasoning puzzles. Evaluating open-ended discovery requires models to autonomously formulate hypotheses and explore unbounded solution spaces, moving past memorized heuristics and narrow grid transformations. By serving as an explicit philosophical counterweight to industry proposals for pacing frontier development, ARC-AGI-4 reaffirms open-source benchmarks as vital for broad scientific progress while exposing the algorithmic limitations of pure pretraining and test-time search.
DISCOVERED
1h ago
2026-09-12
PUBLISHED
1h ago
2026-09-12
RELEVANCE
AUTHOR
fchollet