ARC-AGI-3 Public Games Aren’t Eval Scores
François Chollet clarifies that ARC-AGI-3’s public games are a demonstration set, not training data or an evaluation set. Scores on them should not be interpreted as evidence of progress on the benchmark’s private tests.
ARC-AGI-3 is deliberately moving benchmark credibility behind private holdouts, making public-game optimization a poor proxy for general intelligence.
- –The 25 public environments demonstrate the format and basic mechanics.
- –Private and semi-private environments test broader, out-of-distribution generalization.
- –Public scores can be inflated through task-specific harnesses, human replays, or targeted training.
- –Developers should use the public set for integration testing and agent iteration, not leaderboard claims.
- –The distinction highlights how easily benchmark results become misleading when test data is exposed.
DISCOVERED
1h ago
2026-08-14
PUBLISHED
1h ago
2026-08-14
RELEVANCE
AUTHOR
fchollet