VISTA Takes Claude to ARC-AGI-3 Perfect Score
MIT’s VISTA harness gives Claude Opus 5.0 lossless visual memory it can actively inspect, enabling a perfect score across all 25 public ARC-AGI-3 games with 57.4% fewer actions than first-time human players. The result required no additional model training.
The breakthrough is less about a smarter model than a better interaction loop: giving multimodal agents durable, revisitable visual context dramatically improves long-horizon performance.
- –VISTA archives raw visual frames so Claude can revisit evidence instead of relying on lossy summaries.
- –The result shows that agent harnesses, memory, tools, and feedback can matter as much as the underlying model.
- –ARC-AGI-3 measures action efficiency, making this a stronger systems result than a simple completion-rate benchmark.
- –The claim is limited to the public set; the authors note that model exposure to those games cannot be ruled out, leaving private-set generalization unresolved.
DISCOVERED
1h ago
2026-10-03
PUBLISHED
1h ago
2026-10-03
RELEVANCE
AUTHOR
mark_k