GPT-6 Astra Hits 99.9% on ARC-AGI-3
ARC Prize reports GPT-6 Astra scoring 99.9% on ARC-AGI-3 Semi-Private with a Provider Adapter harness, versus 62.7% under the standard harness. Astra also used fewer actions than median human participants on 96% of levels, marking a major interactive-reasoning milestone without proving AGI.
This is both a model breakthrough and a harness breakthrough: persistent provider-managed reasoning state dramatically changes the result, complicating supposedly apples-to-apples benchmark comparisons.
- –The standard-harness score remains the cleaner cross-provider comparison; the 99.9% result depends on OpenAI-specific context preservation and compaction.
- –Astra reportedly builds compact symbolic world models, tracks mechanics, and plans actions with unusually high information density.
- –Provider Adapter runs were 3.66× faster and used 49% fewer total tokens across shared evaluations.
- –ARC-AGI-3 uses deterministic, closed-ended environments, so benchmark saturation demonstrates strong generalization within scope—not open-ended human-level intelligence.
- –The result makes context management, memory compression, and agent orchestration first-class parts of model evaluation.
DISCOVERED
1h ago
2026-09-03
PUBLISHED
1h ago
2026-09-03
RELEVANCE
AUTHOR
Wes Roth