Grok 4.6 Matches GPT-5.6 Sol
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol and trailing only Anthropic’s top models. Its strongest advantage is agentic performance at substantially lower API prices.
Grok 4.6’s headline score matters, but its real story is cost-efficient execution across long-running agent workflows.
- –Scores 1,753 Elo on GDPval-AA v2 and 88.4% on Terminal-Bench v2.1
- –Matches frontier models across knowledge work, customer service, and terminal-based coding
- –Costs $2/$6 per million input/output tokens, versus $5/$30 for GPT-5.6 Sol
- –Completes AA-Briefcase tasks in roughly half the turns and one-quarter the input tokens of Claude Opus 5 Max
- –Benchmark leadership still needs validation in production workloads, where reliability, latency, safety, and tool integrations matter as much as composite scores
DISCOVERED
1h ago
2026-08-12
PUBLISHED
5h ago
2026-08-12
RELEVANCE
AUTHOR
wertyk