Kimi K3 scores second on AA-Briefcase benchmark
Artificial Analysis evaluated Moonshot AI's 2.8-trillion parameter Kimi K3 on the AA-Briefcase benchmark, where it placed second overall with a 1543 Elo score. While showing strong analytical performance, its high compute demands drove average task costs to $10.57 and execution times to 56.4 minutes.
Kimi K3 demonstrates that massive turn depth can achieve near-frontier agentic reasoning, but its extreme latency and cost profiles pose major hurdles for commercial deployment.
- –Frontier reasoning capability: A +727 Elo jump over Kimi K2.6 puts Kimi K3 on par with top-tier Western frontier models on complex analytical tasks.
- –High cost and speed bottlenecks: Averaging nearly an hour (56.4 mins) and $10.57 per task makes Kimi K3 ~2.5x slower than Claude Fable 5 and expensive for regular workflows.
- –Presentation quality gap: Lower presentation Elo (1471) compared to GPT-5.6 Sol shows that while raw data processing is strong, output design and formatting need polish.
DISCOVERED
2h ago
2026-07-22
PUBLISHED
6h ago
2026-07-22
RELEVANCE
AUTHOR
wertyk