GPT-6 Astra leads the BrowseComp benchmark for agentic web research, though Kimi K3 offers near-identical performance at a fraction of the cost.
The latest BrowseComp leaderboard reveals GPT-6 Astra at the top with a 91.5% score, closely trailed by Kimi K3 at 91.2% and Claude Opus 5 at 90.8%. However, Kimi K3 costs significantly less at $14.25 per million output tokens compared to Astra's $50 per million. The post highlights the viability of using routing tools like Merge Gateway to dynamically allocate tasks between models based on their performance and cost efficiency.
The tightening gap in flagship model capabilities makes cost the primary differentiator for scaled applications.
- –GPT-6 Astra holds the performance edge but demands a premium price.
- –Kimi K3 provides an exceptional price-to-performance ratio for web research tasks.
- –Claude Opus 5 sits competitively in the middle ground of both cost and capability.
- –Gateway routing solutions are increasingly necessary to optimize unit economics in LLM applications.
DISCOVERED
1d ago
2026-09-22
PUBLISHED
1d ago
2026-09-22
RELEVANCE
AUTHOR
merge_api