Merge Gateway finds GLM 5.3 beats Claude
Merge Gateway evaluated five open-weight models—GLM 5.3, GLM 5.3 Flash, DeepSeek V4 Pro, DeepSeek V4 Flash, and Kimi K3—against Anthropic’s Claude Sonnet 5 across 20 real coding challenges, with GLM 5.3 solving 14 tasks versus Claude’s 12 at $0.317 per successful solve versus $3.415 and a faster median completion time (170s versus 197s). DeepSeek V4 Flash matched Claude’s 12/20 solve rate at $0.032 per successful task—106x cheaper but slower at 261 seconds—and all tested models are accessible through Merge Gateway’s unified API.
Open-weight models have caught up with frontier proprietary models on coding agent tasks, radically undermining closed-source pricing power.
• The economic moat of proprietary frontier models is breaking: DeepSeek V4 Flash delivering the exact same task completion rate as Claude Sonnet 5 at 106x lower cost makes premium API pricing hard to sustain for high-throughput coding agents.
• GLM 5.3 establishes a new benchmark standard: By outperforming Claude Sonnet 5 on task resolution (14/20 vs. 12/20), cost ($0.317 vs. $3.415), and speed (170s vs. 197s), GLM 5.3 proves that open-weight options can lead across the entire Pareto frontier.
• Speed remains the compromise for ultra-cheap inference: DeepSeek V4 Flash was the slowest model tested at 261s per solve, whereas GLM 5.3 Flash led raw execution speed at 111s despite a lower task completion rate (9/20).
• Dynamic routing becomes indispensable: With cost and latency variances spanning orders of magnitude across capable models, unified LLM gateways are essential for optimizing agent workflows in production.
DISCOVERED
1h ago
2026-09-15
PUBLISHED
1h ago
2026-09-15
RELEVANCE
AUTHOR
merge_api