Qwen3.8-27B tops Opus 4.6 Max in computer use
New benchmark figures indicate that the 27-billion-parameter open-weight model Qwen3.8-27B substantially outperforms proprietary frontier model Claude Opus 4.6 Max at computer and device control tasks. Evaluated on OSWorld-Verified, Qwen3.8-27B scored 84.3 compared to Opus 4.6 Max's 72.7, while on AndroidWorld it achieved 81.9 versus 62.0. These wide margins demonstrate that targeted open-weight architectures are becoming formidable contenders against the largest closed models in practical GUI grounding, screen comprehension, and automated desktop and mobile operating system navigation.
Smaller open-weight models dominating computer-use benchmarks will rapidly accelerate the shift toward private, locally deployed autonomous agents. Leading by 11.6 points on OSWorld-Verified and nearly 20 points on AndroidWorld proves open-weight models can excel at complex, multi-step interface workflows. A 27B parameter model can also be served locally on accessible consumer or workstation hardware, eliminating latency and privacy concerns tied to streaming screen captures to closed APIs. Furthermore, superior performance in agentic execution by open weights reduces reliance on costly closed-source subscriptions for OS automation.
DISCOVERED
1h ago
2026-09-12
PUBLISHED
1h ago
2026-09-12
RELEVANCE
AUTHOR
Oluwaphilemon1