Ox Alpha scores 80% on DeepSWE
Ox Alpha reportedly scored 80% on a 10-task DeepSWE sample while offering a one-million-token context window and free preview access through OpenRouter. The result is promising but too small to establish a reliable benchmark lead. AI Primer (https://www.ai-primer.com/engineer/stories/ox-alpha-openrouter-release)
Ox Alpha looks like an unusually compelling coding-model experiment, but its anonymity and limited evaluation data make the hype outrun the evidence.
- –The 80% result means eight of ten tasks passed, leaving substantial statistical uncertainty.
- –Its 1M-token context, multimodal inputs, tool support, and free preview are highly attractive for agentic coding workflows.
- –OpenRouter usage indicates immediate interest from coding agents, including Claude Code and Hermes Agent.
- –Provider identity, architecture, official benchmarks, and post-preview pricing remain undisclosed.
- –Privacy claims require scrutiny: OpenRouter’s general ZDR policy differs from the model listing’s reported provider-retention terms. OpenRouter ZDR documentation (https://openrouter.ai/docs/guides/features/zdr)
DISCOVERED
1d ago
2026-08-22
PUBLISHED
1d ago
2026-08-22
RELEVANCE
AUTHOR
WorldofAI