Marathoner Makes 9B Agents Run 10+ Hours
Marathoner is a post-training pipeline built on Qwen3.5-9B for ultra-long-horizon software work. Trained on tasks synthesized from 100,000 major-release PRs across 10,000 GitHub repositories, it reportedly sustains 10+ hour runs and 1,000+ tool calls while improving across five coding benchmarks. [arXiv](https://arxiv.org/abs/2609.34378)
The breakthrough is less a magical 9B model than a training regime designed around endurance, verification, and recovery during long coding sessions. The catch: Marathoner remains a research result without released weights or code.
- –GitHub PRs become realistic repository-level tasks, with unit tests serving as executable verifiers.
- –Multi-Task Chaining combines five tasks into frontier-scale workloads with execution limits up to 40 hours.
- –Marathoner-9B scores 77.5 on SWE-bench Verified versus 43.8 for its Qwen3.5-9B base, and beats Gemini-3.1-Pro on FrontierSWE and SWE-Marathon.
- –The headline endurance result needs context: FrontierSWE runs average 3.56 hours and 648 tool calls, while the 10+ hour, 1,000+ call figure reflects highly challenging executions.
- –No weights, training code, or dataset are linked, limiting immediate reproducibility for developers.
DISCOVERED
1h ago
2026-10-01
PUBLISHED
1h ago
2026-10-01
RELEVANCE
AUTHOR
mark_k