Harness-of-Harness lifts coding agents 52%
Harness-of-Harness wraps existing coding-agent harnesses in persistent planning, implementation, and independent QA loops, improving results across GameCraft-Bench, FrontierSWE, and ProgramBench. The Shanghai Artificial Intelligence Laboratory paper also reports a 70-plus-iteration run that produced Fusepoint, a playable FPS from a PRD and empty workspace.
The important idea isn’t simply running an agent longer; it’s creating an evidence-carrying control loop that preserves verified progress between sessions. The results are promising, but the headline gains remain research claims rather than proof of production-ready autonomy.
- –Reports a 52.25% average relative gain and up to 82.86% after three iterations across three benchmarks
- –Separates planning, development, and independent testing to reduce regressions and self-evaluation bias
- –Demonstrates long-horizon capability by building a playable FPS across more than 70 autonomous loops
- –Keeps versioned artifacts, issue histories, and evidence packets so each iteration can build on validated work
- –HoH-lite and reproducibility materials are planned, which should make the claims easier to independently verify
DISCOVERED
1h ago
2026-09-04
PUBLISHED
1h ago
2026-09-04
RELEVANCE
AUTHOR
Discover AI