YOU ARE VIEWING ONE ITEM FROM THE AICRIER FEED

Harness-of-Harness lifts coding agents 52%

AICrier tracks AI developer news across Product Hunt, GitHub, Hacker News, YouTube, X, arXiv, and more. This page keeps the article you opened front and center while giving you a path into the live feed.

// WHAT AICRIER DOES

7+

TRACKED FEEDS

24/7

SCRAPED FEED

Short summaries, external links, screenshots, relevance scoring, tags, and featured picks for AI builders.

Harness-of-Harness lifts coding agents 52%
OPEN LINK ↗
// 1h agoRESEARCH PAPER

Harness-of-Harness lifts coding agents 52%

Harness-of-Harness wraps existing coding-agent harnesses in persistent planning, implementation, and independent QA loops, improving results across GameCraft-Bench, FrontierSWE, and ProgramBench. The Shanghai Artificial Intelligence Laboratory paper also reports a 70-plus-iteration run that produced Fusepoint, a playable FPS from a PRD and empty workspace.

// ANALYSIS

The important idea isn’t simply running an agent longer; it’s creating an evidence-carrying control loop that preserves verified progress between sessions. The results are promising, but the headline gains remain research claims rather than proof of production-ready autonomy.

  • Reports a 52.25% average relative gain and up to 82.86% after three iterations across three benchmarks
  • Separates planning, development, and independent testing to reduce regressions and self-evaluation bias
  • Demonstrates long-horizon capability by building a playable FPS across more than 70 autonomous loops
  • Keeps versioned artifacts, issue histories, and evidence packets so each iteration can build on validated work
  • HoH-lite and reproducibility materials are planned, which should make the claims easier to independently verify
// TAGS
harness-of-harnessagentcoding-agentai-codingevaluationbenchmarkframework

DISCOVERED

1h ago

2026-09-04

PUBLISHED

1h ago

2026-09-04

RELEVANCE

10/ 10

AUTHOR

Discover AI