Terminal Agents Survey Reframes Harness Benchmarks
This 52-page survey defines terminal agents as systems whose progress depends on command execution, textual feedback, and stateful environments. It argues that model, interface, harness, runtime, and environment jointly determine results, explaining why benchmark comparisons often conflict.
The paper’s sharpest point is that “agent performance” is a property of the entire execution system, not just the underlying model.
- –Proposes a seven-dimensional profile covering action, feedback interpretation, state tracking, verification, recovery, and side-effect control
- –Shows that benchmark families expose different process signals, making leaderboard rankings difficult to generalize
- –Calls for reporting runtime conditions, harness details, and replayable traces alongside final success rates
- –Highlights how weak recovery and governance metrics obscure failures in mutable, real-world environments
- –Gives developers a stronger framework for evaluating terminal agents beyond pass/fail outcomes
DISCOVERED
2h ago
2026-08-24
PUBLISHED
4h ago
2026-08-24
RELEVANCE
AUTHOR
omarsar0
