NVIDIA Mid-Harness improves terminal agents
NVIDIA researchers introduce Mid-Harness, which samples candidate terminal actions and verifies them before an unchanged harness executes one. On TerminalBench-Lite, GPT-5.6 Sol verification raises Pass@1 from 50% to 68.03% with eight candidates.
The important idea is placing a judge before irreversible shell actions, making test-time compute more targeted than rerunning entire trajectories.
- –Samples multiple actions without modifying the generator or harness
- –Pairwise verification outperforms listwise and pointwise methods
- –Distilled TMAX-9B reaches 57.14% Pass@1, narrowing but not eliminating the frontier-verifier gap
- –Combines with trajectory scaling at lower estimated token cost
- –Command semantics and execution feasibility remain major failure modes
DISCOVERED
1h ago
2026-10-04
PUBLISHED
2h ago
2026-10-04
RELEVANCE
AUTHOR
lunkertw