HarnessSecurity-Bench exposes coding-agent security gaps
HarnessSecurity-Bench evaluates 10 security mechanisms across coding-agent harnesses and finds that auto-approval can raise attack success from 29.2% to 95.6%. Its 2,500-trial benchmark shows that stronger restrictions often trade meaningful utility for safety. [Paper](https://arxiv.org/abs/2610.07639)
The paper makes a compelling case that the harness—not just the model—is the real security boundary for coding agents.
- –Auto-approve delivers convenience by dramatically increasing exposure to malicious instructions
- –Network isolation and read-only mode reduce attack impact but can cripple legitimate workflows
- –Command allowlisting appears to offer a more practical security–utility balance
- –Alternative execution paths can bypass narrow tool or command restrictions
- –Developers should evaluate attack effects, task success, and execution cost together
DISCOVERED
1h ago
2026-10-07
PUBLISHED
2h ago
2026-10-07
RELEVANCE
AUTHOR
dani_avila7