Passing Test Study Exposes Detector Gaps
Zhuowen Liu’s study evaluates 15 prompt-injection detectors across AgentDojo, tau-bench, and BIPIA, finding that public benchmark rankings poorly predict agent deployment performance. The strongest BIPIA detector caught just 2% of AgentDojo injections at a 1% false-positive rate. ([arXiv](https://arxiv.org/abs/2610.03448))
This paper shows that prompt-injection detection is suffering from benchmark overfitting: detectors often recognize familiar input formats rather than robustly identify attacks.
- –Detector rankings transfer poorly across benchmarks and real agent tool outputs.
- –False-positive rates range from zero to over 90%, making usability as important as recall.
- –Agent-style training data appears more predictive than exposure to benchmark attack strings.
- –Developers should evaluate detectors on their own tool-call outputs at a fixed low false-positive rate.
- –Detection should remain one defense layer alongside tool permissions, output isolation, and constrained agent actions.
DISCOVERED
56m ago
2026-10-06
PUBLISHED
1h ago
2026-10-06
RELEVANCE
AUTHOR
ktwu01