SRE-Bench exposes AI’s binary-analysis gap
Columbia researchers introduce SRE-Bench, a contamination-free benchmark testing AI agents on real-scale binary reverse engineering without source code. Its 262 protected binaries and 1,572 deterministic tasks show frontier agents still struggle to recover program semantics reliably.
SRE-Bench is a much-needed reality check for cybersecurity agents: strong source-code performance does not automatically transfer to compiled software.
- –19 private programs and 44 anti-analysis mechanisms reduce memorization and shortcutting
- –The strongest initial model fully solved only 31.5% of instances
- –Tasks span malware, firmware, proprietary software, and other realistic reverse-engineering scenarios
- –Agents’ relative insensitivity to optimization and static linking suggests pattern matching remains a major weakness
- –Deterministic grading makes the benchmark more meaningful than evaluations based on plausible explanations
DISCOVERED
1h ago
2026-10-06
PUBLISHED
1h ago
2026-10-06
RELEVANCE
AUTHOR
AI Search
