Claude Opus 5 tops vulnerability recall benchmark
Aikido Security benchmarked five frontier AI models across 32 post-cutoff CVEs to evaluate automated vulnerability discovery. Claude Opus 5 achieved the highest recall by finding 26 CVEs through exhaustive exploration, though Sol delivered higher overall precision and F1 score with fewer false positives.
Deep investigation persistence in security agents maximizes vulnerability discovery but introduces a significant triage overhead that must be balanced with precision filtering.
- –**Unmatched Persistence**: Opus 5 utilized at least 27 of 30 actions in 24.5% of runs (compared to <2.3% for other models), spending extensive budget exploring adjacent attack vectors like CSRF and target flaw verification.
- –**Top Recall & Tail Bug Discovery**: Reaching 81.3% pass@3 recall (26/32 CVEs), Opus 5 was the only model capable of discovering three specific CVEs in the benchmark suite.
- –**Precision & Noise Penalty**: Opus 5 generated 121 rejected candidate findings versus Sol's 53, resulting in lower precision and giving Sol the higher overall F1 score.
- –**Optimization Vector**: Future agent tuning should focus on teaching models to self-assess when prolonged digging yields diminishing returns rather than forcing artificial early termination.
DISCOVERED
1h ago
2026-07-29
PUBLISHED
1h ago
2026-07-29
RELEVANCE
AUTHOR
AikidoSecurity