
Claude Opus 5 Tops SlopCodeBench Eval
Humanlayer evaluated Claude Opus 5, Sonnet 5, and Opus 4.8 on a 17-checkpoint subset of SlopCodeBench, finding that Opus 5 led with a 24% strict pass rate compared to 6% for predecessor models. However, no model completed a multi-checkpoint problem without defects, and Opus 5 exhibited significant code inflation by writing five times as many functions as Opus 4.8.
Current AI models remain incapable of running "lights-off" on complex software projects without active human steering and context management.
- –**Evolving Requirement Signal:** SlopCodeBench addresses the key flaw of traditional single-shot coding benchmarks by testing how agents modify and maintain codebases across sequential checkpoints.
- –**Opus 5 Marginally Leads:** Opus 5 achieved top performance among tested models (24% strict pass rate), but still failed to finish a single full problem without introducing bugs or breaking prior functionality.
- –**Code Inflation and Slop:** High correctness in Opus 5 coincided with extreme verbosity (writing 29,000+ lines of code across the subset), proving that models still struggle with architectural clean-code discipline over time.
DISCOVERED
2h ago
2026-07-28
PUBLISHED
3h ago
2026-07-27
RELEVANCE
AUTHOR
dhorthy