SWE-Bench ProMax Raises Refactoring Bar
SWE-Bench ProMax introduces 170 expert-curated, multilingual refactoring tasks across seven programming languages, averaging 11.4 modified files per task. The top evaluated agent resolves just 41.2%, exposing a tougher frontier for real-world coding automation.
SWE-Bench ProMax is a welcome corrective to benchmarks that reward narrow bug fixes or suffer from flawed tests—but 170 tasks still makes leaderboard conclusions fragile.
- –Covers Python, Java, TypeScript, Go, C, C++, and Rust, reducing Python-centric evaluation bias
- –Targets behavior-preserving, cross-file refactors averaging 261.6 changed lines, closer to professional maintenance work
- –Manually reviewed specifications and tests aim to improve evaluation reliability
- –A 41.2% best score suggests coding agents remain far from dependable autonomous refactoring
- –Developers should treat results as evidence about agent workflows and verification, not just model intelligence
DISCOVERED
2h ago
2026-08-12
PUBLISHED
1d ago
2026-08-11
RELEVANCE
AUTHOR
_akhaliq