Specific Labs launches Real-SWE enterprise coding benchmark
Specific Labs has released Real-SWE, a software engineering benchmark evaluating frontier AI coding agents on private enterprise production codebases featuring complex business rules, multi-service dependencies, and infrastructure emulators. Across 640 scored rollouts, the benchmark reveals enterprise software engineering remains a steep challenge for AI agents, with Anthropic's Fable 5.1 and Claude Code achieving the top pass@1 resolution rate at just 38.8%, followed by OpenAI's GPT-6 Astra at 33.8% and Google's Gemini 3.8 Flash at 31.2%.
Saturated public benchmarks have fueled an illusion of agentic engineering maturity, but Real-SWE shows that frontier models quickly unravel when forced to navigate proprietary architectures without public training data. The sub-40% ceiling across all leading models proves that enterprise development hinges on deciphering bespoke business rules, side effects, and multi-service interactions rather than isolated bug patching. Benchmarking native harness pairings (such as Claude Code, Codex CLI, and Gemini CLI) provides a far more representative assessment of agent utility in authentic developer environments, while sourcing evaluation environments directly from private enterprise codebases introduces a durable defense against dataset contamination and leakage.
DISCOVERED
1h ago
2026-09-13
PUBLISHED
1h ago
2026-09-13
RELEVANCE
AUTHOR
AI Search