Claude Opus 5 Feels Smarter, Harder to Trust
A developer argues that Opus 5 is more capable on benchmarks yet less pleasant in real coding work because it makes bold assumptions, rewrites plans, and asks fewer clarifying questions. The gap highlights how benchmark optimization can conflict with the judgment and restraint developers want from coding agents.
Opus 5’s problem may be less intelligence than calibration: it is optimized to complete underspecified tasks, while real software work often requires recognizing ambiguity and waiting for direction.
- –Strong benchmark performance rewards confident assumptions, even when production requirements are incomplete or politically constrained
- –Coding agents need to preserve user intent, surface tradeoffs, and ask questions—not merely maximize the chance of a technically valid answer
- –Extra babysitting shifts the cost of model errors onto developers through review, rollback, and prompt-harness maintenance
- –The critique suggests evaluating models on interruption quality, assumption tracking, and plan fidelity alongside coding benchmarks
- –Opus 5 may still be valuable for well-specified, long-running tasks, but less suitable as an autonomous collaborator when requirements evolve interactively
DISCOVERED
2h ago
2026-08-14
PUBLISHED
4h ago
2026-08-14
RELEVANCE
AUTHOR
numeri