Claude Opus 5’s Verbosity Outruns Standard Benchmarks
An analysis of tens of thousands of Arena outputs reportedly finds Claude Opus 5 uses roughly three times more em dashes than earlier Opus models. Anthropic’s own prompting guidance acknowledges that Opus 5 produces longer default user-facing responses.
The em-dash count is a funny but revealing proxy for a real developer complaint: capability gains can arrive alongside worse output discipline.
- –Punctuation frequency measures style, not intelligence, and varies with prompts, system instructions, and task mix.
- –Anthropic confirms Opus 5 tends to produce longer visible responses than prior Opus models.
- –Developers should add verbosity, formatting, and instruction-adherence checks to model regression suites.
- –Explicit concision prompts help, but Anthropic says effort controls reasoning more reliably than final response length.
- –Better evaluations pair task success with token use, latency, verbosity, and user preference.
DISCOVERED
2h ago
2026-08-13
PUBLISHED
2h ago
2026-08-13
RELEVANCE
AUTHOR
MartinSzerment