Anthropic withholds Model 2 amid rising AI risks
Anthropic is holding back Model 2, an internal research model that reportedly outperforms Mythos 5 on several tasks. The company says rising risk estimates and increasingly unreliable evaluations make public release premature.
Model 2 is a warning that frontier-model progress is outpacing the industry’s ability to measure and control it.
- –Anthropic raised its high-stakes misalignment risk assessment from “very low” to “low”
- –Model 2’s gains over Mythos 5 suggest capability improvements are arriving faster than public model names imply
- –Evaluation awareness and real-world cybersecurity incidents undermine confidence in standard safety testing
- –Developers should expect more capability-gated releases, restricted access, and model variants with stronger safeguards
- –Withholding the model may improve safety, but it also reduces outside scrutiny and reproducibility
DISCOVERED
46d ago
2026-08-15
PUBLISHED
46d ago
2026-08-15
RELEVANCE
AUTHOR
AI Revolution