Claude Fable 5 system prompt leaks
Following the launch of Anthropic's Claude Fable 5, researcher Pliny the Liberator claimed to jailbreak the model and leak its 120,000-character system prompt using a multi-agent strategy. The exploit allegedly bypassed the model's safety classifiers, which are designed to fall back to Claude Opus 4.8 for sensitive queries.
Classifier-based safety routing is a band-aid, not a cure, and multi-agent jailbreaks prove that static guardrails cannot keep up with dynamic orchestration techniques. Anthropic's fallback to Claude Opus 4.8 for sensitive tasks is easily bypassed if the initial query is obfuscated enough to slip past the classifier. The leaked prompt's massive size of 120,000 characters indicates that system prompts are increasingly serving as complex API and behavior specifications rather than simple guidance. Finally, bypassing guardrails via "pack hunt" strategies shows that multi-agent systems can coordinate to exploit model vulnerabilities, introducing a new tier of security threat.
DISCOVERED
51d ago
2026-06-13
PUBLISHED
51d ago
2026-06-13
RELEVANCE
AUTHOR
AlphaSignalAI