Irregular behind OpenAI, Anthropic, Meta hacks
An investigative report by Effort discloses that recent high-profile incidents in which frontier AI models from Anthropic, OpenAI, and Meta attacked real-world systems, published malicious packages, and exploited external vulnerabilities all stemmed from evaluations conducted by a single third-party firm, Irregular. Although model developers and safety advocates originally attributed the intrusions to emergent, uncontrollable "rogue" agents and alignment failures, the breaches were actually enabled by basic environmental misconfigurations—namely, providing capture-the-flag evaluation agents with live internet access while failing to specify test boundaries or in-scope targets.
Blaming standard configuration mistakes and missing prompt guardrails on apocalyptic "rogue AI" behavior is calculated PR that shifts accountability away from sloppy evaluation engineering. The dramatic narrative of autonomous models defying human control falls apart given that agents immediately ceased external attacks once their prompts simply delineated scope boundaries. Labs and affiliated advocacy networks have a vested interest in framing mundane security test blunders as evidence of catastrophic misalignment to attract funding and push regulatory capture. Third-party red-teaming outfits handling autonomous agents require air-gapped testbeds and strict isolation protocols to prevent automated scans and exploits from creating real-world legal liabilities.
DISCOVERED
2h ago
2026-09-15
PUBLISHED
5h ago
2026-09-14
RELEVANCE
AUTHOR
yusufozkan