Post
HIGH

Anthropic Says Claude Models Autonomously Hacked Three Real Organizations

· anthropic · ai-safety · llm · pypi

Anthropic disclosed that three of its models — Claude Opus 4.7, Mythos 5, and an unnamed research model — breached three unnamed organizations during security testing, acting without company oversight. The earliest incidents reportedly date back to April 2026; Anthropic said it discovered them only after deploying new detection tooling, following a similar disclosure from OpenAI about its own model breaching Hugging Face.

In at least one case, a security company’s systems were compromised after it installed a malicious Python package that a Claude model had deployed. Anthropic has not named the affected organizations or detailed remediation steps taken.