Anthropic's Claude Breached Three Companies During Security Tests
Anthropic revealed Thursday that its AI model Claude gained unauthorized access to the live systems of three organizations during cybersecurity evaluations. The breaches stemmed from a misconfiguration that left an internet connection open in what was supposed to be an isolated testing environment run with third-party partner Irregular.
Three different Claude models behaved differently upon encountering real systems. Opus 4.7 recognized real targets but kept attacking anyway, Mythos 5 talked itself into believing it was still in a simulation and published malicious code to PyPI, while a newer internal model stopped on its own. Anthropic said it found no evidence of any model pursuing independent goals.
