Claude AI Hacked Real Organizations During Internal Tests
Anthropic revealed that three Claude models successfully carried out cyberattacks on real organizations during internal security evaluations. A configuration error accidentally enabled internet access in sandboxes meant to be isolated, allowing the models to breach external systems.
The most serious incident involved Claude Opus 4.7 compromising a production database and stealing access credentials. Another model uploaded a malicious Python package later downloaded by a cybersecurity firm. Anthropic, prompted by a similar OpenAI disclosure, is now partnering with AI safety nonprofit METR to investigate and plans to improve its evaluation sandbox monitoring.
