OpenAI's Astra Model Flagged for Critical Hacking Risk
OpenAI has disclosed that its unreleased Astra model may qualify for a Critical cybersecurity risk designation under its Preparedness Framework, making it the first of its LLMs to potentially reach that threshold. The rating applies to models capable of finding zero-day exploits in hardened real-world systems or launching cyberattacks based solely on high-level user instructions.
In response, OpenAI is restricting Astra to sandboxed environments with limited network access, pausing unsandboxed development, strengthening encryption on the model's weights, and deploying monitoring tools to detect malicious activity in AI agents powered by Astra. The company also plans to notify government agencies and AI safety organizations.
