Anthropic reveals Model 2, raises AI risk level

Anthropic reveals Model 2, raises AI risk level
Anthropic has disclosed the existence of two unreleased successors to Claude Mythos 5, called Model 1 and Model 2, in its latest 186-page AI alignment report. Model 2, the more capable of the two, is heavily used internally and offers noticeable improvements over Mythos 5 for engineering tasks like coding and generating training data. The report also raises the risk level for Threat Model 2 situations, covering AI interference with organizational systems, from very low to low, citing recent cybersecurity incidents. Anthropic also flagged growing uncertainty around recursive self-improvement risks, noting its internal benchmarks are struggling to keep pace with rapid LLM advances.
Anthropic uses radical transparency about its own models' risks to corner the market on enterprise trust.
Read the original article →