Anthropic said in its second company-wide risk report, published on Aug. 14, that it raised its rating for “catastrophic harm caused by misalignment in high-risk scenarios” from “very low” to “low” under Responsible Scaling Policy v3.4. The report also disclosed an unreleased internal system called Model 2 for the first time.
The rating change was not tied to a failed safety test
Anthropic said the revision did not happen because a model failed a safety evaluation. According to the company, the main reason was that recent cybersecurity evaluation disclosures increased overall uncertainty. It also pointed to a key measurement problem: the internal benchmark used to determine whether a model has crossed the most dangerous capability threshold has become saturated, with scores effectively maxing out and no longer capturing incremental capability gains.
In practical terms, Anthropic said the issue is less about a specific model malfunction and more about the fact that its current measurement tools can no longer cleanly distinguish higher capability levels. That pushed the company toward a more conservative risk rating.
Model 2 is slightly stronger than Mythos 5 and remains internal
The report said Anthropic has an internal system called Model 2 that is “slightly more capable” than its frontier model Mythos 5. The company said the model is already used extensively inside Anthropic, but it has not yet completed the full set of evaluations normally required before launch. There are also no current plans to release it publicly.
For a company moving toward an IPO, the combination of acknowledging a stronger internal model, saying its standard safety review is not yet complete, and voluntarily raising a related risk rating puts fresh scrutiny on the gap between frontier capability and safety evaluation.

