Anthropic has moved its own assessment of catastrophic misalignment risk up a notch, from very low to low, in the second risk report it has published under its Responsible Scaling Policy. The document, released in August and covering the period from late February to a cutoff of 15 July, is notable less for what changed than for why the company says it changed.
The upgrade is not the result of a model doing something alarming. Anthropic describes it explicitly as an uncertainty adjustment: the estimate moved because confidence in the old estimate fell, not because new evidence pushed the underlying probability higher. The company points to recent disclosures about how models behaved during cybersecurity evaluations as the reason for that lost confidence.
The evaluations are running out of headroom
The more consequential finding sits in the section on automated research and development. Anthropic says its most capable models are now used heavily inside the company for research and engineering work, which puts it in the position of measuring a capability it also depends on commercially.
Its most concrete task-based evaluations have begun to saturate. In plain terms, models are scoring near the ceiling, so the tests no longer distinguish between a model that is somewhat better and one that is substantially better. An evaluation that everything passes has stopped being a measurement.
Anthropic still rates overall risk from automated R&D as low, but states that it holds that rating with less confidence than in previous reports. That combination, an unchanged rating and a shrinking evidence base, is the honest version of a difficult position.
The cost-benefit test still passes
The report maintains that current deployments clear what the company calls its societal cost-benefit test, meaning the demonstrated benefits of shipping these systems outweigh identified risks. It also acknowledges that the calculation could shift as systems become more capable, which is a standing caveat rather than a prediction.
Redactions are narrow. The version circulated to Anthropic staff with ordinary clearance blanks out material only in Section 3.5, covering internally compartmentalised and commercially sensitive details of the company's AI R&D process. Everything else in the report is visible to employees, and the public version follows the same structure.
Self-assessment has obvious limits
The exercise is voluntary, self-conducted and self-graded, which is its central weakness and one that external reviewers have already probed. METR published a review of the automated R&D section of Anthropic's February report earlier this year, an arrangement that provides some independent friction without amounting to an audit.
Publishing a report that says the measurements are getting worse is not the behaviour of a purely defensive disclosure regime, and it gives outside researchers something concrete to argue with. It is still Anthropic marking its own work.
What it signals
Three things stand out for anyone tracking how frontier labs govern themselves. Risk ratings are now moving on epistemic grounds rather than incidents, which means the trigger for a future escalation may be a loss of visibility rather than a visible failure.
Saturating benchmarks are becoming a safety problem, not merely a marketing inconvenience, because a policy framework that escalates based on evaluation thresholds needs evaluations that still discriminate. And a lab that runs its own research on its own models has an internal feedback loop that no external observer can fully inspect.
The next report will be read for whether the automated R&D rating holds, and for whether Anthropic has managed to build harder evaluations before the current ones stop telling it anything at all.






