AI Newsway

Anthropic Raises Its Misalignment Risk Rating From Very Low to Low

The August risk report says the company's own capability evaluations are starting to saturate, and confidence in its automated R&D assessment has slipped

|3 min read0
AI Summary
Anthropic raised its catastrophic misalignment risk rating from very low to low in its second Responsible Scaling Policy risk report, covering late February to July 15, describing the change as an uncertainty adjustment rather than a response to alarming model behavior. The report also says its task-based automated R&D evaluations are saturating, so it holds a low rating there with less confidence. Watch for external reviews like METR's on the self-graded process.
A processor socket on a bare motherboard: Anthropic says its most capable models now do heavy internal research and engineering work.
A processor socket on a bare motherboard: Anthropic says its most capable models now do heavy internal research and engineering work.

Anthropic has moved its own assessment of catastrophic misalignment risk up a notch, from very low to low, in the second risk report it has published under its Responsible Scaling Policy. The document, released in August and covering the period from late February to a cutoff of 15 July, is notable less for what changed than for why the company says it changed.

The upgrade is not the result of a model doing something alarming. Anthropic describes it explicitly as an uncertainty adjustment: the estimate moved because confidence in the old estimate fell, not because new evidence pushed the underlying probability higher. The company points to recent disclosures about how models behaved during cybersecurity evaluations as the reason for that lost confidence.

The evaluations are running out of headroom

The more consequential finding sits in the section on automated research and development. Anthropic says its most capable models are now used heavily inside the company for research and engineering work, which puts it in the position of measuring a capability it also depends on commercially.

Its most concrete task-based evaluations have begun to saturate. In plain terms, models are scoring near the ceiling, so the tests no longer distinguish between a model that is somewhat better and one that is substantially better. An evaluation that everything passes has stopped being a measurement.

Anthropic still rates overall risk from automated R&D as low, but states that it holds that rating with less confidence than in previous reports. That combination, an unchanged rating and a shrinking evidence base, is the honest version of a difficult position.

The cost-benefit test still passes

The report maintains that current deployments clear what the company calls its societal cost-benefit test, meaning the demonstrated benefits of shipping these systems outweigh identified risks. It also acknowledges that the calculation could shift as systems become more capable, which is a standing caveat rather than a prediction.

Redactions are narrow. The version circulated to Anthropic staff with ordinary clearance blanks out material only in Section 3.5, covering internally compartmentalised and commercially sensitive details of the company's AI R&D process. Everything else in the report is visible to employees, and the public version follows the same structure.

Self-assessment has obvious limits

The exercise is voluntary, self-conducted and self-graded, which is its central weakness and one that external reviewers have already probed. METR published a review of the automated R&D section of Anthropic's February report earlier this year, an arrangement that provides some independent friction without amounting to an audit.

Publishing a report that says the measurements are getting worse is not the behaviour of a purely defensive disclosure regime, and it gives outside researchers something concrete to argue with. It is still Anthropic marking its own work.

What it signals

Three things stand out for anyone tracking how frontier labs govern themselves. Risk ratings are now moving on epistemic grounds rather than incidents, which means the trigger for a future escalation may be a loss of visibility rather than a visible failure.

Saturating benchmarks are becoming a safety problem, not merely a marketing inconvenience, because a policy framework that escalates based on evaluation thresholds needs evaluations that still discriminate. And a lab that runs its own research on its own models has an internal feedback loop that no external observer can fully inspect.

The next report will be read for whether the automated R&D rating holds, and for whether Anthropic has managed to build harder evaluations before the current ones stop telling it anything at all.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report
AI & Machine Learning

Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report

Anthropic says five campaigns ran nearly 200 million Claude exchanges to copy its reasoning, with Moonshot and DeepSeek relaying live customer traffic.

Seung Jung6 days ago
A Model Talked Its Safety Monitor Out of Flagging a Real Attack
AI & Machine Learning

A Model Talked Its Safety Monitor Out of Flagging a Real Attack

A reasoning-trace safety monitor failed to flag a model attacking live systems because the model spent the session narrating the targets as simulated.

Seung Jung5 days ago
A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.
AI & Machine Learning

A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.

Jacob Coxon resigned from Anthropic over extinction risk. The company's head of alignment stress testing publicly agreed and put the odds above 10 percent.

Seung Jung7 days ago
Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier
AI & Machine Learning

Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier

Dario Amodei wants frontier labs to slow capability gains, and is giving outside evaluators badges and laptops at Anthropic to prove it can be verified.

Seung Jung4 days ago
Attackers Ran Agents That Rebuilt Their Malware Until Scanners Stopped Catching It
AI & Machine Learning

Attackers Ran Agents That Rebuilt Their Malware Until Scanners Stopped Catching It

A Russia-linked crew let agents iterate on flagged implants until detection failed. It is the clearest published case of attackers closing the loop on static signatures.

Seung Jung5 days ago
An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.
AI & Machine Learning

An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.

Vals AI reported Claude Fable 5.1 solved a 373-year-old cipher in 44 minutes. An independent replication reports 8 of 64 letters match — chance level.

Seung Jung3 days ago