Anthropic published its second Risk Report on 14 August, under version 3.4 of its Responsible Scaling Policy. It covers the period from the previous report on 24 February to a coverage date of 15 July 2026.
The rating that moved
For Autonomy threat model 1 — misalignment in high-stakes settings — Anthropic now rates overall risk Low, explicitly an increase from very low. A frontier lab revising its own risk assessment upward is rare enough to be the story on its own.
Why the coverage gets the cause wrong
Most write-ups tie the increase to Model 2, the strongest of three unreleased internal models the report discloses. The report does not. It attributes the rise to general increased uncertainty around recent incident disclosures relating to model behaviour in cybersecurity evaluations, and handles Model 2 in a separate section. "Anthropic shelves a model over safety" overstates it too: the stated reasons are incomplete predeployment assessment and no current release plan, not a safety verdict.
What Model 2 is
Anthropic describes it as somewhat more capable than frontier model Claude Mythos 5 and a noticeable improvement on many internal tasks — but not a jump of the size seen from Claude Opus 4.6 to Mythos Preview. The three undisclosed systems are Claude Opus 5, Model 1 and Model 2.
The admission buried in the methodology
Anthropic keeps automated AI R&D risk at Low but says it is less confident than before, because its most concrete task-based evaluations have saturated — they no longer capture increases in capability. It reports internal research moving significantly faster with AI assistance, though not yet at twice the rate. A lab losing the instrument it uses to measure its own acceleration is a more durable problem than any single model.
The classifier gap
Separately, the report rates non-novel chemical and biological weapons risk low, but higher than our previous estimate, after discovering that all human-feedback vendor traffic had run without blocking biological classifiers. Anthropic says the gap is remediated, with no evidence of misuse and no customer impact. Note the timing throughout: the coverage date is 15 July, so the whole assessment describes a position a month before it was published.
