AI Safety
Anthropic Raised Its Own Misalignment Risk Rating — and Not Because of Model 2
The second Risk Report moves high-stakes misalignment from 'very low' to 'low', discloses three unreleased internal models, and admits its clearest capability tests have stopped working.
3h ago
