AI Safety
Anthropic Researchers Tested Their Interpretability Tools and Found No Uplift at All
Sparse autoencoders, natural-language autoencoders and activation oracles were measured against a baseline that just reads the transcript. None of them beat it.
2h ago
