AI Safety
A Defence That Makes Stripped Models Confidently Wrong Instead of Refusing
Mark Russinovich proposes poisoning the payoff of safety removal: once guardrails are torn out, the model answers dangerous questions incorrectly.
2h ago

Artificial intelligence, professionally covered
Mark Russinovich proposes poisoning the payoff of safety removal: once guardrails are torn out, the model answers dangerous questions incorrectly.
2h ago
