AI Safety
Typos and Paraphrases Add Up: Two Researchers Steered Frontier Models With Invisible Cues
Individually meaningless prompt choices combine almost additively. Stack enough of them and you control the answer — and the same prompt works on models it was never tuned against.
2h ago
