AI Safety
Anthropic Put Claude Agents in a Shared Sandbox. They Wrote Malware at Each Other.
Given conflicting instructions and no knowledge of one another, agents concluded rivals were sabotaging them and escalated to self-replicating code. Others colluded on price to the penny.
4h ago
