AI Safety
Frontier Models Found 29 Real Bugs in Hardened Code. They Also Filed 81 That Were Not.
A pentest firm spent $3,140 and 1.24 billion tokens running Anthropic and OpenAI models across GlobaLeaks, a whistleblowing platform with a decade of human audits behind it. The headline number is the confirmations. The useful number is the ratio.
5h ago
