An Italian offensive-security firm pointed frontier models from Anthropic and OpenAI at the source code of GlobaLeaks, an open-source whistleblowing platform that has absorbed roughly a decade of independent human audits, and published what came back: 29 confirmed vulnerabilities, 12 denial-of-service issues and 42 hardening observations.

The target was not a soft one

That matters more than the count. GlobaLeaks is used by newsrooms and anti-corruption bodies and has been reviewed repeatedly by people paid to break it. Finding anything there is a stronger claim than finding bugs in an unexamined codebase.

The ratio the headline hides

The models produced 110 candidate findings. Human triage confirmed 29. That is a precision of roughly 26 percent — about three quarters of what the models raised did not survive review. The syndicated headline, that AI can identify real software vulnerabilities, is true and omits the tax.

What the $77 leaves out

The study reports about $77 per confirmed finding, from a total spend of roughly $3,140 across 12,000 requests and 1.24 billion tokens. Its own text states that figure is before human validation. The expert triage that turned 110 guesses into 29 findings is the step that produced the confirmations, and it is not in the number. The study's stated conclusion concedes that experts retain the decisive role.

Read the cost split, not the total

One figure travels further than the rest: the advanced reasoning model consumed 62 percent of the budget while processing 7 percent of the tokens. That is the real economics of reasoning models in review work — spend concentrates in a thin slice of the job. Two caveats on provenance. This is a vendor self-report from a firm that sells exactly this service. And the 29 figure carries no published severity breakdown, so it should not be read as 29 criticals; the remaining 54 items are denial-of-service and hardening suggestions, not exploitable holes.