A systematic analysis published in PLOS Digital Health on 19 August examined every AI/ML-enabled medical device cleared or approved by the FDA through 5 December 20251,357 of them — and asked how many had been tested on whether patients actually do better.

The numbers

34 devices (2.5%) are linked to a registered prospective trial. 12 (0.9%) have posted results. 12 (0.9%) have a peer-reviewed publication. And 3 (0.2%) were evaluated on patient-centered outcomes — mortality, morbidity or readmissions. Of the registered trials, 62% were observational rather than randomised. The device population is heavily concentrated: 1,059 radiology devices (78%), 122 cardiovascular (9%), 68 neurology (5%).

What the common telling gets wrong

The coverage is running this as "AI medical devices are unproven" or "untested." That is not what the paper measures. Almost all 1,357 went through 510(k) substantial-equivalence clearance, which by statute does not require an outcomes trial — it requires the device to be equivalent to an existing predicate. "Only 3 of 1,357" is not a finding that the FDA broke its own rules; it is a finding about what the rules ask for.

Nor were these devices untested. They were tested — on diagnostic accuracy, AUC, reader agreement, standalone performance against a reference standard. What is missing is the downstream link from a better read to a better patient. And with 78% of the population in radiology, the headline is largely a statement about imaging triage software, not about AI making treatment decisions.

Who was not in the studies

Most of the included studies excluded pregnant women, adults over 75 and non-English speakers, and ran in well-resourced health systems. The authors — from Toronto, MIT Critical Data, Harvard Chan, Johns Hopkins, Mbarara University of Science and Technology and Bergen — declare no competing interests.

A rule for reading clearance news

The practical takeaway is a test any reader can apply: when a company announces an FDA clearance, check whether the release contains an outcome number. 99.8% of the time nobody has produced one — which makes its absence the norm rather than a scandal, and its presence genuinely notable.