A prospective, single-centre crossover reader study published on 21 August in Academic Radiology tested four commercially available chest-radiography AI products against unaided reading. Five readers with one to six years of experience assessed 1,200 consecutive patients and 1,861 radiographs for pulmonary infiltrates, pleural effusions, mediastinal masses, pneumothorax and pulmonary nodules, under five conditions, with a 14-day washout between sessions and randomised case order. The reference standard was the finalised clinical report, supplemented by confirmatory CT where available.

The primary outcome

"AI assistance did not improve diagnostic accuracy." For pleural effusions and pulmonary nodules, accuracy decreased in several reader-tool pairings, and the mechanism named is increased false positives — the tool flags something, the reader agrees, the finding is not there.

What the common framing gets wrong

This is the exact inverse of how radiology AI results are normally presented. Every marketable number here is a secondary outcome and every one of them improved: three of five residents read faster, with median reductions of 6 to 17 seconds per case (p≤0.031); four of five reported higher confidence; senior consultations fell in selected pairings and CT escalation dropped in one. Meanwhile the primary outcome, diagnostic performance, did not move — and went backwards for two findings. Higher confidence alongside lower accuracy is the textbook signature of automation bias, and the authors name it explicitly. A write-up leading on "AI cuts reading time and boosts radiologist confidence" would be accurate in every particular and wrong in aggregate.

Why "commercially available" is the load-bearing phrase

These are not research prototypes. They are marketed, cleared products already deployed in clinical workflows, tested head to head in a real-world setting — which is a substantially stronger evidence design than the retrospective, single-tool studies that underpin most procurement decisions.

What the study does not establish

It is monocentric, the readers are relatively junior, and the reference standard is the finalised report rather than patient outcomes. The authors close by saying clinical trials are still needed to address whether patients do better, and that deployment requires careful local adaptation to avoid automation bias while capturing the efficiency gains. The efficiency gains, notably, are real — they are simply not the thing the tools are sold as delivering.