Nature Medicine published two documents on the morning of 3 September, registered seconds apart: a peer-reviewed Matters Arising challenging the June 2026 paper that claimed general-purpose frontier LLMs outperform specialised clinical AI tools, and the original authors' Reply. The journal released them as a deliberate pair.
What the critique argues
That the benchmarks producing the study's largest effects are confounded in ways the authors themselves acknowledge, and that the remaining evaluation is limited in scope, overstates its precision, and relies on inference settings that are not adequately described or controlled. The critique includes a figure showing recoverable MedQA benchmark content from a frontier model — test-set memorisation, demonstrated rather than asserted.
What the authors concede
The Reply agrees that possible MedQA contamination and HealthBench evaluator affinity may limit generalisability — evaluator affinity being the tendency of an LLM judge to prefer output from its own model family — and states that the paper treats those two benchmarks as supplementary. The authors maintain the primary conclusion under their study and deployment conditions. That primary conclusion rests on the real clinical query benchmark: 100 de-identified physician queries, reviewed blind by 12 US clinicians.
What the common framing gets wrong
The June result circulated as "ChatGPT beats the doctors' reference tools" — a general capability ranking. It was never that, and after 3 September it is narrower still: two of the three legs are conceded as contaminated or self-favouring, and the surviving leg is 100 queries at one institution. The second correction matters for how this gets reported: a Matters Arising is not a correction, a retraction, or an expression of concern. The original paper stands. What changed is the scope its own authors now claim for it, in print.
Why it lands where it does
This is the paper being cited in hospital procurement arguments against paying for specialised clinical AI subscriptions. The exchange sets out, in the journal of record, how much weight that argument can carry — and it puts benchmark contamination and LLM-as-judge preference leakage at the centre of medical purchasing decisions rather than leaving them as machine-learning methodology concerns.
