AI models are getting good at reading medical images — and dangerously bad at knowing when they are wrong. A new benchmark detailed on July 19, RadLE 2.0, or "Radiology's Last Exam," finds that frontier models will confidently misdiagnose X-rays rather than defer, a failure mode that matters more in medicine than raw accuracy does.

How the test works

Built by the CRASH Lab at India's Ashoka University, RadLE 2.0 evaluated 16 frontier models — including Claude Fable 5, Google's Gemini 3 Pro, Meta's Muse Spark 1.1 and Grok 4.5 — on 200 X-ray cases. Crucially, it uses a confidence-weighted scoring system that rewards honest uncertainty and penalizes confident wrong answers, rather than simply counting correct guesses.

The gap to humans

Under that scoring, human radiologists scored 988.7 out of 2,000, well ahead of the best AI model's 758. Notably, on raw accuracy alone Gemini 3 Pro nearly matched humans — meaning the models' weakness is not only what they get wrong, but how they behave when uncertain.

The overconfidence problem

That is the report's central warning. Many models produced highly confident misdiagnoses, and the authors note they "would have scored much better if they had stayed quiet more often instead of guessing." A model that doesn't know when to hand off to a human is a specific clinical hazard, because a confident wrong read can be more harmful than an admitted "I'm not sure."

Why the scoring matters

Most benchmarks reward the right answer and ignore the manner of getting there. By penalizing confident errors, RadLE 2.0 measures something closer to what safe clinical deployment actually requires — calibrated uncertainty — and its results suggest today's models are not there yet, even where their raw accuracy looks close to human.