AI Research
New Benchmark Finds AI Reading X-Rays Can Be Dangerously Confident When It's Wrong
RadLE 2.0 tested 16 frontier models on 200 X-ray cases with a scoring system that penalizes confident errors; human radiologists scored 988.7 out of 2,000 to the best AI's 758.
1d ago
