Qi Chen, Wenxuan Li, Pedro R. A. S. Bassi, Xinze Zhou, Jakob Wasserthal, Ibrahim Ethem Hamamci, Sezgin Er, Ashwin Kumar et al.
This paper proposes a large-scale benchmark to systematically evaluate tumor-detection AI models for biases across patient demographics and imaging protocols.
AI models for medical imaging often exhibit inconsistent performance in real-world clinical settings due to variations in patient demographics and imaging protocols. There is a lack of standardized methods to quantitatively evaluate model performance on rare or underrepresented patient subgroups.
The authors curated a benchmark of 85,355 CT scans and used large language models (LLMs) to extract and organize subgroup information from clinical data. They systematically evaluated 12 tumor-detection AI models across dimensions like tumor size, location, patient subgroups (age, sex, race), and imaging protocols (e.g., contrast phases).
The benchmark reveals that state-of-the-art AI models, optimized for average accuracy, perform poorly on rare subgroups such as young, female African Americans. This work highlights the critical need for rigorous, subgroup-level evaluation in medical imaging and computer vision, providing a foundation for building more reliable and robust AI models.