Fan Bai, Yuxin Du, Tiejun Huang, Max Q.‐H. Meng, Bo Zhao
We propose M3D-LaMed, a multimodal large language model for 3D medical image analysis, along with a large-scale dataset (M3D-Data) and benchmark (M3D-Bench).
Existing medical image analysis research has primarily focused on 2D images, leaving the rich spatial information of 3D medical images underutilized in multimodal large language models. Additionally, there is a lack of large-scale datasets and evaluation benchmarks specifically designed for 3D medical images.
The proposed method outperforms existing solutions. All code, data, and models are publicly available to facilitate further research. This pioneering work applies multimodal large language models to 3D medical image analysis, contributing to clinical diagnosis and treatment.