C. Liang, Pu Tian, Caitlyn Heqi Yin, Yao Yua, An-Hou Wei, Ming Li, Xinyuan Song, Tianyang Wang et al.
This is a comprehensive survey paper that systematically analyzes and guides the architectures, training, applications, and challenges of Multimodal Large Language Models (MLLMs) for vision-language tasks.
The field of MLLMs is advancing rapidly, but there has been a lack of systematic organization and insights regarding architecture design, training pipelines, evaluation methodologies, and key challenges like scalability, robustness, and inference cost.
The survey analyzes core MLLM components such as visual encoders, language model backbones, and connector modules. It reviews training pipelines including contrastive pre-training, instruction tuning, and preference alignment. It also highlights fundamental constraints like information bottlenecks, data-processing limits, and statistical co-occurrence bias, and examines prominent MLLM implementations through task-level analysis and system-level case studies.
It presents a unified taxonomy of the MLLM design space and a comparative overview of representative models and evaluation benchmarks. By addressing key challenges in scalability, memory, energy use, inference cost, robustness, and cross-modal learning, and discussing ethical considerations, the survey offers theoretical frameworks and practical insights for researchers and practitioners at the intersection of NLP and computer vision.