Yuhao Dong, Zuyan Liu, Shulin Tian, Yongming Rao, Ziwei Liu
A unified framework introducing multi-agent architecture and spatial-temporal reinforcement learning algorithms (ST-GRPO, J-GRPO) for long-chain visual reasoning in multimodal LLMs.
Multimodal LLMs lack performance in complex visual reasoning tasks requiring long reasoning chains, and there is a critical scarcity of high-quality long-chain reasoning data and optimized training pipelines.
1) A scalable data generation pipeline with multi-granularity assessment autonomously synthesizes structured, complex reasoning trajectories across image and video domains. 2) A dual-agent architecture comprising a reasoning agent and a summary agent separates the reasoning process, and an iterative self-improvement loop leverages feedback from the summary agent. 3) To overcome limitations of DPO, we introduce ST-GRPO for enhanced spatial-temporal reasoning and J-GRPO for improved evaluative robustness.
Significant performance gains on image and video reasoning benchmarks using base models like LLaVA-NeXT and Qwen2.5-VL, while maintaining strong performance on existing perception-oriented tasks.