Zizhao Yuan, Zhengtu Liang, Taowen Wang, Qiwei Liang, Yichi Wang, Yunheng Wang, Yuetong Fang, Lusong Li et al.
This paper proposes a world model for high-degree-of-freedom robot control that structures action conditioning and integrates semantic priors.
Existing action-conditioned world models compress entire action sequences into a single representation, which is ineffective for modeling the fine-grained, complex motions required for high-degree-of-freedom dexterous control, leading to optimization imbalance and poor action fidelity.
The authors redefine action conditioning as a structured process instead of global compression. They preserve dimension-level semantics via action tokenization and align action signals with visual dynamics through local refinement and global modulation. A semantic branch is further introduced to provide rich object-scene priors, enabling the model to capture dynamic visual details.
Experiments on EgoDex and EgoVerse show that combining the semantic branch with structured action conditioning significantly improves FID, FVD, and PCK metrics, demonstrating gains in visual-temporal realism and action-following consistency. The work validates the scalability of the structured action-conditioning design, suggesting a new direction for scaling world models to high-degree-of-freedom control.