Zhenyu Liu, Yunxin Li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Haolan Chen, Minjun Zhao, Fanbo Meng et al.
This paper identifies the root cause of performance degradation in full-duplex Spoken Language Models (SLMs) as inherent modality interference and proposes a novel end-to-end framework based on hierarchical parameter separation to resolve it.
State-of-the-art full-duplex SLMs suffer from severe modality interference when acoustic and semantic modalities share a deep parameter space, leading to knowledge degradation and compromised semantic integrity, which makes the models feel unnatural and unintelligent.
Through a fine-grained analysis of model optimization dynamics, the authors reveal that modality interference stems from inherent gradient conflicts between acoustic and semantic modeling. They propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers while preserving cross-modality coherence via a dedicated semantic alignment channel.
Extensive experiments on multiple full-duplex benchmarks show significant state-of-the-art improvements in both speech intelligence (+7.4% on Spoken QA) and full-duplex interaction fluidity (+28.5% on FullDuplexBench 1.5) without compromising inference efficiency. This work is the first to uncover and elucidate the root cause of modality interference in full-duplex SLMs and to design an elegant hierarchical model with a practical solution for seamless, high-performance, native intelligent full-duplex SLMs.