TL;DR
MOSS-TTS Family is an open-source model family for high-fidelity, high-expressiveness speech and sound generation, supporting various real-world scenarios including long-form speech, multi-speaker dialogue, voice/character design, environmental sound effects, and real-time streaming TTS.
Key features
High-quality speech synthesis: Stable long-form speech generation, multi-speaker dialogue support, voice cloning and character design.
Environmental sound generation: MOSS-SoundEffect-v2.0 uses a DiT backbone and Flow Matching objective to generate bilingual sound effects at 48kHz, up to 30 seconds.
Real-time streaming TTS: MOSS-TTS-Realtime and MOSS-TTS-Nano (approx. 100M parameters) can stream output even on a 4-core CPU.
Multilingual support: Multilingual synthesis via language tags, punctuation-based prosody control, explicit pause control ([pause X.Ys]).
Open source and extensibility: Models and code released on Hugging Face, ModelScope, GitHub; API documentation provided.
When to use it
Applications requiring high-quality speech synthesis (e.g., audiobooks, virtual assistants, game NPCs).
Scenarios needing multi-speaker dialogue or voice character design.
Environmental sound effect generation (e.g., games, movies, VR).
Services requiring real-time speech synthesis (e.g., real-time interpretation, streaming broadcasts).