Qwen-Audio-3.1-TTS-Plus
Qwen-Audio-3.1-TTS-Plus is a high-performance text-to-speech generation model developed by Alibaba as part of the Qwen-Audio-3.1 audio suite. Tuned specifically for high-fidelity narration and professional audio production, the model prioritizes timbre fidelity, prosodic naturalness, and expressive voice generation. It supports broad multilingual coverage spanning 16 languages and multiple regional dialects, delivering advanced capabilities for localized content creation, audiobooks, and narrative media.
The model incorporates natural-language style control, enabling users to steer emotion, speaking pace, role, and situational delivery using simple textual instructions rather than manual acoustic adjustments. It supports fine-grained control tags to trigger paralinguistic elements such as laughter, breathing, or shifts in mood at precise points within the text alongside robust zero-shot voice cloning capabilities.
When designing prompts for the model, providing clear style directions alongside the input script helps shape the desired emotional tone, while clean reference audio clips of up to 30 seconds are recommended for optimal voice identity preservation.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Qwen-Audio-3.1-TTS-Plus ranks
Qwen-Audio-3.1-TTS-Plus is highlighted in the table below. Switch the metric to see how the ordering changes.