Qwen-Audio-3.0-TTS-Plus
Qwen-Audio-3.0-TTS-Plus is a proprietary text-to-speech model developed by Alibaba's Qwen team and distributed through Alibaba Cloud Model Studio. It belongs to the Qwen-Audio-TTS model family and is positioned as a higher-quality alternative to the lower-latency qwen-audio-3.0-tts-flash model.
The model supports real-time and non-real-time speech synthesis. Real-time synthesis is provided through a streaming WebSocket interface, while HTTP interfaces are available for standard and batch generation. Supported output settings include speech rate, pitch, volume, bitrate, audio format, and sampling rates of up to 48 kHz.
Qwen-Audio-3.0-TTS-Plus includes natural-language instruction control for modifying speaking style and delivery. Instructions can be used to request characteristics such as slower pacing, increased energy, a softer tone, or other expressive variations. The model also supports inline tags for emotions and vocal effects, including laughter, sighing, whispering, shouting, singing, excitement, anger, curiosity, and empathy.
Speech can be generated using system voices supplied by Alibaba Cloud or custom voices created through voice cloning. Voice cloning uses a reference recording between 3 and 30 seconds in length. The voice-enrollment service can apply preprocessing operations such as noise reduction, audio enhancement, and volume normalization. Language hints may also be supplied to improve the processing of multilingual reference recordings.
The model supports Mandarin Chinese, English, and a broader multilingual set through custom voice cloning. Supported languages and varieties include several Chinese dialects, Japanese, Korean, Russian, French, German, Portuguese, Thai, Indonesian, Vietnamese, Spanish, Italian, Malay, Filipino, and Arabic. Language availability varies by voice and service configuration.
Qwen-Audio-3.0-TTS-Plus supports instruction-guided expression and voice cloning, but does not provide the text-prompt-based Voice Design capability associated with some Qwen3-TTS models. Support for SSML, timestamps, and other synthesis features depends on the selected voice and API interface.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Qwen-Audio-3.0-TTS-Plus ranks
Qwen-Audio-3.0-TTS-Plus is highlighted in the table below. Switch the metric to see how the ordering changes.