Breeze TTS 2
Breeze TTS 2 is an open-weight, 3-billion-parameter text-to-speech model developed by BreezeBlue for real-time, expressive voice synthesis. It is designed for low-latency performance, reaching a time-to-first-audio of under 40 milliseconds. The model is bilingual, supporting English and Chinese, and has demonstrated competitive performance against both open-source and proprietary frontier systems on industry leaderboards.
The model's architecture integrates a T5-Gemma-2 text encoder with a Qwen3 backbone and a 15-step depth decoder. This structure supports three distinct control modes: Voice Clone, which replicates a speaker's timbre and style from a reference clip; Voice Design, which creates new voices from text descriptions; and Voice Direction, which allows users to steer the emotion and delivery of a voice via natural language instructions.
Additionally, Breeze TTS 2 supports inline vocal events, allowing the synthesis of non-speech sounds like laughs, sighs, and coughs by inserting parenthetical tags in English or bracketed tags in Chinese. It also supports real-time streaming at a factor of approximately 3.1x faster than playback speed. The model weights are provided for research and non-commercial use, while the accompanying inference code is released under an Apache 2.0 license.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Breeze TTS 2 ranks
Breeze TTS 2 is highlighted in the table below. Switch the metric to see how the ordering changes.