Flux TTS
Flux TTS is a conversation-native text-to-speech model developed by Deepgram, specifically engineered for low-latency, real-time voice agent applications. Unlike traditional TTS systems that process text as isolated utterances, Flux TTS is designed to handle the complexities of live dialogue, including streaming inputs from Large Language Models (LLMs), natural interruptions, and multi-turn persistence. It achieves a time-to-first-audio (TTFA) as low as 80 milliseconds, making it suitable for high-performance conversational AI.
The model's architecture introduces a turn-based lifecycle via a dedicated WebSocket endpoint, moving away from one-shot text-to-audio pipelines. This approach allows the model to maintain conversational state and voice consistency across multiple turns, ensuring that short interjections or responses retain the emotional register and tone established earlier in the session. It also features native interruption awareness; when a user "barges in," the API reports the exact text spoken and remaining, allowing developers to keep the LLM context perfectly synchronized with the user's experience.
Flux TTS emphasizes natural expressiveness without the need for manual tuning. It automatically adjusts pacing, emphasis, and emotional inflection based on the textual context, eliminating the requirement for SSML (Speech Synthesis Markup Language), prompt engineering, or style tags. Key voices in the family, such as "flux-alexis-en" and "flux-kit-en," are optimized for human-like interaction in customer service, telephony, and assistive AI roles.
For optimal performance, the model supports mid-stream configuration, enabling developers to adjust parameters like speaking speed dynamically without reconnecting the session. It is compatible with modern NVIDIA GPU architectures (Ampere or later) and is typically deployed through Deepgram's hosted API or as a self-hosted containerized solution for enterprise environments requiring data sovereignty.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Flux TTS ranks
Flux TTS is highlighted in the table below. Switch the metric to see how the ordering changes.