Bland Speech v3
Bland Speech v3 is a proprietary text-to-speech (TTS) model developed by Bland AI, specifically engineered for conversational AI and real-time phone interactions. Unlike general-purpose TTS models that prioritize polished, studio-quality narration, v3 is designed to capture the "imperfections" of human speech. It naturally incorporates breaths, stumbles, hesitations, and varying cadences, which are essential for making synthetic voices sound authentic during live telephone conversations.
The model was trained on a dataset comprising over 100 million real human conversations, allowing it to handle complex linguistic nuances such as names, numbers, and emotional shifts. It delivers high-fidelity audio at a 44.1 kHz sample rate in PCM16 WAV format. For developers, the model supports low-latency streaming via HTTP chunked transfer and WebSockets, facilitating near-instant responses in interactive voice applications.
Users can direct the model's performance through both plain-language instructions and technical parameters. Delivery can be modulated using expressiveness and stability controls (0.0 to 1.0) to adjust emotional intensity and consistency. Additionally, the model supports performance tags for precise timing, such as <|0.4|> to insert a 400ms pause, and a "Director" feature that enables prompting for specific tones like "warmer" or "more professional."
Bland Speech v3 has been ranked as a top performer on the Audio Realism Bench, an Elo-based blind listening benchmark. It achieved scores second only to actual human recordings, outperforming many general-purpose and studio-oriented voice models. Key features also include Instant Voice Cloning from 10 seconds of audio and Professional Voice Cloning for higher-tier realism using 30 minutes of verified source material.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Bland Speech v3 ranks
Bland Speech v3 is highlighted in the table below. Switch the metric to see how the ordering changes.