StepAudio 2.5 TTS
StepAudio 2.5 TTS is a contextual text-to-speech model developed by the Shanghai-based AI lab StepFun. Released as part of the Step-Audio 2.5 multimodal family, it represents a transition from traditional tag-based speech synthesis to a context-aware generation pipeline. The model is designed to produce high-fidelity, expressive speech that naturally incorporates paralinguistic cues—such as sighs, laughter, and subtle intonations—interpreting the emotional intent behind the text rather than providing a flat reading.
A defining feature of the model is its dual-level contextual control, which utilizes both Global Context and Inline Context. Users can define an overall emotional register or character persona through a global instruction parameter, while simultaneously using inline commands to sculpt local delivery details. StepAudio 2.5 TTS also supports zero-shot voice cloning, allowing developers to generate speech in any target voice using a short reference audio sample without requiring specialized fine-tuning.
Technically, the model is integrated into a unified multimodal foundation (Step-Audio-2.5) that shares a backbone designed for both speech understanding and generation. This architecture allows the model to maintain character consistency across long-turn interactions, a capability optimized through roleplay-specific Reinforcement Learning from Human Feedback (RLHF). In benchmark evaluations, such as the Artificial Analysis Speech Arena, the model has been noted for its ability to compete with top-tier commercial systems in terms of prosodic naturalness and emotional accuracy.
To control the output, users can provide natural language instructions. When using the model's API, text placed inside parentheses, such as (excited) or (whispering), is treated as a delivery instruction and is not spoken. Global instructions can be used to establish complex personas, specifying traits like "warm and patient" or "authoritative and deep," which the model adheres to throughout the generation process.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow StepAudio 2.5 TTS ranks
StepAudio 2.5 TTS is highlighted in the table below. Switch the metric to see how the ordering changes.