Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Speech/StepAudio 2.5 TTS
StepFun logoStepFun·Speech Generation

StepAudio 2.5 TTS

View rankingsstepfun.com
AA Arena#7
ReleasedApr 2026

StepAudio 2.5 TTS is a contextual text-to-speech model developed by the Shanghai-based AI lab StepFun. Released as part of the Step-Audio 2.5 multimodal family, it represents a transition from traditional tag-based speech synthesis to a context-aware generation pipeline. The model is designed to produce high-fidelity, expressive speech that naturally incorporates paralinguistic cues—such as sighs, laughter, and subtle intonations—interpreting the emotional intent behind the text rather than providing a flat reading.

A defining feature of the model is its dual-level contextual control, which utilizes both Global Context and Inline Context. Users can define an overall emotional register or character persona through a global instruction parameter, while simultaneously using inline commands to sculpt local delivery details. StepAudio 2.5 TTS also supports zero-shot voice cloning, allowing developers to generate speech in any target voice using a short reference audio sample without requiring specialized fine-tuning.

Technically, the model is integrated into a unified multimodal foundation (Step-Audio-2.5) that shares a backbone designed for both speech understanding and generation. This architecture allows the model to maintain character consistency across long-turn interactions, a capability optimized through roleplay-specific Reinforcement Learning from Human Feedback (RLHF). In benchmark evaluations, such as the Artificial Analysis Speech Arena, the model has been noted for its ability to compete with top-tier commercial systems in terms of prosodic naturalness and emotional accuracy.

To control the output, users can provide natural language instructions. When using the model's API, text placed inside parentheses, such as (excited) or (whispering), is treated as a delivery instruction and is not spoken. Global instructions can be used to establish complex personas, specifying traits like "warm and patient" or "authoritative and deep," which the model adheres to throughout the generation process.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How StepAudio 2.5 TTS ranks

StepAudio 2.5 TTS is highlighted in the table below. Switch the metric to see how the ordering changes.

ELO arena ranking for speech synthesis quality based on blind pairwise human preference comparisons. Data from Artificial Analysis

#
1
AlibabaAlibaba
Alibaba
Qwen-Audio-3.0-TTS-Plus
Alibaba
1,234±16Jul 2026
2
SpeechifyAISpeechifyAI
SpeechifyAI
Simba 3.2
SpeechifyAI
1,230±16Jul 2026
3
GoogleGoogle
Google
Gemini 3.1 Flash TTS
Google
1,215±13Apr 2026
4
InworldInworld
Inworld
Inworld TTS 1.5 Max
Inworld
1,212±19Jan 2026
5
CartesiaCartesia
Cartesia
Sonic 3.5
Cartesia
1,208±14May 2026
6
AlibabaAlibaba
Alibaba
Fun-Realtime-TTS
Alibaba
1,205±15May 2026
7
StepFunStepFun
StepFun
StepAudio 2.5 TTS
StepFun
1,199±16Apr 2026
8
Smallest.aiSmallest.ai
Smallest.ai
Lightning V3.1 Pro (Jul 2026)
Smallest.ai
1,198±17Jul 2026
9
InworldInworld
Inworld
Realtime TTS 1.5 Max
Inworld
1,198±14Jan 2026
10
InworldInworld
Inworld
Realtime TTS-2
Inworld
1,197±14May 2026
11
VUI Labs
Luna TTS
VUI Labs
1,194±16Jun 2026
12
AlibabaAlibaba
Alibaba
Fun-Realtime-TTS-Preview
Alibaba
1,193±15May 2026
13
MiniMaxMiniMax
MiniMax
Speech 2.8 HD
MiniMax
1,178±12Feb 2026
14
StepFunStepFun
StepFun
StepAudio 2.5 TTS
StepFun
1,175±18Apr 2026
15
ElevenLabsElevenLabs
ElevenLabs
Eleven v3
ElevenLabs
1,175±12Jun 2025
16
asyncasync
async
Async Flash v1.5
async
1,164±15May 2026
17
InworldInworld
Inworld
Inworld TTS 1 Max
Inworld
1,161±13Aug 2025
18
InworldInworld
Inworld
Inworld TTS 1.5 Mini
Inworld
1,159±16Jan 2026
19
MiniMaxMiniMax
MiniMax
Speech 2.8 Turbo
MiniMax
1,150±12Jan 2026
20
Smallest.aiSmallest.ai
Smallest.ai
Lightning V3.1 Pro TTS (Jun 2026)
Smallest.ai
1,146±15Jun 2026
21
StepFunStepFun
StepFun
Step TTS 2 (Mar 2026)
StepFun
1,142±14Mar 2026
22
asyncasync
async
Async Pro v1.0
async
1,140±15May 2026
23
Fish AudioFish Audio
Fish Audio
Fish Audio S2.1 Pro
Fish Audio
1,138±15Jun 2026
24
MiniMaxMiniMax
MiniMax
Speech 2.6 HD
MiniMax
1,138±12Oct 2025
25
InworldInworld
Inworld
Realtime TTS 1.5 Mini
Inworld
1,137±13Jan 2026
26
MiniMaxMiniMax
MiniMax
Speech 2.6 Turbo
MiniMax
1,128±12Oct 2025
27
Microsoft AzureMicrosoft Azure
Microsoft Azure
Azure HD 2.5
Microsoft Azure
1,126±13Nov 2025
28
SpeechifySpeechify
Speechify
SIMBA 3.0
Speechify
1,122±14Feb 2026
29
InworldInworld
Inworld
Inworld TTS 1
Inworld
1,121±12Aug 2025
30
Fish AudioFish Audio
Fish Audio
Fish Audio S2 Pro
Fish Audio
1,120±14Mar 2026
31
StepFunStepFun
StepFun
Step Audio EditX (Mar 2026)
StepFun
1,113±14Mar 2026
32
MiniMaxMiniMax
MiniMax
Speech-02-HD
MiniMax
1,112±12May 2025
33
SpaceXAISpaceXAI
SpaceXAI
xAI Text to Speech
SpaceXAI
1,108±21Mar 2026
34
StepFunStepFun
StepFun
Step TTS 2open weights
StepFun
1,105±19Aug 2025
35
OpenAIOpenAI
OpenAI
TTS-1 HD
OpenAI
1,103±12Nov 2023
36
ElevenLabsElevenLabs
ElevenLabs
Turbo v2.5
ElevenLabs
1,102±11Jul 2024
37
ElevenLabsElevenLabs
ElevenLabs
Multilingual v2
ElevenLabs
1,102±11Aug 2023
38
ElevenLabsElevenLabs
ElevenLabs
ElevenLabs v3 - Alpha
ElevenLabs
1,095±11Jun 2025
39
Gradium
Gradium TTS
Gradium
1,094±16Mar 2026
40
Smallest.aiSmallest.ai
Smallest.ai
Lightning V3.1 Pro
Smallest.ai
1,090±14Mar 2026
41
Maya ResearchMaya Research
Maya Research
Maya 2 Flash
Maya Research
1,089±17Jul 2026
42
Gradium
Gradium TTS (Jun 2026)
Gradium
1,088±15Jun 2026
43
ElevenLabsElevenLabs
ElevenLabs
Flash v2.5
ElevenLabs
1,086±11Dec 2024
44
MiniMaxMiniMax
MiniMax
Speech-02-Turbo
MiniMax
1,084±12Apr 2025
45
Resemble AI
Chatterbox HD
Resemble AI
1,083±14May 2025
46
OpenAIOpenAI
OpenAI
TTS-1
OpenAI
1,083±11Nov 2023
47
MistralMistral
Mistral
Voxtral TTS
Mistral
1,075±14Mar 2026
48
GoogleGoogle
Google
Gemini 2.5 Flash Lite TTS
Google
1,075±13Jul 2025
49
GoogleGoogle
Google
Studio
Google
1,072±12Oct 2022
50
CartesiaCartesia
Cartesia
Sonic 3
Cartesia
1,070±12Oct 2025
51
Fish AudioFish Audio
Fish Audio
OpenAudio S1
Fish Audio
1,068±12Jun 2025
52
OpenAIOpenAI
OpenAI
GPT-Realtime-2
OpenAI
1,066±15May 2026
53
XiaomiXiaomi
Xiaomi
MiMo-V2.5-TTS
Xiaomi
1,064±14Apr 2026
54
SpeechifySpeechify
Speechify
SIMBA 1.6
Speechify
1,064±14Nov 2024
55
AmazonAmazon
Amazon
Polly Generative
Amazon
1,064±12May 2024
56
StepFunStepFun
StepFun
Step Audio EditXopen weights
StepFun
1,061±18Nov 2025
57
MiniMaxMiniMax
MiniMax
T2A-01-HD
MiniMax
1,061±12Jan 2025
58
NVIDIANVIDIA
NVIDIA
Magpie-Multilingual 357M (Feb 2026)
NVIDIA
1,060±14Feb 2026
59
KokoroKokoro
Kokoro
Kokoro 82M v1.0
Kokoro
1,057±11Jan 2025
60
Hume AIHume AI
Hume AI
Octave 2
Hume AI
1,056±13Oct 2025
61
OpenAIOpenAI
OpenAI
GPT-4o mini TTS
OpenAI
1,056±22Mar 2025
62
HithinkHithink
Hithink
Speech 2.6
Hithink
1,055±15Jul 2026
63
GoogleGoogle
Google
Chirp 3: HD
Google
1,054±12Apr 2025
64
Maya ResearchMaya Research
Maya Research
Maya 2 Global
Maya Research
1,048±17Jul 2026
65
AmazonAmazon
Amazon
Polly Long-Form
Amazon
1,047±14Nov 2023
66
Maya ResearchMaya Research
Maya Research
Maya1
Maya Research
1,046±12Nov 2025
67
asyncasync
async
Async Flash v1.0
async
1,046±12Jul 2025
68
Fish AudioFish Audio
Fish Audio
OpenAudio S1 Mini
Fish Audio
1,045±20Jun 2025
69
GoogleGoogle
Google
Gemini 2.5 Flash TTS
Google
1,044±14Sep 2025
70
GoogleGoogle
Google
Journey
Google
1,043±14Dec 2023
71
RimeRime
Rime
Coda
Rime
1,042±14May 2026
72
CartesiaCartesia
Cartesia
Sonic English (Oct '24)
Cartesia
1,038±12Oct 2024
73
GoogleGoogle
Google
Gemini 2.5 Flash TTS (Dec 2025)
Google
1,037±13Dec 2025
74
GoogleGoogle
Google
Gemini 2.5 Pro TTS
Google
1,036±14Dec 2025
75
GoogleGoogle
Google
Gemini 2.5 Pro (Dec 2025)
Google
1,035±13Dec 2025
76
SpeechifySpeechify
Speechify
Simba
Speechify
1,035±12Jun 2024
77
Microsoft AzureMicrosoft Azure
Microsoft Azure
MAI-Voice-1
Microsoft Azure
1,029±14Apr 2026
78
Microsoft AzureMicrosoft Azure
Microsoft Azure
Azure Neural
Microsoft Azure
1,027±19Sep 2018
79
Smallest.aiSmallest.ai
Smallest.ai
Lightning v3.1
Smallest.ai
1,025±14Mar 2026
80
Hume AIHume AI
Hume AI
Octave TTS
Hume AI
1,025±12Feb 2025
81
MiniMaxMiniMax
MiniMax
T2A-01-Turbo
MiniMax
1,022±12Jan 2025
82
XiaomiXiaomi
Xiaomi
MiMo-V2-TTS
Xiaomi
1,017±15Mar 2026
83
Resemble AI
Chatterbox
Resemble AI
1,013±12Jun 2025
84
Fish AudioFish Audio
Fish Audio
Fish Speech 1.5
Fish Audio
1,010±12Dec 2024
85
NVIDIANVIDIA
NVIDIA
Magpie-Multilingual 357M
NVIDIA
1,005±12Aug 2025
86
RimeRime
Rime
Arcana v3
Rime
1,003±14Feb 2026
87
ZyphraZyphra
Zyphra
Zonos-v0.1
Zyphra
1,000±0Feb 2025
88
Murf AIMurf AI
Murf AI
Murf Speech Gen 2
Murf AI
972.0±12Mar 2024
89
LMNTLMNT
LMNT
LMNT
LMNT
970.0±13Sep 2023
90
Microsoft AzureMicrosoft Azure
Microsoft Azure
VibeVoice 1.5B
Microsoft Azure
966.0±14Aug 2025
91
StepFunStepFun
StepFun
Step TTS Miniopen weights
StepFun
959.0±10Feb 2025
92
Microsoft AzureMicrosoft Azure
Microsoft Azure
VibeVoice 7B
Microsoft Azure
957.0±14Aug 2025
93
OpenVoiceOpenVoice
OpenVoice
OpenVoice v2
OpenVoice
956.0±14Apr 2024
94
NVIDIANVIDIA
NVIDIA
Magpie Multilingual
NVIDIA
940.0±15Mar 2025
95
NeuphonicNeuphonic
Neuphonic
Neuphonic TTS
Neuphonic
934.0±14Oct 2024
96
AlibabaAlibaba
Alibaba
Qwen3 TTS Flash
Alibaba
931.0±15Sep 2025
97
AlibabaAlibaba
Alibaba
Qwen3 TTS
Alibaba
917.0±14Jan 2026
98
CoquiCoqui
Coqui
XTTS v2
Coqui
916.0±15Oct 2023
99
GoogleGoogle
Google
WaveNet
Google
902.0±12Sep 2016
100
StyleTTS StyleTTS
StyleTTS
StyleTTS 2
StyleTTS
890.0±16Jun 2023
101
RimeRime
Rime
Mist V2
Rime
885.0±15Feb 2025
102
GoogleGoogle
Google
Neural2
Google
883.0±12Jun 2022
103
AmazonAmazon
Amazon
Polly Neural
Amazon
881.0±14Jul 2019
104
GoogleGoogle
Google
Standard
Google
876.0±12Mar 2018
105
Murf AIMurf AI
Murf AI
Falcon (Beta)
Murf AI
857.0±15Nov 2025
106
NoizNoiz
Noiz
Noiz TTS
Noiz
853.0±17Nov 2025
107
MetaVoiceMetaVoice
MetaVoice
MetaVoice v1
MetaVoice
835.0±18Feb 2024
108
AmazonAmazon
Amazon
Polly Standard
Amazon
810.0±16Nov 2016

Artificial Analysis Arena data from Artificial Analysis· updated Jul 26, 2026