Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Speech/StepAudio 2.5 TTS
StepFun logoStepFun·Speech Generation

StepAudio 2.5 TTS

View rankingsstepfun.com
AA Arena#6
ReleasedApr 2026

StepAudio 2.5 TTS is a contextual text-to-speech model developed by the Shanghai-based AI lab StepFun. Released as part of the Step-Audio 2.5 multimodal family, it represents a transition from traditional tag-based speech synthesis to a context-aware generation pipeline. The model is designed to produce high-fidelity, expressive speech that naturally incorporates paralinguistic cues—such as sighs, laughter, and subtle intonations—interpreting the emotional intent behind the text rather than providing a flat reading.

A defining feature of the model is its dual-level contextual control, which utilizes both Global Context and Inline Context. Users can define an overall emotional register or character persona through a global instruction parameter, while simultaneously using inline commands to sculpt local delivery details. StepAudio 2.5 TTS also supports zero-shot voice cloning, allowing developers to generate speech in any target voice using a short reference audio sample without requiring specialized fine-tuning.

Technically, the model is integrated into a unified multimodal foundation (Step-Audio-2.5) that shares a backbone designed for both speech understanding and generation. This architecture allows the model to maintain character consistency across long-turn interactions, a capability optimized through roleplay-specific Reinforcement Learning from Human Feedback (RLHF). In benchmark evaluations, such as the Artificial Analysis Speech Arena, the model has been noted for its ability to compete with top-tier commercial systems in terms of prosodic naturalness and emotional accuracy.

To control the output, users can provide natural language instructions. When using the model's API, text placed inside parentheses, such as (excited) or (whispering), is treated as a delivery instruction and is not spoken. Global instructions can be used to establish complex personas, specifying traits like "warm and patient" or "authoritative and deep," which the model adheres to throughout the generation process.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How StepAudio 2.5 TTS ranks

StepAudio 2.5 TTS is highlighted in the table below. Switch the metric to see how the ordering changes.

ELO arena ranking for speech synthesis quality based on blind pairwise human preference comparisons. Data from Artificial Analysis

#
1
AlibabaAlibaba
Alibaba
Qwen-Audio-3.0-TTS-Plus
Alibaba
1,233±15Jul 2026
2
SpeechifyAISpeechifyAI
SpeechifyAI
Simba 3.2
SpeechifyAI
1,233±15Jul 2026
3
VUI Labs
Luna TTS
VUI Labs
1,215±14Jun 2026
4
GoogleGoogle
Google
Gemini 3.1 Flash TTS
Google
1,213±13Apr 2026
5
InworldInworld
Inworld
Inworld TTS 1.5 Max
Inworld
1,212±19Jan 2026
6
StepFunStepFun
StepFun
StepAudio 2.5 TTS
StepFun
1,207±15Apr 2026
7
AlibabaAlibaba
Alibaba
Fun-Realtime-TTS
Alibaba
1,205±15May 2026
8
CartesiaCartesia
Cartesia
Sonic 3.5
Cartesia
1,202±13May 2026
9
InworldInworld
Inworld
Realtime TTS 1.5 Max
Inworld
1,195±13Jan 2026
10
Smallest.aiSmallest.ai
Smallest.ai
Lightning V3.1 Pro (Jul 2026)
Smallest.ai
1,194±16Jul 2026
11
AlibabaAlibaba
Alibaba
Fun-Realtime-TTS-Preview
Alibaba
1,193±15May 2026
12
InworldInworld
Inworld
Realtime TTS-2
Inworld
1,190±13May 2026
13
MiniMaxMiniMax
MiniMax
Speech 2.8 HD
MiniMax
1,176±12Feb 2026
14
StepFunStepFun
StepFun
StepAudio 2.5 TTS
StepFun
1,175±18Apr 2026
15
ElevenLabsElevenLabs
ElevenLabs
Eleven v3
ElevenLabs
1,172±12Jun 2025
16
InworldInworld
Inworld
Inworld TTS 1 Max
Inworld
1,161±13Aug 2025
17
InworldInworld
Inworld
Inworld TTS 1.5 Mini
Inworld
1,159±16Jan 2026
18
Gradium
Gradium TTS Beta
Gradium
1,155±16Jul 2026
19
asyncasync
async
Async Flash v1.5
async
1,153±14May 2026
20
Murf AIMurf AI
Murf AI
Falcon 2
Murf AI
1,153±16Mar 2026
21
MiniMaxMiniMax
MiniMax
Speech 2.8 Turbo
MiniMax
1,153±12Jan 2026
22
StepFunStepFun
StepFun
Step TTS 2 (Mar 2026)
StepFun
1,142±13Mar 2026
23
Fish AudioFish Audio
Fish Audio
Fish Audio S2.1 Pro
Fish Audio
1,141±14Jun 2026
24
Smallest.aiSmallest.ai
Smallest.ai
Lightning V3.1 Pro TTS (Jun 2026)
Smallest.ai
1,138±14Jun 2026
25
MiniMaxMiniMax
MiniMax
Speech 2.6 HD
MiniMax
1,136±12Oct 2025
26
InworldInworld
Inworld
Realtime TTS 1.5 Mini
Inworld
1,131±12Jan 2026
27
asyncasync
async
Async Pro v1.0
async
1,129±14May 2026
28
Microsoft AzureMicrosoft Azure
Microsoft Azure
Azure HD 2.5
Microsoft Azure
1,129±13Nov 2025
29
Fish AudioFish Audio
Fish Audio
Fish Audio S2 Pro
Fish Audio
1,124±13Mar 2026
30
MiniMaxMiniMax
MiniMax
Speech 2.6 Turbo
MiniMax
1,124±11Oct 2025
31
SpeechifySpeechify
Speechify
SIMBA 3.0
Speechify
1,123±13Feb 2026
32
InworldInworld
Inworld
Inworld TTS 1
Inworld
1,121±12Aug 2025
33
SpaceXAISpaceXAI
SpaceXAI
xAI Text to Speech
SpaceXAI
1,115±17Mar 2026
34
StepFunStepFun
StepFun
Step Audio EditX (Mar 2026)
StepFun
1,112±13Mar 2026
35
MiniMaxMiniMax
MiniMax
Speech-02-HD
MiniMax
1,112±12May 2025
36
StepFunStepFun
StepFun
Step TTS 2open weights
StepFun
1,105±19Aug 2025
37
OpenAIOpenAI
OpenAI
TTS-1 HD
OpenAI
1,103±12Nov 2023
38
ElevenLabsElevenLabs
ElevenLabs
Turbo v2.5
ElevenLabs
1,101±11Jul 2024
39
ElevenLabsElevenLabs
ElevenLabs
Multilingual v2
ElevenLabs
1,101±11Aug 2023
40
ElevenLabsElevenLabs
ElevenLabs
ElevenLabs v3 - Alpha
ElevenLabs
1,095±11Jun 2025
41
Gradium
Gradium TTS
Gradium
1,094±16Mar 2026
42
Gradium
Gradium TTS (Jun 2026)
Gradium
1,091±14Jun 2026
43
Smallest.aiSmallest.ai
Smallest.ai
Lightning V3.1 Pro
Smallest.ai
1,090±14Mar 2026
44
Resemble AI
Chatterbox HD
Resemble AI
1,090±13May 2025
45
OpenAIOpenAI
OpenAI
TTS-1
OpenAI
1,090±11Nov 2023
46
GoogleGoogle
Google
Gemini 2.5 Flash Lite TTS
Google
1,083±12Jul 2025
47
MiniMaxMiniMax
MiniMax
Speech-02-Turbo
MiniMax
1,083±12Apr 2025
48
ElevenLabsElevenLabs
ElevenLabs
Flash v2.5
ElevenLabs
1,083±11Dec 2024
49
GoogleGoogle
Google
Studio
Google
1,080±12Oct 2022
50
HithinkHithink
Hithink
Speech 2.6
Hithink
1,075±14Jul 2026
51
Fish AudioFish Audio
Fish Audio
OpenAudio S1
Fish Audio
1,075±11Jun 2025
52
Maya ResearchMaya Research
Maya Research
Maya 2 Flash
Maya Research
1,074±15Jul 2026
53
MistralMistral
Mistral
Voxtral TTS
Mistral
1,074±13Mar 2026
54
CartesiaCartesia
Cartesia
Sonic 3
Cartesia
1,072±12Oct 2025
55
NVIDIANVIDIA
NVIDIA
Magpie-Multilingual 357M (Feb 2026)
NVIDIA
1,068±13Feb 2026
56
AmazonAmazon
Amazon
Polly Generative
Amazon
1,068±12May 2024
57
GoogleGoogle
Google
Journey
Google
1,068±13Dec 2023
58
OpenAIOpenAI
OpenAI
GPT-Realtime-2
OpenAI
1,066±14May 2026
59
SpeechifySpeechify
Speechify
SIMBA 1.6
Speechify
1,066±13Nov 2024
60
MiniMaxMiniMax
MiniMax
T2A-01-HD
MiniMax
1,063±11Jan 2025
61
XiaomiXiaomi
Xiaomi
MiMo-V2.5-TTS
Xiaomi
1,061±13Apr 2026
62
StepFunStepFun
StepFun
Step Audio EditXopen weights
StepFun
1,061±18Nov 2025
63
KokoroKokoro
Kokoro
Kokoro 82M v1.0
Kokoro
1,060±11Jan 2025
64
Hume AIHume AI
Hume AI
Octave 2
Hume AI
1,056±12Oct 2025
65
OpenAIOpenAI
OpenAI
GPT-4o mini TTS
OpenAI
1,056±22Mar 2025
66
GoogleGoogle
Google
Chirp 3: HD
Google
1,054±12Apr 2025
67
GoogleGoogle
Google
Gemini 2.5 Flash TTS (Dec 2025)
Google
1,053±12Dec 2025
68
RimeRime
Rime
Coda
Rime
1,052±13May 2026
69
Maya ResearchMaya Research
Maya Research
Maya 2 Global
Maya Research
1,051±15Jul 2026
70
asyncasync
async
Async Flash v1.0
async
1,049±12Jul 2025
71
Fish AudioFish Audio
Fish Audio
OpenAudio S1 Mini
Fish Audio
1,048±20Jun 2025
72
Maya ResearchMaya Research
Maya Research
Maya1
Maya Research
1,045±12Nov 2025
73
GoogleGoogle
Google
Gemini 2.5 Flash TTS
Google
1,044±14Sep 2025
74
AmazonAmazon
Amazon
Polly Long-Form
Amazon
1,043±14Nov 2023
75
Boson AIBoson AI
Boson AI
Higgs Audio V3 TTS
Boson AI
1,041±14Jun 2026
76
CartesiaCartesia
Cartesia
Sonic English (Oct '24)
Cartesia
1,041±12Oct 2024
77
SpeechifySpeechify
Speechify
Simba
Speechify
1,038±11Jun 2024
78
GoogleGoogle
Google
Gemini 2.5 Pro (Dec 2025)
Google
1,037±13Dec 2025
79
GoogleGoogle
Google
Gemini 2.5 Pro TTS
Google
1,036±14Dec 2025
80
Microsoft AzureMicrosoft Azure
Microsoft Azure
MAI-Voice-1
Microsoft Azure
1,031±13Apr 2026
81
MiniMaxMiniMax
MiniMax
T2A-01-Turbo
MiniMax
1,029±11Jan 2025
82
Microsoft AzureMicrosoft Azure
Microsoft Azure
Azure Neural
Microsoft Azure
1,029±19Sep 2018
83
Hume AIHume AI
Hume AI
Octave TTS
Hume AI
1,028±12Feb 2025
84
Smallest.aiSmallest.ai
Smallest.ai
Lightning v3.1
Smallest.ai
1,022±13Mar 2026
85
Resemble AI
Chatterbox
Resemble AI
1,019±12Jun 2025
86
XiaomiXiaomi
Xiaomi
MiMo-V2-TTS
Xiaomi
1,015±14Mar 2026
87
Fish AudioFish Audio
Fish Audio
Fish Speech 1.5
Fish Audio
1,010±12Dec 2024
88
RimeRime
Rime
Arcana v3
Rime
1,005±13Feb 2026
89
NVIDIANVIDIA
NVIDIA
Magpie-Multilingual 357M
NVIDIA
1,004±12Aug 2025
90
ZyphraZyphra
Zyphra
Zonos-v0.1
Zyphra
1,000±0Feb 2025
91
Murf AIMurf AI
Murf AI
Murf Speech Gen 2
Murf AI
978.0±12Mar 2024
92
Microsoft AzureMicrosoft Azure
Microsoft Azure
VibeVoice 1.5B
Microsoft Azure
973.0±14Aug 2025
93
LMNTLMNT
LMNT
LMNT
LMNT
972.0±13Sep 2023
94
Microsoft AzureMicrosoft Azure
Microsoft Azure
VibeVoice 7B
Microsoft Azure
966.0±14Aug 2025
95
StepFunStepFun
StepFun
Step TTS Miniopen weights
StepFun
959.0±10Feb 2025
96
OpenVoiceOpenVoice
OpenVoice
OpenVoice v2
OpenVoice
953.0±14Apr 2024
97
NVIDIANVIDIA
NVIDIA
Magpie Multilingual
NVIDIA
945.0±14Mar 2025
98
AlibabaAlibaba
Alibaba
Qwen3 TTS Flash
Alibaba
941.0±14Sep 2025
99
NeuphonicNeuphonic
Neuphonic
Neuphonic TTS
Neuphonic
936.0±14Oct 2024
100
AlibabaAlibaba
Alibaba
Qwen3 TTS
Alibaba
922.0±13Jan 2026
101
CoquiCoqui
Coqui
XTTS v2
Coqui
917.0±15Oct 2023
102
GoogleGoogle
Google
WaveNet
Google
909.0±12Sep 2016
103
RimeRime
Rime
Mist V2
Rime
893.0±15Feb 2025
104
StyleTTS StyleTTS
StyleTTS
StyleTTS 2
StyleTTS
889.0±15Jun 2023
105
GoogleGoogle
Google
Neural2
Google
887.0±12Jun 2022
106
AmazonAmazon
Amazon
Polly Neural
Amazon
882.0±14Jul 2019
107
GoogleGoogle
Google
Standard
Google
878.0±12Mar 2018
108
Murf AIMurf AI
Murf AI
Falcon (Beta)
Murf AI
870.0±14Nov 2025
109
NoizNoiz
Noiz
Noiz TTS
Noiz
863.0±16Nov 2025
110
MetaVoiceMetaVoice
MetaVoice
MetaVoice v1
MetaVoice
841.0±18Feb 2024
111
AmazonAmazon
Amazon
Polly Standard
Amazon
815.0±16Nov 2016

Artificial Analysis Arena data from Artificial Analysis· updated Aug 14, 2026