Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Speech/Realtime TTS-2
Inworld logoInworld·Speech Generation

Realtime TTS-2

View rankingsinworld.ai
AA Arena#2
ReleasedAug 2026

Realtime TTS-2 is a specialized speech synthesis model developed by Inworld, designed for high-fidelity, conversational text-to-speech (TTS) applications. Unlike traditional one-shot synthesis engines, this model is built with conversational context-awareness, meaning it can analyze the audio of previous turns in an exchange to adapt its emotional state, pacing, and tone to match the ongoing interaction. It is optimized for sub-200ms latency to enable natural turn-taking in real-time environments.

A primary feature of the model is its support for natural-language steering. Developers can influence the delivery of speech using bracketed instructions within the text, such as [whisper], [say excitedly], or [calm and reassuring]. This allows for precise control over the prosody and emotional nuances of the output without requiring complex parameter tuning. Additionally, the model supports non-verbal vocalizations including [laugh], [sigh], [breathe], [cough], and [yawn], which are rendered as realistic human sounds rather than spoken text.

The model is highly multilingual, supporting over 100 languages and locales. It utilizes a cross-lingual architecture that allows a single voice identity to remain consistent across different languages, facilitating the creation of localized characters that retain their unique vocal personality. Realtime TTS-2 also includes capabilities for high-quality instant voice cloning, enabling the creation of custom voices from brief reference audio samples or written descriptions.

For optimal performance, the model supports persistent steering, where an instruction like [shouting] remains active until a [reset] tag is encountered or a new instruction is provided. It provides granular controls for temperature (to adjust expressiveness), speaking rate, and text normalization to handle abbreviations and dates according to the application's needs.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Realtime TTS-2 ranks

Realtime TTS-2 is highlighted in the table below. Switch the metric to see how the ordering changes.