Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Speech/Realtime TTS-2 Flash - Research Preview
Inworld logoInworld·Speech Generation

Realtime TTS-2 Flash - Research Preview

View rankingsinworld.ai
AA Arena#6
ReleasedAug 2026

Realtime TTS-2 Flash is a high-speed, low-latency audio synthesis model developed by Inworld AI as part of its second-generation speech model family. Released as a Research Preview, the Flash variant is optimized for high-volume workloads and latency-critical applications such as interactive voice agents and live customer support. It is characterized by its significant reduction in "time to first audio" (TTFA), achieving a server-side median latency of approximately 20–25ms.

Architecturally, the model is built on a SpeechLM framework that supports conversational awareness. Unlike traditional text-to-speech systems that process sentences in isolation, the TTS-2 family can utilize prior audio turns as input context, allowing the model to adapt its delivery based on the user's tone, pacing, and emotional state. While the Flash version is roughly 2.5x faster than the primary TTS-2 model, it achieves this by trading off some advanced natural-language steering capabilities in favor of raw performance and lower character costs.

The model features a catalog of approximately 95 ready-made voices and supports crosslingual delivery across more than 100 languages. This allows a single voice identity to be maintained even when switching languages within a single generation. It also supports instant voice cloning, enabling developers to create unique voices from short reference audio clips. To enhance realism, the model renders non-verbal human sounds via bracketed tags like [laugh], [breathe], [sigh], [cough], and [yawn], which are processed as organic audio cues rather than spoken text.

To optimize output quality, Inworld recommends "writing for the ear" using specific formatting cues. Users can emphasize terms through capitalization (e.g., "We NEED to go") or by wrapping words in single asterisks. Precise timing can be managed through SSML break tags (e.g., <break time="1s" />) and punctuation, with periods and commas directly influencing the length of natural pauses. While the Flash variant ignores complex style-steering prompts like [shouting], it remains highly expressive through its use of conversational context and non-verbal tags.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Realtime TTS-2 Flash - Research Preview ranks

Realtime TTS-2 Flash - Research Preview is highlighted in the table below. Switch the metric to see how the ordering changes.