Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Speech/Raon SpeechLM
Krafton logoKrafton·Speech GenerationOpen weights

Raon SpeechLM

View rankingsHugging Faceraon.krafton.ai
AA Arena#74
Parameters9B
ReleasedApr 2026

Raon SpeechLM is a series of bilingual large speech-language models developed by KRAFTON AI Research, designed to handle both speech understanding and generation in English and Korean. Launched as part of the Raon AI brand, the model family aims to bridge the gap between text-based large language models (LLMs) and real-time audio interaction, moving beyond traditional cascaded systems that rely on separate speech-to-text and text-to-speech modules.

The core architecture of the standard Raon-Speech 9B model is built upon a Qwen3 backbone, integrated with specialized components including a Mimi codec for audio quantization and an ECAPA-TDNN speaker encoder. The model underwent a staged training process involving speech module alignment, end-to-end pre-training using approximately 1.38 million hours of curated audio-text data, and multi-reward Direct Preference Optimization (DPO) to refine prosody and response quality. This design allows the model to preserve the strong reasoning capabilities of its underlying LLM while natively processing and generating complex audio nuances.

Raon SpeechLM supports a unified multi-task interface that includes speech-to-text (STT), text-to-speech (TTS), and spoken question-answering (TextQA). Key capabilities include speaker voice conditioning, which enables voice cloning via reference audio, and TTS continuation, allowing the model to generate speech that naturally follows the prosody of a provided audio clip. An extended version, known as Raon-SpeechChat, introduces full-duplex capabilities, enabling real-time, bidirectional conversations with features like backchanneling and interruption handling.

In July 2026, KRAFTON expanded the family with A.X K2 Raon-Speech, a 21-billion-parameter variant developed in collaboration with SK Telecom. This larger model was optimized for high-performance benchmarks, ranking at the top of several Korean and English evaluation suites for open-weight speech models. The technology is primarily intended for use in interactive gaming environments, powering autonomous characters capable of natural, intent-aware voice communication.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Raon SpeechLM ranks

Raon SpeechLM is highlighted in the table below. Switch the metric to see how the ordering changes.