Raon SpeechLM
Raon SpeechLM is a series of bilingual large speech-language models developed by KRAFTON AI Research, designed to handle both speech understanding and generation in English and Korean. Launched as part of the Raon AI brand, the model family aims to bridge the gap between text-based large language models (LLMs) and real-time audio interaction, moving beyond traditional cascaded systems that rely on separate speech-to-text and text-to-speech modules.
The core architecture of the standard Raon-Speech 9B model is built upon a Qwen3 backbone, integrated with specialized components including a Mimi codec for audio quantization and an ECAPA-TDNN speaker encoder. The model underwent a staged training process involving speech module alignment, end-to-end pre-training using approximately 1.38 million hours of curated audio-text data, and multi-reward Direct Preference Optimization (DPO) to refine prosody and response quality. This design allows the model to preserve the strong reasoning capabilities of its underlying LLM while natively processing and generating complex audio nuances.
Raon SpeechLM supports a unified multi-task interface that includes speech-to-text (STT), text-to-speech (TTS), and spoken question-answering (TextQA). Key capabilities include speaker voice conditioning, which enables voice cloning via reference audio, and TTS continuation, allowing the model to generate speech that naturally follows the prosody of a provided audio clip. An extended version, known as Raon-SpeechChat, introduces full-duplex capabilities, enabling real-time, bidirectional conversations with features like backchanneling and interruption handling.
In July 2026, KRAFTON expanded the family with A.X K2 Raon-Speech, a 21-billion-parameter variant developed in collaboration with SK Telecom. This larger model was optimized for high-performance benchmarks, ranking at the top of several Korean and English evaluation suites for open-weight speech models. The technology is primarily intended for use in interactive gaming environments, powering autonomous characters capable of natural, intent-aware voice communication.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Raon SpeechLM ranks
Raon SpeechLM is highlighted in the table below. Switch the metric to see how the ordering changes.