Stable Audio 3 Medium
Stable Audio 3 is a family of text-to-audio and music generation latent diffusion models developed by Stability AI. Released in May 2026, the model series is designed for variable-length audio generation, supporting tracks up to six minutes and twenty seconds long, as well as specialized sound effects and foley generation. Trained entirely on fully licensed and Creative Commons data, Stable Audio 3 introduces a novel semantic-acoustic autoencoder (SAME) and adversarial post-training to dramatically accelerate inference speeds while preserving high audio fidelity.
The model architecture utilizes a diffusion transformer operating on compact latent spaces, coupled with advanced text encoders. Stable Audio 3 is available in multiple variants, ranging from compact open-weight small and medium checkpoints to larger proprietary variants, with parameter counts spanning roughly 0.6B to 2.7B. These checkpoints cater to diverse deployment needs, enabling efficient local inference on consumer-grade hardware or fast generation via cloud infrastructure.
Key capabilities include full-song structure synthesis, audio inpainting for targeted editing, and LoRA fine-tuning support using custom user samples. The models are capable of producing high-quality stereo audio at 44.1 kHz, offering precise prompt adherence and rich musical arrangements across varied genres, tempos, and soundscapes.
Run Stable Audio 3 Medium in Crafiq
Ready to use in the studio. No API keys, no setup.
Open the studioHow Stable Audio 3 Medium ranks
Stable Audio 3 Medium is highlighted in the table below. Switch the metric to see how the ordering changes.