Melodia v3
Melodia v3 is a latent diffusion model developed by Sonauto for text-to-music generation. Unlike many competing audio generation architectures that rely heavily on autoregressive language models and token sequences, Melodia applies diffusion techniques adapted from image generation directly to audio, starting from pure noise and iteratively denoising it into structured compositions.
The model is capable of generating fully produced songs complete with vocals and instrumentation from plain-text prompts, style tags, or custom lyrics. Melodia v3 supports over thirty vocal languages and provides granular creative features such as section-level inpainting, track extension, and stem separation. Its underlying diffusion mechanics maintain atmospheric continuity across multi-minute audio outputs.
Architecture and Capabilities
Operating through iterative denoising steps, Melodia v3 translates descriptive prompts, tempos, and key signatures into structured audio. The architecture handles complex multi-instrumental arrangements and diverse genre combinations, allowing users to modify specific sections of a track independently while preserving overall coherence.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Melodia v3 ranks
Melodia v3 is highlighted in the table below. Switch the metric to see how the ordering changes.