MiniMax Music 3.0
MiniMax Music 3.0 is a music generation model developed by MiniMax that produces complete, full-length songs from creative prompts and optional lyrics in a single generation. Designed to handle complex audio structures and extended durations of up to five minutes, the model focuses on maintaining consistent vocal identity, clear instrumentation, and evolving musical progression throughout an entire track.
Architecture
The system utilizes a hierarchical hybrid language model architecture combined with continuous audio synthesis. It pairs an 8B global LLM, which manages long-range musical structure and global continuity, with a 0.6B local LLM responsible for frame-level acoustic details. Audio representation is handled through an eight-layer residual vector quantization (RVQ) tokenizer, separating core semantic structure from acoustic residuals, while a synthesis stack built on flow matching and a Flow-VAE reconstructs 32 kHz, 16-bit stereo WAV audio.
Capabilities and Prompt Guide
MiniMax Music 3.0 accepts two primary inputs: a structured natural-language caption describing the genre, tempo, instrumentation, and vocal characteristics, and optional lyrics marked with structural section tags such as [Verse] and [Chorus]. Creators can guide the generation by specifying details like performance techniques, regional genres, emotional progression, and production profiles. The model supports both vocal tracks and a dedicated instrumental mode, allowing for flexible production workflows across various creative applications.
Run MiniMax Music 3.0 in Crafiq
Ready to use in the studio. No API keys, no setup.
Open the studioHow MiniMax Music 3.0 ranks
MiniMax Music 3.0 is highlighted in the table below. Switch the metric to see how the ordering changes.