Robots Atlas>ROBOTS ATLAS
MiniMax-Music3

MiniMax-Music3

3
Music generation model by MiniMax. From lyrics and a style description it produces complete songs up to five minutes long with vocals, as 32 kHz stereo audio.
✓ Active✓ Public access⚖ Open weightsAudio
Context window
5K tokenów (prompt)
tokens
Parameters
8B Global LLM + 0.6B Local LLM (+ 2.4B Flow Matching, 123M Flow-VAE)
parameters
Release date
7 August 2026
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud

Overview

MiniMax Music 3 (MiniMax-Music3) is a music generation model developed by the Chinese company MiniMax. Conditioned on lyrics (optionally with section tags such as [Verse] or [Chorus]) and a style description, it produces complete, structurally coherent songs of up to five minutes with vocals and evolving arrangements. The output is 32 kHz, 16-bit stereo WAV audio.

The model combines a hierarchical autoregressive architecture with a synthesis system based on Flow Matching and Flow-VAE. An 8B Global LLM (initialized from Qwen3-8B) models the song's long-range structure, while a 0.6B Local LLM restores frame-level acoustic detail. The synthesis module (2.4B Flow Matching and a 123M Flow-VAE decoder) generates a continuous audio representation. The training tokenizer uses eight layers of Residual Vector Quantization (RVQ).

The model weights are publicly available on Hugging Face under the MiniMax-Music3 Community License. Inference requires an NVIDIA GPU (CUDA) and is supported by SGLang-Omni, diffusers and ComfyUI, among others. Only non-streaming generation is supported and the tokenized text prompt is limited to 5,000 tokens.

Classification
Audio
Applications
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open weights
Key parameters
📏 Context: 5K tokenów (prompt)
🧩 Parameters: 8B Global LLM + 0.6B Local LLM (+ 2.4B Flow Matching, 123M Flow-VAE)
📥 Input: text

Technical specification

Context window
5K tokenów (prompt)
tokens
Parameters
8B Global LLM + 0.6B Local LLM (+ 2.4B Flow Matching, 123M Flow-VAE)
parameters
License
MiniMax-Music3 Community License
Hardware requirements
Inference requires an NVIDIA GPU (CUDA). Supported via SGLang-Omni/SGLang, diffusers and ComfyUI. Non-streaming generation only.
Modalities
⬇ Input
text
⬆ Output
audio

Capabilities and applications

Native model capabilities
Audio generation
Generating audio, including audio synchronized with video.
Category: audio
Music generation
Generating complete music tracks (with vocals and arrangement) from lyrics and a style description.
Category: audio
Application domains

Technical architecture