Audio & Speech

New AI framework generates full songs from lyrics with flow-matching rendering

Generate complete songs from just text descriptions and lyrics using hierarchical planning.

Deep Dive

The paper introduces a comprehensive framework for full-song generation that combines hierarchical autoregressive planning with flow-matching rendering. Built around four core components—a semantic-aware tokenizer, a hybrid language model (hybird-LM), FullDiT, and a two-level melody module—the system can handle three tasks: lyrics-to-song generation, instrumental music generation, and cover song reinterpretation. The tokenizer encodes audio into 8-codebook residual vector quantized (RVQ) tokens, enabling efficient discrete representation. The hybird-LM then performs hierarchical autoregressive modeling on these tokens for full-song structure planning. To enhance audio fidelity, FullDiT applies a flow-matching process in a continuous VAE latent space, conditioned on codec tokens, lyrics, and text captions. For cover songs, the melody module extracts and discretizes melody cues from a reference audio track to preserve the original melodic content while allowing style changes.

To further improve generation quality, the researchers explore reward-based post-training strategies including DPO, GRPO, and OPD for hybird-LM, and flow-based GRPO for FullDiT. These techniques are applied to optimize both musicality and rendering quality. Experimental results on a multilingual automatic benchmark, as well as the Artificial Analysis Music with Vocals leaderboard, show the framework achieves competitive performance. The paper is published on arXiv under the subject Sound and Artificial Intelligence, authored by Junyu Dai, Xinyue Fan, and 14 others. This work pushes the frontier of AI-generated music by enabling full-length, high-fidelity song generation from simple inputs like lyrics and text descriptions.

Key Points
  • Uses 8-codebook RVQ tokens for efficient discrete music representation and hierarchical autoregressive planning.
  • Employs FullDiT with flow-matching in a VAE latent space conditioned on codec tokens, lyrics, and text captions.
  • Supports three tasks: lyrics-to-song, instrumental music, and cover song generation with melody preservation.

Why It Matters

Enables professional-grade full-song generation from lyrics and text, transforming music production workflows.

📬 Get the top 10 AI stories daily