Audio & Speech

MPEcho generates accurate cover songs with explicit phoneme control

New AI model reduces lyric errors in cover song generation using phoneme-level conditioning.

Deep Dive

Cover song generation (CSG) aims to recreate a song's melody and lyrics while altering other musical elements. The previous state-of-the-art model, SongEcho, used F0 sequences and voiced/unvoiced (V/UV) tags for conditioning, but implicit linguistic information led to high phoneme error rates (PER), degrading lyric accuracy. Researchers from Taiwan (Wei-Jaw Lee et al.) propose MPEcho, which integrates a phoneme encoder and a length regulator (LR) from singing voice synthesis (SVS) into the SongEcho framework. This provides explicit phoneme-level conditioning and precise temporal boundaries, significantly reducing PER. To support this, they developed Phonsa, a Whisper-based automatic transcription model that delivers high-precision phoneme annotations for singing voices, overcoming the scarcity of high-quality audio-phoneme pairs.

Experimental results validate Phonsa's alignment accuracy and MPEcho's end-to-end CSG performance, showing better preservation of both melody and lyrics compared to prior methods. The paper has been accepted by the 27th International Society for Music Information Retrieval (ISMIR). Audio samples, code, and model weights are publicly available, enabling researchers and developers to experiment with controllable cover song generation. This approach bridges the gap between generative audio and linguistic fidelity, opening doors for more accurate vocal synthesis in music production and AI-assisted cover creation.

Key Points
  • MPEcho adds a phoneme encoder and length regulator to SongEcho, enabling explicit control over lyrics and reducing phoneme error rate.
  • Phonsa, a Whisper-based transcription model, provides high-precision phoneme annotations for singing voices, solving data scarcity issues.
  • Accepted at ISMIR 2026; code and weights are open-sourced for community use.

Why It Matters

Enables precise, controllable cover song generation with accurate lyrics, benefiting music producers and AI researchers.

📬 Get the top 10 AI stories daily