Audio & Speech

ReGen AI generates studio-quality audio from ultra-compressed 12.5Hz waveforms

Trained on just 4 GPUs in one day, ReGen matches models using 10x more compute.

Deep Dive

A team of researchers led by Sang-Hoon Lee and Ha-Yeong Choi have published ReGen, a novel hierarchical multi-prompt representation generation framework for efficient waveform diffusion models. The key innovation is addressing a problem in representation alignment (REPA): regularizing intermediate representations in diffusion Transformers (DiT) can implicitly entangle latents and limit generative capacity. ReGen jointly estimates multiple vector fields for both representations and data within a single diffusion model, using a new generalized flow matching (GFM) technique to improve conditional flow matching generalization.

ReGen is validated on single-stage waveform diffusion models including neural audio codec and Wave-VAE, significantly improving generation quality from highly compressed latent representations at 12.5 Hz. The team also presents ReGenVoice, a latent diffusion model (LDM)-based text-to-speech system that operates at 6.25 Hz with rich semantic and acoustic latent representation. This enables remarkably efficient training — just 1 day on 4 GPUs — and fast inference with a real-time factor (RTF) of 0.08. Despite the small training footprint, ReGenVoice achieves strong speech intelligibility (WER) and speaker similarity (SIM). The paper has been accepted to ICML 2026, and audio samples are available online.

Key Points
  • ReGen uses hierarchical multi-prompt representation generation and generalized flow matching to avoid representation entanglement in diffusion Transformers.
  • Generates high-quality waveforms from highly compressed latents at 12.5 Hz, improving upon neural audio codec and Wave-VAE baselines.
  • ReGenVoice TTS trains in 1 day on 4 GPUs with an RTF of 0.08, achieving strong WER and speaker similarity on small datasets.

Why It Matters

Enables high-quality audio generation on limited hardware, democratizing TTS and waveform synthesis for startups and researchers.

📬 Get the top 10 AI stories daily