Research & Papers

PD-GS uses phonemes to fix 'leaky mouth' in AI talking heads

New 3DGS method cuts lip-sync errors to LMD 2.66 on HDTF benchmark.

Deep Dive

3D Gaussian Splatting (3DGS) has enabled fast, photorealistic talking-head rendering, but accurate lip articulation remains a stubborn challenge. Standard regression objectives from continuous acoustic embeddings tend to average mouth configurations, causing over-smoothed motion and violating hard articulatory constraints like bilabial closures—the notorious 'leaky mouth' artifact. To fix this, Ao Fu and Yi Zhou developed PD-GS, which augments a 3DGS talker with time-aligned phoneme tokens extracted from an automatic ASR and forced-alignment pipeline.

At the core is a Linguistic Fusion Module (LFM) that adaptively blends continuous audio context with discrete phoneme embeddings through a learned gate. This lets the model preserve natural audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments like stops and bilabials. Trained purely from monocular video using image reconstruction and lip landmark supervision, PD-GS achieves the best lip geometry among baselines on HDTF (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences. Accepted at ACM MM 2026, PD-GS points to a more linguistically faithful generation of digital humans for film, gaming, and real-time avatars.

Key Points
  • PD-GS combines 3D Gaussian Splatting with ASR-derived phoneme tokens and forced alignment for explicit lip guidance.
  • A learned Linguistic Fusion Module (LFM) gates between continuous audio and discrete phoneme embeddings, reducing averaged mouth poses on articulation-critical frames.
  • Sets a new lip geometry milestone on HDTF with LMD 2.66, cutting 'leaky mouth' closure violations in complex phoneme sequences.

Why It Matters

Better lip articulation brings digital avatars closer to real human speech, improving dubbing, virtual assistants, and metaverse presence.

📬 Get the top 10 AI stories daily