Audio & Speech

Face-to-Speech AI generates a person's voice from just a photo

New research achieves UTMOS scores of 3.7-4.0, matching real speech quality.

Deep Dive

A new paper from Carlos Muñoz-Romero and Jose A. Gonzalez-Lopez introduces a zero-shot Face-to-Speech (F2S) framework that can predict a plausible voice from a single static facial image. This overcomes a key limitation of traditional voice cloning (text-to-speech), which requires a short audio sample. The approach uses a lightweight Face Adapter and soft-tuning of the face encoder's upper blocks to align face-recognition features with the latent style space of a frozen StyleTTS 2 model. The model is trained only on facial images and text, leaving the TTS backbone untouched.

Evaluated on held-out identities from the LRS3 audiovisual corpus (English TED talks), the synthesized speech achieves naturalness scores of UTMOS 3.7-4.0—matching or even exceeding the 3.61 of the original ground-truth recordings. Face-to-voice retrieval also performed consistently above chance, confirming that generated voices are consistent with the target speaker. Remarkably, without any retraining, the English-trained adapter produced fluent Spanish speech, suggesting the face-to-style mapping is largely language-agnostic. This opens up applications for historical figures, video-game characters, or any scenario where only visual data is available.

Key Points
  • UTMOS scores of 3.7-4.0 match or exceed ground-truth voice quality (3.61) on LRS3.
  • Zero-shot voice synthesis from a single photo, no audio sample needed.
  • Cross-lingual transfer works: English-trained adapter generates fluent Spanish without retraining.

Why It Matters

Voice cloning without audio unlocks historical recreations, game characters, and accessibility tools from just photos.

📬 Get the top 10 AI stories daily