AI Voices Just Learned to Laugh and Sigh Like Real People
That fake-sounding AI narrator is about to get a lot more convincing.
Researchers built a way to generate controllable non-verbal vocalizations — the sounds speech models make beyond words — in continuous autoregressive speech models, using preference optimization with feedback from a Large Audio-Language Model. To avoid human preference annotation, they combined NVV-injected real transcripts with LLM-generated prompts, ran stochastic model rollouts, and had a LALM rank candidates into chosen–rejected pairs. Their two-stage strategy pairs Rejection Sampling Fine-Tuning with Anchored Flow-DPO. On the official 1,600-utterance NVVSpeech Challenge Track 2 test set, the method scored 75.80 (79.39 ZH / 72.21 EN), beating the VoxCPM2 baseline by +1.84, driven mainly by higher NVV Accuracy and NVV Perceptual Effect while Overall Quality stayed stable.
- AI speech is getting good at wordless human sounds: sighs, gasps, giggles, and 'mm-hmm' noises that signal you're listening.
- Researchers skipped expensive human ratings by having a second AI rank recordings, then training the voice generator on the winners — scoring 75.80 on a 1,600-clip test, beating the old best by 1.84 points.
- Real-world impact: more lifelike audiobooks, dubbing, and assistants — but also more believable AI voice scams to watch out for.
Why It Matters
More human-sounding AI voices improve assistants, audiobooks, and dubbing — but also make phone scams much harder to detect.