Audio & Speech

New Tech Scrambles Your Voice So It Can't Be Identified

Your voice is like a fingerprint — this tool hides it without garbling your words.

Deep Dive

Your voice is one of the most personal things about you — and one of the easiest to track. Banks, call centers, smart speakers and even scammers can identify you from just a few seconds of audio. A team of researchers has now built a system called Decaf that removes that identifying quality before your voice ever leaves your device. The name is a deliberate nod to decaffeinated coffee: they take out the part that gives you away, but leave the flavour of your words intact.

Here's how it works in plain terms. Rather than sending your actual voice, the software boils your speech down to just the words, strips the recording of any hint of who is speaking, and ships that lean package across the internet. On the other end, a shared, generic "stock voice" is used to rebuild the sentence out loud. Everyone using the system sounds the same, but they can still be perfectly understood.

The results are striking. The compressed audio needs only 0.5 kilobits per second — roughly a hundredth of what a typical phone call uses — so it works on flaky mobile signals and cheap devices. When security systems tried to figure out who was speaking, they failed up to 43.5% of the time, meaning identification was effectively broken. Meanwhile, automated transcription got 33.2% more accurate than a leading alternative, so nothing is lost in comprehension.

The catch: Decaf hides who you are, not what you say. If you speak your name, address or card number aloud, those words come through clearly and still need protecting. And both sides of the conversation must agree on the same replacement voice in advance — which limits it to apps and services that opt in, rather than any call you make today.

Key Points
  • Decaf removes your voice's identity but keeps every word clear, so you can be understood without being recognised.
  • It runs on 0.5 kilobits per second — about a hundredth of a normal phone call — so it works on weak signals and cheap hardware.
  • Voice-recognition systems failed up to 43.5% of the time trying to identify speakers, while transcription accuracy actually improved by 33.2%.

Why It Matters

Could let you use voice assistants and helplines without your voiceprint being stored, sold or stolen.

📬 Get the top 10 AI stories daily