Audio & Speech

AI Struggles to Find Specific Sounds in Long Audio, New Test Shows

If you want AI to search a podcast for one sound, it's nowhere near ready yet.

Deep Dive

You can ask an AI to describe what's in a photo, or to tell you what a dog bark sounds like. But can it listen to a two-hour podcast and tell you exactly when every dog bark happens? That skill is called "temporal audio grounding," and a new research paper shows AI is really bad at it.

The team built a test called TAG-Bench with 1,750 audio clips and natural-language questions — for example, "find every time someone sneezes". They gave the test to 21 popular AI audio models. The clips ranged from 7 seconds to 20 minutes, and some sounds appeared multiple times in one clip. The goal was simple: spot every matching moment.

Results were sobering. The best model scored only 31.2 out of 100 on the main accuracy measure, and even it basically failed on longer recordings. On the strictest check, it correctly found only about 1 out of every 5 matching sounds. Nine models scored below 5 out of 100. Even more striking: every single model undercounted repeats. If a sound happened five times, AI tended to report two or three — but never all five.

Why does this matter? Tools that could search audio for specific moments would save journalists, researchers, and everyday users hours of scrubbing through recordings. Think of searching meeting notes for when someone mentioned a deadline, or finding every moment a baby cries in a night camera feed. This benchmark reveals current AI is only good at roughly locating sounds in short clips, not reliably finding every occurrence in long, messy audio. The researchers are releasing their data and code so other teams can improve, but for now, don't ask your AI assistant to find that one funny noise in last week's call — you'll be waiting a while.

Key Points
  • 21 AI audio models were tested on a new benchmark called TAG-Bench to see how well they can pinpoint when specific sounds occur.
  • Even the best model only found about 31% of the right audio segments, and all models badly undercounted repeated sounds.
  • The result shows AI can't yet reliably search long recordings for particular events, which limits future features like audio search or meeting summaries.

Why It Matters

AI may describe what it hears, but it can't yet tell you when it happened - so audio search tools are still far from dependable.

📬 Get the top 10 AI stories daily