Audio & Speech

Audio-Language Models Like CLAP Fail at Negation, New Study Finds

New benchmark shows CLAP can't tell 'dog barking' from 'no dog barking.'

Deep Dive

Researchers Chun-Yi Kuan and Hung-yi Lee introduce NegEval-Audio, a framework revealing that audio-language embedding models like CLAP struggle with negation. On AudioCaps and Clotho datasets, negation-type multiple-choice accuracy falls below chance, as models map affirmative and negated captions to nearly identical representations. Even recent multimodal LLM-based embeddings fail. A training-free steering method marginally improves MCQ-Neg but not retrieval, suggesting the need for explicit negation-aware training.

Key Points
  • Audio-language embedding models like CLAP map negated captions to the same representation as affirmative ones, completely ignoring negation.
  • On AudioCaps and Clotho, negation-type multiple-choice accuracy drops below random chance (e.g., well below 50%).
  • A training-free steering method improves MCQ-Neg but fails to fix retrieval, highlighting a need for explicit negation-aware training.

Why It Matters

This blind spot in audio AI could lead to dangerous misunderstandings in safety-critical applications like surveillance or autonomous driving.

📬 Get the top 10 AI stories daily