Audio & Speech

Text-Prompted CLAP enhances audio AI with query-conditioned learning

A new fusion module lets audio models understand text prompts better than ever before.

Deep Dive

Researchers Mohan Li, Rama Doddipatla, and Philip C. Woodland have unveiled Text-Prompted CLAP (TP-CLAP), a parameter-efficient upgrade to the Contrastive Language-Audio Pretraining (CLAP) model. While CLAP encodes text and audio independently into a shared embedding space, it struggles with cross-modal semantics for complex understanding and retrieval tasks. TP-CLAP solves this by adding a cross-attention-based fusion module that lets audio features be conditioned on textual prompts. The model is trained using an audio multiple-choice question (audio-MCQ) framework, where it learns to align query-conditioned audio representations with correct answer text embeddings via contrastive learning.

Results show TP-CLAP achieves competitive performance on audio question answering (audio-QA) while being substantially smaller than current large audio-LLMs. It also improves the base CLAP model on audio-text retrieval and zero-shot classification benchmarks. Further fine-tuning for attribute-focused audio-to-audio retrieval shows TP-CLAP consistently outperforms standard CLAP, particularly in music retrieval. The work demonstrates that lightweight prompt conditioning can unlock powerful cross-modal reasoning without scaling model size.

Key Points
  • Adds a cross-attention fusion module to incorporate text prompts into audio features, parameter-efficient.
  • Trained via audio multiple-choice question (audio-MCQ) framework, aligning query-conditioned audio with correct answers.
  • Outperforms standard CLAP on audio-text retrieval, zero-shot classification, and attribute-focused music retrieval.

Why It Matters

This lightweight approach brings powerful text-guided audio understanding closer to production, rivaling much larger models.

📬 Get the top 10 AI stories daily