Audio & Speech

VISA system tops audio reasoning leaderboard using visual cues

Ranks 2nd overall with 66.23% Rubrics score and highest accuracy at 77.40%.

Deep Dive

Audio reasoning—the ability to perform multi-step, evidence-grounded inference on temporally dynamic and mixed acoustic signals—remains a frontier beyond traditional tasks like ASR or captioning. To address this, a team of researchers from multiple institutions (including authors Wenming Tu, Jian Gao, and others) presents VISA, their submission to the Interspeech 2026 Audio Reasoning Challenge (ARC) Agent Track. VISA operates under a 'LALM as a Tool' paradigm, meaning it strengthens large audio language models with auxiliary multi-modal evidence without heavy orchestration. The system comprises three key components: multi-modal feature extraction that captures complementary audio and acoustic-visual clues; model-voting inference with consistency checking for stable predictions; and fine-grained category-aware routing to resolve disagreements and select reasoning chains aligned with evaluation rubrics.

On the official Agent Track leaderboard, VISA ranks 2nd overall with a 66.23% Rubrics score—a metric measuring correctness and reasoning quality per the MMAR Rubrics. Notably, it also achieves 77.40% Accuracy, the highest among all systems listed across both the Single Model and Agent tracks. This demonstrates that incorporating visual information can significantly boost audio reasoning performance, especially in complex, real-world scenarios where sound sources have visual correlates. The work was submitted to INTERSPEECH 2026 and is available on arXiv (2606.07264). VISA’s success suggests a promising direction for building more robust, multi-modal AI agents capable of sophisticated audio understanding.

Key Points
  • Ranks 2nd overall on the Interspeech 2026 ARC Agent Track leaderboard with a 66.23% Rubrics score.
  • Achieves 77.40% Accuracy—the highest across both Single Model and Agent tracks in the challenge.
  • Uses a three-component architecture: multi-modal feature extraction, model-voting with consistency checking, and category-aware routing.

Why It Matters

Combining visual cues with audio reasoning pushes AI toward human-like, multi-step inference in noisy environments.

📬 Get the top 10 AI stories daily