Sonic Stage uses spatial audio to help blind viewers follow dialogue scenes
New system turns video dialogue into 3D soundscapes for blind audiences
Blind and low-vision (BLV) viewers often miss critical visual information during dialogue-heavy scenes because traditional audio description (AD) avoids overlapping speech. To solve this, researchers from HKUST, Columbia University, and other institutions developed Sonic Stage, a system that automatically generates interactive spatial soundscapes from dialogue videos. Instead of narrating actions, Sonic Stage uses three techniques: spatialized dialogue (placing voices in 3D space to indicate positions), diegetic sounds (e.g., footsteps, object noises) to represent actions, and interactive descriptions that users can trigger for context-specific details. The system works with any video and outputs an immersive auditory experience that allows BLV audiences to intuitively follow character movements and scene layouts.
In a user study with 12 BLV participants, Sonic Stage significantly outperformed traditional AD on measures of video comprehension (understanding what characters did during dialogue), spatial presence (feeling of being in the scene), and narrative engagement. The system was accepted to UIST 2026, a top HCI conference. The researchers highlight opportunities to extend this approach to different video genres (movies, TV shows, educational content) and to integrate with existing accessibility tools. Sonic Stage represents a major leap forward in making visual media truly inclusive for blind audiences.
- Sonic Stage uses three auditory techniques: spatialized dialogue, diegetic sounds, and interactive descriptions to convey actions during dialogue.
- In a study with 12 blind/low-vision viewers, the system significantly improved video comprehension and spatial presence compared to traditional audio description.
- Accepted to UIST 2026, the system addresses a critical gap where audio description cannot describe actions that overlap with speech.
Why It Matters
Makes dialogue-driven videos accessible to blind viewers, improving media inclusivity and narrative engagement.