Audio & Speech

Aurchestra lets you remix real-world audio like a sound engineer

University of Washington's new system separates up to 5 overlapping sounds in real time on earbuds

Deep Dive

University of Washington researchers (Seunghyun Oh, Malek Itani, Aseem Gauri, Shyamnath Gollakota) have unveiled Aurchestra, a system that transforms hearables from blunt noise-cancellers into programmable audio mixers. Current hearing aids and earbuds offer only global noise suppression or single-target focus, but real-world soundscapes contain multiple simultaneous sources (e.g., a person talking, a dog barking, traffic). Aurchestra’s key innovation is a real-time, on-device multi-output extraction network that generates separate audio streams for up to five overlapping sound classes. It runs on compute-limited hardware using 6 ms streaming chunks, and a dynamic interface only displays active sound categories, reducing cognitive load. In tests across unseen indoor and outdoor environments, the system achieved substantial improvements in target-class enhancement and interference suppression compared to baselines.

The system's architecture is optimized for multiple low-power platforms, making it deployable on current-generation earbuds and smart glasses. Users can independently adjust per-class volume—like an audio engineer mixing tracks—enabling scenarios such as amplifying a conversation while dimming traffic noise without fully muting it. The paper, published at ACM MobiSys 2026, demonstrates that the world need not be heard as a single stream: Aurchestra makes the soundscape truly programmable. This could revolutionize hearing aids, AR/VR headsets, and smart assistants by providing nuanced control over one's auditory environment, with potential applications in accessibility, immersive experiences, and personalized hearing.

Key Points
  • Dynamic interface surfaces only active sound classes for intuitive control
  • On-device multi-output extraction network handles up to 5 overlapping target sounds in real time
  • Processes 6 ms streaming audio chunks on resource-constrained hearables

Why It Matters

Turns hearing augmentation into a programmable experience, enabling nuanced control over real-world audio environments.

📬 Get the top 10 AI stories daily