Audio & Speech

BiTSE framework brings speaker isolation to noisy AR glasses settings

Uses direction-of-arrival and voice activity to extract one voice from a crowded room

Deep Dive

A new paper from researchers Selani A. Indrapala and Wageesha N. Manamperi presents BiTSE, a binaural target speaker extraction framework designed specifically for AR glass microphone arrays. The problem: in noisy multi-talker scenarios, ordinary beamforming and denoising models struggle to isolate a single desired speaker. BiTSE tackles this by using both spatial and temporal cues—specifically the direction-of-arrival (DoA) of the target speaker and their voice activity information—to guide the extraction process. Built on a binaural signal denoising architecture, the model integrates three key enhancements: a DoA-aware attention mechanism that uses cyclic positional embeddings to encode speaker direction, a timestamp-based masking strategy that suppresses non-target speech segments using speaker activity, and a novel two-stage loss optimization that first trains for robust denoising, then fine-tunes for improved perceptual audio quality.

Evaluated on the SPeech Enhancement for Augmented Reality (SPEAR) challenge dataset, BiTSE consistently outperforms conventional approaches, delivering enhanced signal fidelity and perceptual quality. This is significant for AR glasses, where arrays of tiny microphones must pick out a single voice amid overlapping conversations, background noise, and reverberation. By combining spatial attention with voice activity gating, BiTSE achieves more accurate target speaker isolation than systems relying on spatial filtering alone. The two-stage training strategy also helps balance noise suppression with natural sound, a key tradeoff in hearing-critical AR applications. The paper, accepted at APSIPA ASC 2026, offers a practical blueprint for improving voice interaction and communication in wearable AR devices.

Key Points
  • DoA-aware attention with cyclic positional embeddings encodes target speaker direction for spatial filtering
  • Timestamp-based masking uses voice activity detection to suppress non-target speech segments
  • Two-stage loss optimization boosts denoising robustness then fine-tunes perceptual quality; tested on SPEAR dataset

Why It Matters

BiTSE makes AR glasses practical in noisy crowds by isolating one speaker, improving voice commands and assisted hearing.

📬 Get the top 10 AI stories daily