Vo & Han's new underwater audio dataset boosts zero-shot detection by 42.6%
Over 1,000 labeled segments across 8 acoustic classes tackle data scarcity.
Machine learning for underwater acoustics has long been hampered by a lack of publicly available labeled datasets—unlike air-acoustic domains with large benchmarks. To fill this gap, researchers Vo and Han curated a new dataset from an open-source maritime sound archive, producing over 1,000 labeled audio segments spanning eight biologically and mechanically relevant classes (e.g., ships, marine life). They also established a lightweight Convolutional Neural Network (CNN) baseline and proposed a margin-enhanced loss technique combined with feature alignment to mitigate class confusion from data imbalance and acoustic similarity.
While the baseline hits 96.35% accuracy in-domain, the real breakthrough comes from cross-domain evaluation on the ShipsEar dataset. The proposed feature alignment improves zero-shot ship detection by 42.60%, demonstrating strong robustness under distribution mismatch. The team released a transparent curation pipeline and reproducible benchmark to support future research in domain adaptation, imbalance mitigation, and data-efficient underwater acoustic classification. The work has been accepted to AVSS 2026 in Lecce, Italy.
- New curated dataset: 1,000+ labeled audio segments across 8 acoustic classes (biologically and mechanically relevant).
- Lightweight CNN baseline achieves 96.35% in-domain accuracy; novel margin-enhanced loss with feature alignment delivers 42.60% improvement in zero-shot ship detection on ShipsEar.
- Open-source curation pipeline and benchmark released to advance domain adaptation and data-efficient underwater classification.
Why It Matters
Enables robust AI-driven underwater surveillance with limited data, critical for defense, maritime safety, and ocean monitoring.