arXiv AI

Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization

arXiv AI
Sep 11

Reliable Near-Field Multi-User Positioning Informed by Two-Stage MUSIC

The paper introduces MUSIC-Net, an end-to-end deep learning framework for near-field multi-user positioning that incorporates a two-stage MUSIC algorithm to isolate line-of-sight signal components and estimate surrogate distances. By embedding these MUSIC-derived objects into training, the method bypasses separate parameter estimation and path/source association, directly recovering user positions even in mixed LoS/NLoS multipath scenarios. Additionally, the authors employ split conformal prediction to provide statistically guaranteed confidence sets for each user’s position, achieving lower mean positioning error and tighter prediction regions compared to existing benchmarks.

By Jiaying Li, Haifeng Wen, Changsheng You, Yuanwei Liu, Hong Xing
arXiv AI
2d ago

Supervising Sound Localization by In-the-wild Egomotion

The paper introduces a method for learning binaural sound localization by using egomotion as a supervisory signal. By tracking how a camera’s direction changes relative to a sound source during a video, the authors train an audio model to predict sound directions that align with visual estimates of camera motion derived from multi‑view geometry. They evaluate this approach on a newly proposed dataset of real‑world audio‑visual videos with egomotion, demonstrating that the model can learn from real data and perform well on sound localization tasks.

By Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens