Hugging Face Trending Papers

NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effectively in dynamic environments. Classical methods such as Multiple Signal Classification (MUSIC) offer strong theoretical foundations but degrade under low signal-to-noise ratios.

arXiv AI
2d ago

Supervising Sound Localization by In-the-wild Egomotion

The paper introduces a method for learning binaural sound localization by using egomotion as a supervisory signal. By tracking how a camera’s direction changes relative to a sound source during a video, the authors train an audio model to predict sound directions that align with visual estimates of camera motion derived from multi‑view geometry. They evaluate this approach on a newly proposed dataset of real‑world audio‑visual videos with egomotion, demonstrating that the model can learn from real data and perform well on sound localization tasks.

By Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens
arXiv Machine Learning
Sep 10

BinauralVAE: Spatial Audio Reconstruction For World Models

BinauralVAE is an open‑source pipeline that reconstructs spatial audio using various Variational Autoencoder architectures, including complex‑valued variants, to learn latent representations of binaural signals. The project builds on realistic acoustic data from a simulated robot navigating an environment, providing a foundation for audio‑centric world models. It aims to map the causal link between navigational actions and their acoustic outcomes, positioning sound as a complementary modality for spatial awareness.

By Luis Vitor Zerkowski, Luiz Velho
arXiv Machine Learning
Aug 18

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

arXiv:2608. 15037v1 Announce Type: cross Abstract: Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference.

By Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi