OmniEcho: Spatial Audio Understanding for Embodied Agents
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.23407v2 Announce Type: replace-cross Abstract: Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challengin...
arXiv:2606. 10738v1 Announce Type: cross Abstract: Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding.
BinauralVAE is an open‑source pipeline that reconstructs spatial audio using various Variational Autoencoder architectures, including complex‑valued variants, to learn latent representations of binaural signals. The project builds on realistic acoustic data from a simulated robot navigating an environment, providing a foundation for audio‑centric world models. It aims to map the causal link between navigational actions and their acoustic outcomes, positioning sound as a complementary modality for spatial awareness.
arXiv:2608. 09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time.
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources.
arXiv:2607. 24786v1 Announce Type: cross Abstract: Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale.