arXiv AI
2d ago

CTAN: Cycle-Temporal Attention Network for Embodied Audio-Visual Navigation

The paper introduces CTAN, a Cycle-Temporal Attention Network for audio‑visual embodied navigation. It proposes an Audio‑Visual Reconstruction Cross‑Attention module that uses bidirectional cycle‑consistency to strengthen spatial semantics across visual and acoustic modalities, and a Temporal Cross‑Modal Memory to fuse real‑time multimodal features with historical context. Experiments on Replica and Matterport3D show that CTAN outperforms prior methods in success rate, SPL, and scene navigation accuracy.

By Teng Liu, Yinfeng Yu
arXiv Machine Learning
Sep 10

BinauralVAE: Spatial Audio Reconstruction For World Models

BinauralVAE is an open‑source pipeline that reconstructs spatial audio using various Variational Autoencoder architectures, including complex‑valued variants, to learn latent representations of binaural signals. The project builds on realistic acoustic data from a simulated robot navigating an environment, providing a foundation for audio‑centric world models. It aims to map the causal link between navigational actions and their acoustic outcomes, positioning sound as a complementary modality for spatial awareness.

By Luis Vitor Zerkowski, Luiz Velho
arXiv AI
Jul 16

A Hybrid Mamba for Audio-Visual Navigation

arXiv:2607. 13110v1 Announce Type: cross Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences.

By Yi Wang, Yinfeng Yu