The paper introduces CTAN, a Cycle-Temporal Attention Network for audio‑visual embodied navigation. It proposes an Audio‑Visual Reconstruction Cross‑Attention module that uses bidirectional cycle‑consistency to strengthen spatial semantics across visual and acoustic modalities, and a Temporal Cross‑Modal Memory to fuse real‑time multimodal features with historical context. Experiments on Replica and Matterport3D show that CTAN outperforms prior methods in success rate, SPL, and scene navigation accuracy.
By Teng Liu, Yinfeng Yu
arXiv:2609.23407v1 Announce Type: cross
Abstract: Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for em...
By Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong
arXiv:2609.23407v2 Announce Type: replace-cross
Abstract: Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challengin...
By Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong
RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.
By Xingyi He, Ziwei Wang, Dongrui Wu
BinauralVAE is an open‑source pipeline that reconstructs spatial audio using various Variational Autoencoder architectures, including complex‑valued variants, to learn latent representations of binaural signals. The project builds on realistic acoustic data from a simulated robot navigating an environment, providing a foundation for audio‑centric world models. It aims to map the causal link between navigational actions and their acoustic outcomes, positioning sound as a complementary modality for spatial awareness.
By Luis Vitor Zerkowski, Luiz Velho
arXiv:2604. 03329v2 Announce Type: replace-cross Abstract: Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible.
By Damith Chamalke Senadeera, Dimitrios Kollias, Gregory Slabaugh