arXiv:2607. 15265v1 Announce Type: cross Abstract: We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language.
By Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman
arXiv:2606. 02724v1 Announce Type: cross Abstract: Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric scene understanding.
By Yaoting Wang, Yun Zhou, Zipei Zhang, Henghui Ding
arXiv:2607. 24786v1 Announce Type: cross Abstract: Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale.
By Hugo Malard, Michel Olvera, Sanjeel Parekh, Ga\"el Richard, Slim Essid, St\'ephane Lathuili\`ere
arXiv:2506.20756v4 Announce Type: replace
Abstract: Recent video depth estimation methods achieve great performance by following the paradigm of image depth estimation, i.e., typically fine-tuning pr...
By Haodong Li, Chen Wang, Jiahui Lei, Kostas Daniilidis, Lingjie Liu
The paper introduces ACF-Net, an optical flow‑guided framework for asymmetric audio‑visual fine‑grained visual categorization (FGVC), addressing challenges where video and audio are not strictly synchronized or matched. ACF-Net comprises Optical Flow‑Guided Motion (OFGM) to capture motion‑sensitive visual cues and suppress background noise, and Asymmetric Cross‑Modal Adaptive Fusion (ACAF) to estimate modality reliability and perform uncertainty‑aware fusion. The authors also present BirdPro, a new bird‑oriented audio‑visual benchmark with 1,919 audio recordings and 11,965 videos across 194 species, and report that ACF‑Net outperforms baselines by 2.97% in fused and 1.92% in mismatched settings.
By Bohan Deng, Shuo Ye, Zitong Yu
arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.
By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie