arXiv AI By Zhaojin Fu, Yuyang Hong, Qi Yang, Zili Wang, Kun Ding, Shiming Xiang, Bin Fan

Audio-Visual Segmentation via Depth-Guided Collaborative Modeling

Read the original on arXiv AI →

arXiv:2608. 16285v1 Announce Type: cross Abstract: Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 27

Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark

The paper introduces ACF-Net, an optical flow‑guided framework for asymmetric audio‑visual fine‑grained visual categorization (FGVC), addressing challenges where video and audio are not strictly synchronized or matched. ACF-Net comprises Optical Flow‑Guided Motion (OFGM) to capture motion‑sensitive visual cues and suppress background noise, and Asymmetric Cross‑Modal Adaptive Fusion (ACAF) to estimate modality reliability and perform uncertainty‑aware fusion. The authors also present BirdPro, a new bird‑oriented audio‑visual benchmark with 1,919 audio recordings and 11,965 videos across 194 species, and report that ACF‑Net outperforms baselines by 2.97% in fused and 1.92% in mismatched settings.

By Bohan Deng, Shuo Ye, Zitong Yu
arXiv Machine Learning
Jul 14

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.

By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie