arXiv:2606. 07033v1 Announce Type: new Abstract: Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training.
By Zhe Yang, Ruyi Zhang, Hongtao Chen, Wenrui Li, Hengyu Man, Wangmeng Zuo, Xiaopeng Fan
The paper investigates audio‑video diffusion models by examining the "attention triangle"—the cross‑attention links among text, audio, and video. It finds that the audio‑video edge is bidirectional and heavily influenced by model biases, leading to semantic leakage when prompts conflict with learned priors. The authors develop attention‑derived diagnostics and inference‑time interventions that improve semantic grounding without sacrificing generation quality.
By Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes
The paper introduces ACF-Net, an optical flow‑guided framework for asymmetric audio‑visual fine‑grained visual categorization (FGVC), addressing challenges where video and audio are not strictly synchronized or matched. ACF-Net comprises Optical Flow‑Guided Motion (OFGM) to capture motion‑sensitive visual cues and suppress background noise, and Asymmetric Cross‑Modal Adaptive Fusion (ACAF) to estimate modality reliability and perform uncertainty‑aware fusion. The authors also present BirdPro, a new bird‑oriented audio‑visual benchmark with 1,919 audio recordings and 11,965 videos across 194 species, and report that ACF‑Net outperforms baselines by 2.97% in fused and 1.92% in mismatched settings.
By Bohan Deng, Shuo Ye, Zitong Yu
arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.
By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie
arXiv:2606. 25225v1 Announce Type: cross Abstract: Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning.
By Revant Teotia, Adrien Bardes, Michael Rabbat, Sumit Chopra, Matthew J. Muckley, Nicolas Ballas
The paper investigates how audio‑video diffusion models use cross‑modal attention, focusing on the "attention triangle" that connects text, audio, and video streams. It finds that the audio‑video edge is bidirectional and heavily influenced by model biases, leading to semantic leakage when prompts conflict with learned priors. By extracting attention signals, the authors develop diagnostic tools and inference‑time interventions that improve cross‑modal alignment without sacrificing generation quality.