arXiv:2606. 07033v1 Announce Type: new Abstract: Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training.
By Zhe Yang, Ruyi Zhang, Hongtao Chen, Wenrui Li, Hengyu Man, Wangmeng Zuo, Xiaopeng Fan
AVTrace is a diagnostic suite designed to evaluate audio‑visual temporal reasoning in omni models. It covers tasks such as onset and span grounding, synchronization, next‑step prediction, cross‑modal localization, chain parsing, and event‑conditioned comprehension, providing 34,114 training examples and balanced development and test splits. Five open omni models were tested, all scoring below the majority‑label baseline on synchronization verification and showing low performance on chain parsing and event‑conditioned tasks, while parameter‑efficient temporal post‑training improved some metrics.
By Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw
OP-CAD introduces a curriculum-based, on-policy clean-audio distillation framework that enhances audio-visual reasoning under environmental noise and competing speech. The method trains a student model from mild to severe noise, using a frozen teacher that provides token-level supervision based on clean audio and verified answers, while selectively weighting positions sensitive to acoustic interference. Experiments show OP‑CAD outperforms existing methods across all noise conditions, preserving clean‑correct answers without sacrificing overall accuracy.
By Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities inde...
OmniReasoning introduces a new benchmark, OmniReasoningBench, that requires both audio and visual evidence for answering 1,150 multiple-choice and open-ended questions across two tasks. The authors also develop OmniQA, a data engine that automatically generates evidence‑grounded QA pairs with time‑stamped clue chains, producing training datasets OmniReasoning‑SFT‑112K and OmniReasoning‑RL‑19K. Finally, they propose Modality‑Factored Self‑Distillation (MFSD), an on‑policy self‑distillation method that assigns token‑level credit by evaluating responses under modality‑specific clue contexts, enabling the OmniReasoning‑30B‑A3B model to achieve significant performance gains on both the new benchmark and existing video benchmarks.
By Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang, Minghao Han, Yunfei Chu, Shun Lei, Xueyao Zhang, Qize Yang, Jin Xu, Yiwu Zhong
arXiv:2608. 09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time.
By Zhi Zeng, Cheng Zhang, Zesheng Yang, Rendong Pi, Jiaying Wu, Di Zhang, Zihan Ma, Guodong Li, Zhou Yang, Yu Xiang, Yifei Zheng, Minnan Luo
LEAP is a framework for long audio‑video question answering that avoids encoding entire recordings by dividing them into fixed‑duration blocks. It performs a lightweight localization pass on each block to score short candidate windows, then pools the highest‑ranked windows for a single bounded answer pass, keeping the answer input and peak context independent of recording length. The method trains both a localization LoRA and an answer LoRA, supports causal streaming queries, and achieves significant performance gains over baseline models on multiple AVQA benchmarks.
By Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, Liang Peng
The paper introduces VISTA, a baseline for dense multi‑label classroom coding that leverages the COPUS protocol as a video benchmark. VISTA applies MiniCPM‑V‑4.5 over sliding windows, refines predictions with an MLP head, and aggregates results onto a 2‑minute COPUS grid, achieving 80.1% macro accuracy on held‑out chemistry lectures. The authors also identify systematic failure modes and provide benchmark tooling and code on GitHub.
By Andrew Franck, Brendan Ng, Ben Fitzgerald, Zane Derrod, Chris Cianci, Chris Craney
arXiv:2605.07593v2 Announce Type: replace
Abstract: Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory str...
By Hengyi Feng, Hao Liang, Mingrui Chen, Bohan Zeng, Meiyi Qiang, Zhengyang Zhao, Zimo Meng, Zeang Sheng, Wentao Zhang
The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.
By Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan
arXiv:2608. 09227v1 Announce Type: new Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive.
By Puneet Mathur, Manan Suri, Dinesh Manocha
arXiv:2607. 24786v1 Announce Type: cross Abstract: Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale.
By Hugo Malard, Michel Olvera, Sanjeel Parekh, Ga\"el Richard, Slim Essid, St\'ephane Lathuili\`ere