The paper introduces an audio-first triage method for budgeted vision‑language captioning of untrimmed egocentric video. By selecting windows for a vision‑language model based on lightweight audio features before any video decoding, the approach reduces costly model calls. It achieves 4.0–10.8 percentage point improvements in action coverage across call rates, cuts 9–20% of VLM calls at matched coverage on EPIC‑KITCHENS‑100, and outperforms uniform sampling and recent visual keyframe selectors on Ego4D.
By Masoud Jalayer, Changyi Li, Yu Xiao
arXiv:2607. 11798v1 Announce Type: cross Abstract: Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film.
By Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin
arXiv:2608. 14016v1 Announce Type: cross Abstract: Live game commentary is scarce: it exists for professional esports broadcasts and almost nowhere else.
By Mathew Varghese
arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.
By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie
arXiv:2606. 02724v1 Announce Type: cross Abstract: Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric scene understanding.
By Yaoting Wang, Yun Zhou, Zipei Zhang, Henghui Ding
arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.
By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
arXiv:2605. 00873v2 Announce Type: replace-cross Abstract: The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods.
By Advait Tilak, Jiwon Choi, Nazifa Mouli, Wei Le
arXiv:2606. 02642v1 Announce Type: cross Abstract: Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination.
By Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu, Tae-Hyun Oh
arXiv:2605.07593v2 Announce Type: replace
Abstract: Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory str...
By Hengyi Feng, Hao Liang, Mingrui Chen, Bohan Zeng, Meiyi Qiang, Zhengyang Zhao, Zimo Meng, Zeang Sheng, Wentao Zhang
arXiv:2608.28699v1 Announce Type: new
Abstract: Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine...
By Dong-Hee Kim, Seonwoo Choi, Changbeen Kim, Jungmyung Wi, Juyeon Ko, Youngju Choi, Il Hyeon Mun, Hyunwoo J. Kim, Donghyun Kim
Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations.
arXiv:2607. 03050v1 Announce Type: cross Abstract: Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-visual inputs, leading to substantial inference cost.
By Shijie Cao, Qingyu Zhang, Boxi Yu, Yuzhong Zhang, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun