arXiv:2608. 19646v1 Announce Type: new Abstract: Visual understanding in sports has emerged as a hot topic in computer vision in recent years.
By Yunhao Zhao, Haoying Sun, Jiarui Li, Zhuming Wang, Ya Jing, Xiangbo Shu, Lifang Wu, Changwen Chen
arXiv:2608.09200v3 Announce Type: replace
Abstract: Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfo...
By Lifang Wu, Yuyang Wu, Yangdong Gao, Fengyu Liu, Ya Jing, Liang Wang
arXiv:2608.23435v1 Announce Type: cross
Abstract: Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge...
By Yirong Hu, Jiayuan Rao, Yu Zhang, Shangzhe Di, Weidi Xie
arXiv:2608. 07932v2 Announce Type: replace Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement.
By Yizhi Li, Jiawei Jiang, Guanhong Wang, Yingcai Wu, Gaoang Wang
arXiv:2609.28049v1 Announce Type: cross
Abstract: Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as ev...
By Sai Varun Kodathala, Prashanth Pollishetty, Jaylen Cargill
Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Ama...
arXiv:2505.01583v2 Announce Type: replace
Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...
By Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou, Vivian Wang, Huayu Wang, Hsiang-Wei Huang, Wenhao Chai, Hou-I Liu, Kuang-Ming Chen, Cheng-Yen Yang, Yi-Ling Chen, Vibhav Vineet, Qin Cai, Jenq-Neng Hwang
arXiv:2606. 09181v1 Announce Type: cross Abstract: Recent advances in video multimodal models have significantly improved VideoQA performance.
By Zhou Du, Hamid Krim, Xiao Wu, Zhaoquan Yuan, Liangwei Li, Keisuke Fujii
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.
By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
EventVL introduces the first generative event-based multimodal large language model (MLLM) designed for explicit semantic understanding of event streams. The framework leverages a newly annotated dataset of nearly 1.4 million event–image/video–text pairs and incorporates an Event Spatiotemporal Representation to capture comprehensive event information, along with Dynamic Semantic Alignment to refine sparse semantic spaces. Experiments demonstrate that EventVL outperforms existing MLLM baselines in event captioning and scene description generation tasks, advancing the field of event vision.
By Pengteng Li, Yunfan Lu, Pinghao Song, Wuyang Li, Huizai Yao, Hui Xiong
TimeBlind is a diagnostic benchmark designed to evaluate fine‑grained spatio‑temporal compositionality in video large language models (LLMs). It categorizes temporal understanding into three levels—atomic event recognition, event property characterization, and reasoning about event interdependencies—and uses a minimal‑pairs paradigm where video pairs share identical static content but differ only in temporal structure. Across 20 state‑of‑the‑art MLLMs tested on 600 curated instances, the best model achieved only 48.2% instance accuracy, far below human performance of 98.2%, highlighting a reliance on static visual shortcuts rather than true temporal reasoning.
By Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius
Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame selection, but these strategies usually optimize either broad temporal coverage or local relevance, making it difficult to preserve both global storyline context and fine-grained evidence.