arXiv:2608.09200v3 Announce Type: replace
Abstract: Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfo...
By Lifang Wu, Yuyang Wu, Yangdong Gao, Fengyu Liu, Ya Jing, Liang Wang
arXiv:2607. 21267v1 Announce Type: new Abstract: Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears.
By Yu Zhang, Jiayuan Rao, Haoning Wu, Weidi Xie
arXiv:2608. 07932v2 Announce Type: replace Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement.
By Yizhi Li, Jiawei Jiang, Guanhong Wang, Yingcai Wu, Gaoang Wang
arXiv:2606. 04806v1 Announce Type: cross Abstract: LLMs and agentic systems are increasingly deployed in social environments, making normative competence critical for safe and appropriate behavior.
By Sichao Li, Sai Ma, Daniel Kilov, Secil Yanik Guyot, Zhuang Li, Seth Lazar
arXiv:2608. 19646v1 Announce Type: new Abstract: Visual understanding in sports has emerged as a hot topic in computer vision in recent years.
By Yunhao Zhao, Haoying Sun, Jiarui Li, Zhuming Wang, Ya Jing, Xiangbo Shu, Lifang Wu, Changwen Chen
arXiv:2607. 02588v1 Announce Type: cross Abstract: Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory.
By Yixin Ji, Fanghua Ye, Juntao Li, Bo Zhao, Zexuan Qiu, Zhaopeng Tu, Liefeng Bo, Min Zhang
arXiv:2608. 19723v1 Announce Type: cross Abstract: Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory.
By Chenxi Shao, Bozhong Wang, Jiaxin Huang, Zhao Liu, Sunwei Zhu, Tianxin Hang, Gaoqi He, Yang Li, Changbo Wang
arXiv:2609.13258v1 Announce Type: cross
Abstract: We present a structured temporal video reasoning pipeline built around a discrete EventGraph, a continuous EventField, and a human-readable EventGlyp...
By Durgendra Narayan Singh
arXiv:2605. 21917v2 Announce Type: replace-cross Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support.
By Han Zhang, Wanting Jiang, Tomasz Kornuta, Tian Zheng, Vidya Murali
WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.
By Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
arXiv:2601. 18157v3 Announce Type: replace-cross Abstract: The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous, longitudinal stream of egocentric video.
By Aniket Rege, Arka Sadhu, Yuliang Li, Kejie Li, Ramya Korlakai Vinayak, Yuning Chai, Yong Jae Lee, Hyo Jin Kim
SocialReasonBench is a new video‑multiple‑choice QA benchmark designed to test socially grounded reasoning in interactive narrative videos. It uses branching gameplay footage from *Detroit: Become Human*, where player choices create alternative social outcomes that can be verified against the game’s script and flowchart. The benchmark includes seven reasoning dimensions—such as intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent—and employs a multi‑agent pipeline to curate clips, ground answer labels, and generate theory‑guided questions with diagnostic distractors.
By Zheyu Huang, Zijing Shi, Haozhe Luo, Huadong Tang, Mingyu Liu, Meng Fang, Ling Chen