arXiv:2607. 24560v1 Announce Type: cross Abstract: We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D.
By Jinjie Mai, Gordon Guocheng Qian, Willi Menapace, Arpit Sahni, Chaoyang Wang, Ashkan Mirzaei, Runjia Li, Sergey Tulyakov, Bernard Ghanem, Peter Wonka, Rameen Abdal
arXiv:2610.01039v1 Announce Type: new
Abstract: While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting st...
By Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim
arXiv:2609.38839v1 Announce Type: new
Abstract: Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retainin...
By Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan
We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form "when X happens, do Y," EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation.
LiveProBench evaluates streaming video models on their ability to interact proactively, assessing whether they respond at appropriate times without explicit cues. The benchmark tests models at one‑second intervals across six subtasks that vary trigger ambiguity and timing tolerance, measuring response accuracy, silence rates, and duplicate responses. Results show that many models issue premature responses more often than missed ones, highlighting a significant shortfall in human‑like temporal decision making.
By Kaixuan Du, Xin Wan, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, YuKun Wang
ProactiveBench evaluates streaming video models on their ability to interact proactively, rather than reactively. It tests models at one‑second intervals without explicit cues, using six subtasks that vary trigger ambiguity, timing tolerance, and response patterns. The benchmark measures both response and silence rates, distinguishing early, in‑window, and missed responses, and penalizes omissions and repetitions.
By Kaixuan Du, Xin Wan, YuKun Wang, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, Ni Li
arXiv:2606. 18586v1 Announce Type: cross Abstract: Physical events are not understood by their names alone, but by the causal state changes that compose them.
By Shang Wu, Haoran Lu, Songling Liu, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Manling Li, Zhaoran Wang, Han Liu
The paper introduces VT-Contrast, a representation-level temporal counterfactual objective designed to improve temporal understanding in Video Language Models (VideoLMs). By supervising late-layer last-frame video tokens and contrasting order-preserving views with reordered counterfactuals graded by Kendall tau distance, VT-Contrast addresses the mismatch between ordered video input and text-based supervision. The method requires no architectural changes, is compatible with various VideoLM training tasks, and demonstrates improved performance on temporal understanding benchmarks.
By Yumeng Shi, Quanyu Long, Yin Wu, Wenya Wang
arXiv:2602.08277v3 Announce Type: replace-cross
Abstract: The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engi...
By Xiangbo Gao, Renjie Li, Xinghao Chen, Yuheng Wu, Suofei Feng, Qing Yin, Zhengzhong Tu
arXiv:2609.40219v1 Announce Type: cross
Abstract: World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experienc...
By Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan
OmniAssistBench is a new benchmark for evaluating omni-modal large language models (Omni-LLMs) as real‑time video assistants that actively guide users toward goals. The benchmark addresses the challenge of dynamic interaction paths by providing models with predefined priors from source videos, forcing them to follow the same routes as users. The dataset was constructed by reverse‑engineering existing Internet videos into multi‑turn clips, a process that required over 1,000 expert person‑hours. Results show that proprietary Gemini‑3‑Pro scores 66.4/100 while open‑source Qwen3‑Omni‑Instruct scores 51.2, revealing that current models often give incorrect or incomplete answers, struggle with visual prompts, and fail to maintain context or delay responses until target events.
By Xianyun Sun, Chaoyou Fu, Zhengye Zhang, Feiyang Duan, Qingyuan Cao, Yonghui Niu, Sihang Yuan, Ge Zhang, Caifeng Shan
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang