WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining, using semantically coherent action events as the atomic unit of learning. Existing WAMs commonly initialize from multimodal or video foundation models and then optimize fixed-length action chunks conditioned directly on the current observation and instruction.
We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instances from other videos or images that depict that state. Unlike action anticipation, which predicts a label, moment retrieval, which localizes an observed event within a video, or video generation, which synthesizes pixels, PSR combines anticipation with cross-instance retrieval across multiple temporal horizons.
arXiv:2608. 07420v1 Announce Type: new Abstract: World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions.
By Xinyi Li, Zaishuo Xia, Chenjie Hao, Yubei Chen
The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.
By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang
SAM3Dual is a training‑free inference extension of pretrained SAM 3 that won third place in the MOSEv2 track of the 8th Large‑scale Video Object Segmentation Challenge. It separates temporal memory into short‑term and long‑term branches, fuses their responses deterministically, and modulates them with previous‑frame confidence, all while keeping SAM 3 parameters frozen. The approach achieved an official J&F score of 64.37, demonstrating competitive long‑term VOS performance without task‑specific training.
By JeongRae Kim, Chaehyun Kim, Changwon Lim
Kairos is a new video dataset designed for fine-grained video-language modeling, featuring long-duration videos from ten minutes to half an hour. Each video is annotated with time-resolved labels that capture ongoing actions, entity appearances, attributes, interactions, and evolving contextual cues throughout the timeline. The dataset supports fine-grained evaluation, long-range modeling, reasoning, instruction data construction, representation learning, and video generation.
By Ruibo Ming, Lei Sun, Deheng Zhang, He Zhang, Jialu Li, Jian Wang, Zhendong Li, Mengshun Hu, Danda Pani Paudel, Luc Van Gool, Jinjin Gu
arXiv:2505.01583v2 Announce Type: replace
Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...
By Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou, Vivian Wang, Huayu Wang, Hsiang-Wei Huang, Wenhao Chai, Hou-I Liu, Kuang-Ming Chen, Cheng-Yen Yang, Yi-Ling Chen, Vibhav Vineet, Qin Cai, Jenq-Neng Hwang
The paper introduces STITCH, a training‑free method that partitions videos into semantically meaningful temporal chunks using a frozen video‑text backbone. By detecting changes in the embedding sequence of short video windows, STITCH produces reusable temporal abstractions that can be applied to multiple tasks such as event boundary detection, language‑based moment retrieval, and frame selection for vision‑language models. Experiments show that STITCH performs competitively with specialized methods while requiring no task‑specific training, especially when processing is limited to a few frames or tokens.
By Etienne Casanova, Sevan Brodjian, Pietro Perona
The paper introduces an adaptive temporal modeling framework for weakly supervised video anomaly detection that addresses the limitations of rigid Multiple Instance Learning approaches. It presents a Temporal Refinement Module using dynamic positional encoding and a learnable class token to capture long‑range dependencies, and an Event Segmentation Module that identifies event boundaries via temporal discontinuity analysis to produce discriminative event‑level representations. An adaptive similarity‑based fusion strategy replaces fixed top‑k heuristics, dynamically integrating snippet‑level and event‑level anomaly scores into video‑level predictions, and the method outperforms state‑of‑the‑art baselines on two benchmarks.
By Changyi Li, Yu Xiao
arXiv:2609.38839v1 Announce Type: new
Abstract: Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retainin...
By Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan
Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames and are therefore expensive to process exhaustively. Existing methods usually construct compact visual inputs from long videos under a limited visual budget.
StarWM introduces a self‑supervised attention routing mechanism that selectively applies reconstruction only to dynamically relevant regions of visual input. By combining a cross‑attention module with a dual‑stream decoder and stop‑gradient barriers, it balances faithful environmental dynamics capture with abstraction of irrelevant content. Experiments on DeepMind Control show that StarWM outperforms both reconstruction‑based and reconstruction‑free baselines, especially under distractor conditions, and preserves state attributes over long‑horizon imagination.
By Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid, Sebastian Gottwald, Arne Peter Raulf, Daniel Alexander Braun