R2M-Bench is a benchmark that evaluates revisit memory in interactive video world models by comparing a revisit pair to two control pairs from the same rollout: a gap‑matched non‑revisit pair and a short‑range pair. It introduces MemoryGain (MG) and Normalized Memory Ratio (NMR) to quantify the revisit advantage over generic temporal stability and normalize it by short‑to‑baseline dynamics. Across 300 instances and seven models, NMR correlates with human judgments and reduces the influence of slow‑motion artifacts, with DreamX‑World‑Memo achieving the highest NMR.
By Qiwen Gu, Bingjie Gao, Rui Chen, Geng Li, Jifan Li, Qishuai Wen, Li Niu, Jing Tang, Xiangxiang Chu, Junqiao Zhao
arXiv:2607. 15621v1 Announce Type: cross Abstract: Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires.
By Yun Li, Jiachen Gong, Simon Thompson, Ehsan Javanmardi, Qunli Zhang, Zifan Zeng, Shiming Liu, Peng Wang, Zixuan Guo, Manabu Tsukada
The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.
By Prakhar Khatri
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
By Jae Joong Lee
CALIPER is a new benchmark that tests whether pretrained visual encoders can infer physical properties such as mass and friction from images. The test involves striking an object twice at known speeds, showing a third strike only up to contact, and asking a linear readout on frozen features to predict how far the object slides. Results show that in clean, fixed‑camera scenes all representations perform similarly, but when camera, lighting, and clutter are varied, only encoders that truly infer physics—like V‑JEPA 2—maintain performance, while random or raw pixel representations fail.
By Aman Mehta, Riya Baviskar
TimeBlind is a diagnostic benchmark designed to evaluate fine‑grained spatio‑temporal compositionality in video large language models (LLMs). It categorizes temporal understanding into three levels—atomic event recognition, event property characterization, and reasoning about event interdependencies—and uses a minimal‑pairs paradigm where video pairs share identical static content but differ only in temporal structure. Across 20 state‑of‑the‑art MLLMs tested on 600 curated instances, the best model achieved only 48.2% instance accuracy, far below human performance of 98.2%, highlighting a reliance on static visual shortcuts rather than true temporal reasoning.
By Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius
GlanceWAM introduces a sparse test‑time imagination approach for world‑action models that decouples visual imagination from control. By asynchronously generating a single lookahead frame on a slow clock and decoding action chunks at a 48 ms control rate purely in latent space, it avoids latency while maintaining high success. The method achieves 72.2 % on the RoboCasa kitchen benchmark and 99.0 % on LIBERO, running 24× faster than synchronous baselines.
By Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Zichong Yang, Yinlin Chen, Lifeng Zhou, Chang-Tien Lu
arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.
By Ali Alavi
arXiv:2608.23565v1 Announce Type: new
Abstract: An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time. The tension is structural: contro...
By Zhifei Chen, Luozhou Wang, Guibao Shen, Dongyu Yan, Shuai Yang, Tianshuo Xu, Yihua Du, Wei Wang, Tianyi Gui, Lianghua Huang, Yingcong Chen
arXiv:2607. 11078v1 Announce Type: cross Abstract: Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode?
By Mohammad Al-Ratrout, Shayla Sharmin, Aditya Raikwar, Roghayeh Leila Barmaki
arXiv:2607. 28362v1 Announce Type: cross Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models.
By Jin Cao, Zian Meng, Kaipeng Zhang
SNF-Bench is an evaluation framework for long‑horizon fixed‑camera video generation that separates static background fidelity from dynamic flow persistence and drift leakage. It reports these three factors independently, using controlled injections of translation, rotation, scale drift, and progressive freezing to validate each metric’s sensitivity. Auditing public checkpoints shows that whole‑frame motion metrics can mislead, while SNF‑Bench reveals the true trade‑offs between motion quality and background stability.
By Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park