arXiv:2603. 10652v3 Announce Type: replace-cross Abstract: In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion.
By Yangfan He, Changgyu Boo, Jaehong Yoon
SYNCR is a synthetic benchmark designed to evaluate multimodal large language models on cross‑video reasoning. It contains 4,000 question‑answer pairs across 4,827 unique videos, covering tasks in temporal alignment, spatial tracking, comparative reasoning, and holistic synthesis. The benchmark reveals a significant performance gap between current models and humans, with models excelling at temporal ordering but struggling with precise physical and spatial reasoning.
By Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami
arXiv:2603. 06828v2 Announce Type: replace-cross Abstract: We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better.
By Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Abdullah Ibne Hanif Arean, Juena Ahmed Noshin
The paper introduces a framework that combines world models, which generate concrete visual rollouts of possible futures, with multimodal large language models (MLLMs) that perform abstract reasoning. It proposes a controlled concrete reasoning approach and a new training method called Privileged‑Future On‑Policy Self‑Distillation (PF‑OPSD), which uses ground‑truth future videos as privileged teacher context during training while the student model never sees true futures at test time. Experiments on two human‑verified benchmarks, VRQABench and OpenWorldQA, show that PF‑OPSD improves performance by about 10–11% over baselines and enhances robustness to noisy or conflicting rollouts.
By Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
By Jae Joong Lee
arXiv:2608. 16829v1 Announce Type: cross Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleatoric uncertainty of specific phenomena untested.
By Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham, Jeffrey Hawke