The paper introduces VWG-Bench, a benchmark covering nine reasoning dimensions and 38 tasks to evaluate video generative models on symbolic reasoning, physical laws, and goal pursuit. It also presents Vid-PRE, a prompt-rewriting framework that offloads reasoning to a VLM, improving logical performance without changing the generator architecture. Experiments show that current models excel at visual quality but struggle with logic-heavy tasks, while Vid-PRE significantly boosts reasoning across multiple generators.
By Meng Luo, Yicheng Liu, Jiahao Wang, Yuanxing Zhang, Xin Tao, Pengfei Wan, Kun Gai, Hao Fei
The paper introduces structured video prompting, a training‑free inference‑time technique that augments input videos with lightweight spatial and temporal structure to provide explicit anchors for evidence organization. By applying this method to two video benchmarks and two open video‑language models, the authors demonstrate performance improvements across several tasks, with gains varying by model and task. The study suggests that failures in video‑language models stem not only from reasoning capacity but also from how video evidence is presented during inference.
By Sadegh Mohammadian
arXiv:2608.19583v2 Announce Type: replace-cross
Abstract: Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet rel...
By Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai
arXiv:2608. 19583v1 Announce Type: cross Abstract: Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames.
By Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai
arXiv:2604.22875v3 Announce Type: replace-cross
Abstract: When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-languag...
By Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri, Trung Bui, Anh Totti Nguyen
CinematicVQA is a new benchmark for evaluating large vision‑language models on film‑grammar reasoning. It introduces the Cinematic Scene Graph, a structured representation linking filming techniques to perceptual effects and narrative functions, and tests models on tasks beyond low‑level technique recognition. The study finds a semantic gap where models excel at describing visuals but struggle to identify underlying techniques, and shows that fine‑tuning improves performance on narrative function and multi‑hop reasoning.
By Shuo Xing, Pooja Verlani, Balu Adsumilli, Zhengzhong Tu
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible.
arXiv:2603. 08652v2 Announce Type: replace Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning.
By Haodong Li, Chunmei Qing, Huanyu Zhang, Dongzhi Jiang, Yihang Zou, Hongbo Peng, Dingming Li, Yuhong Dai, ZePeng Lin, Juanxi Tian, Yi Zhou, Siqi Dai, Jingwei Wu, Pheng-Ann Heng
arXiv:2608. 04726v1 Announce Type: new Abstract: Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels.
By Yongxin Wang, Ruizhe Zhou, Yueling Tang, Yingying Zhu, Xuemin Zhao, Xiaojun Chang, Xiaodan Liang
The paper investigates how visual presentation affects vision‑language models (VLMs) on the SPaRC spatial planning benchmark. By adding lightweight input‑side scaffolds that keep the visual modality but make spatial structure clearer, the authors achieve up to a 34.0‑percentage‑point accuracy boost across multiple VLMs, and an additional 4.6 points when combined with GRPO training. Analyses reveal that these improvements stem mainly from reduced grounding errors, while rule‑based reasoning remains difficult, highlighting visual presentation as a key determinant of whether VLM benchmarks test grounded perception, downstream reasoning, or both.
By Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas, Jan Philip Wahle, Daniel Kurzawe, Bela Gipp
arXiv:2608.22174v1 Announce Type: new
Abstract: Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding...
By Yubo Zhu, Zhehan Kan, Jingyi Yang, Miaolin Chen, Jinbo Xing, Kai Zhu, Zijian Wang, Sheng Zhong, Wei Tong
VisionCoach is an input‑adaptive reinforcement learning framework that enhances spatio‑temporal grounding in video reasoning by using visual prompting during training. The system selectively applies visual prompts to challenging inputs, amplifying question‑relevant evidence and suppressing distractors, and then internalizes these improvements through self‑distillation so that inference can be performed on raw videos without prompts. Experiments on multiple benchmarks (V‑STAR, VideoMME, World‑Sense, VideoMMMU, PerceptionTest, and Charades‑STA) show that VisionCoach achieves state‑of‑the‑art performance while maintaining a single efficient inference pathway.
By Daeun Lee, Shoubin Yu, Yue Zhang, Mohit Bansal