arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.
By Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
arXiv:2608. 02150v2 Announce Type: replace-cross Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities.
By Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou, Ting Yu, Zhibo Wang, Zhan Qin
The paper introduces a new benchmark for vision‑language models that tests their ability to decide whether to answer a physics question immediately or to request additional experimental evidence. Each problem presents one measurement image and four possible physical worlds defined by two masses and two values of another property; the model must either stop and answer or choose the cheapest experiment that resolves the question. Across six open models and 144 parameter sets, the models almost always repeat the same action even when the optimal choice changes, and only a single model gets both decisions correct on 5.9% of cases.
By Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws.
PhysVista is a new benchmark that evaluates physical intelligence in Vision‑Language Models (VLMs) by integrating perception, reasoning, and plausibility assessment into a closed cognitive loop. It distinguishes between event‑level and scale‑level reasoning and tests models on both real‑world and AI‑generated videos to provide a holistic, fine‑grained analysis of physical understanding. Experiments show significant gaps in VLMs’ physical reasoning and plausibility assessment, underscoring the need for more principled, physically grounded multimodal designs.
By Xinge Peng, Yiting Lu, Tianwu Zhi, Wen Wen, Jianzhao Liu, Xin Li, Zhibo Chen
arXiv:2608.28623v1 Announce Type: cross
Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before a...
By Mahir Numayeer Islam, Gakuto Okuyama, Nikolaus Siauw, Shivank Garg, Madhur Panwar, Vasu Sharma