arXiv AI

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning

arXiv:2509. 12263v3 Announce Type: replace Abstract: Large multimodal models (LMMs) encode physical laws observed during training, such as momentum conservation, as parametric knowledge.

arXiv AI
Jul 14

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.

By Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
arXiv Machine Learning
Sep 11

New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

The paper introduces a new benchmark for vision‑language models that tests their ability to decide whether to answer a physics question immediately or to request additional experimental evidence. Each problem presents one measurement image and four possible physical worlds defined by two masses and two values of another property; the model must either stop and answer or choose the cheapest experiment that resolves the question. Across six open models and 144 parameter sets, the models almost always repeat the same action even when the optimal choice changes, and only a single model gets both decisions correct on 5.9% of cases.

By Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang
Hugging Face Trending Papers
Aug 3

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws.

arXiv Computer Vision
2d ago

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

PhysVista is a new benchmark that evaluates physical intelligence in Vision‑Language Models (VLMs) by integrating perception, reasoning, and plausibility assessment into a closed cognitive loop. It distinguishes between event‑level and scale‑level reasoning and tests models on both real‑world and AI‑generated videos to provide a holistic, fine‑grained analysis of physical understanding. Experiments show significant gaps in VLMs’ physical reasoning and plausibility assessment, underscoring the need for more principled, physically grounded multimodal designs.

By Xinge Peng, Yiting Lu, Tianwu Zhi, Wen Wen, Jianzhao Liu, Xin Li, Zhibo Chen
arXiv AI
6d ago

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

The paper introduces LIFT, a lightweight vector‑intervention technique that transfers reasoning capability from a base large language model (LLM) to a vision‑language model (VLM) without retraining the VLM backbone. LIFT defines Reasoning Vectors as differences in hidden states between a reasoning path with an explicit trace and a solver path without it, and injects these vectors into the VLM’s language‑side activations. Experiments on two VLMs across six reasoning benchmarks show that vectors derived from the base LLM consistently outperform those derived from the aligned VLM, indicating that the base LLM is a more effective source for recovering degraded reasoning. "whyItMatters":"The study demonstrates that a simple, frozen‑backbone intervention can partially restore reasoning abilities in multimodal models, highlighting the value of leveraging the original language model’s reasoning power."

By Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang
arXiv Computer Vision
Sep 4

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

VKnowU is a benchmark that tests multimodal large language models (MLLMs) on their grasp of visual knowledge—intuitive, human-like understanding of physical and social principles in videos. The benchmark contains 1,680 questions across 1,249 videos, covering eight core types of visual knowledge, and shows that current state‑of‑the‑art MLLMs still lag behind human performance, especially on world‑centric tasks. To address this gap, the authors release VKnowQA and VideoKnow+, a baseline model that incorporates visual knowledge via a See‑Think‑Answer framework and reinforcement learning, improving performance on VKnowU and related datasets.

By Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
arXiv Computation and Language
Sep 14

MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation

arXiv:2512.14691v3 Announce Type: replace Abstract: Modern multimodal generative models can synthesize visually compelling images and videos, but it remains unclear whether this visual fluency reflec...

By Zefan Cai, Haoyi Qiu, Tianyi Ma, Haozhe Zhao, Gengze Zhou, Tingting Liao, Xinyan Velocity Yu, Kung-Hsiang Huang, Ke Wan, Shawn Lin, Parisa Kordjamshidi, Minjia Zhang, Wen Xiao, Jiuxiang Gu, Nanyun Peng, Junjie Hu
arXiv AI
Sep 12

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

ReactHuman is a physics‑grounded benchmark that tests whether multimodal large language models (MLLMs) can make immediate, safety‑critical decisions in simulated humanoid scenarios involving sudden household hazards. The benchmark includes 17 event families, over 1,000 reproducible scenes generated from 240 Hz rigid‑body simulation, and a five‑metric suite evaluating reactions on reasonableness, safety, and physical grounding. Evaluation of seven MLLMs reveals that reactive safety remains unsolved, with models frequently mishandling hazards, relying on appearance over motion, and missing key interception points.

By Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen, Zicheng Zhao, Dekun Wu, Dongqing Zhang, Bang Liu