arXiv Machine Learning

Can Vision Language Models Learn Intuitive Physics from Interaction?

arXiv:2602. 06033v2 Announce Type: replace Abstract: Pre-trained vision language models do not have good intuitions about the physical world.

arXiv Computation and Language
Aug 25

Decoupled Physical Modeling and Execution for Physics Reasoning

The paper introduces a framework that separates physical modeling from execution in physics reasoning tasks. It uses a two‑stage post‑training approach: supervised fine‑tuning to build structured models and reinforcement learning with rubric‑based feedback to refine them. Experiments on PhysReason, PhyX, and SeePhys show that this explicit modeling improves reasoning performance by about 3% on average for small LLMs.

By Ye Zhang, Xuehang Guo, Rui Pan, Pengfei Yu, Denghui Zhang, Manling Li, Qingyun Wang
arXiv Computer Vision
Aug 31

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

arXiv:2608.27549v1 Announce Type: new Abstract: Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can...

By Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi Duan, Jun Gao, Jiangran Lyu, Fangfu Liu, Jialong Wu
arXiv AI
Jul 14

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.

By Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
Hugging Face Trending Papers
Aug 3

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws.