The paper introduces a framework that separates physical modeling from execution in physics reasoning tasks. It uses a two‑stage post‑training approach: supervised fine‑tuning to build structured models and reinforcement learning with rubric‑based feedback to refine them. Experiments on PhysReason, PhyX, and SeePhys show that this explicit modeling improves reasoning performance by about 3% on average for small LLMs.
By Ye Zhang, Xuehang Guo, Rui Pan, Pengfei Yu, Denghui Zhang, Manling Li, Qingyun Wang
arXiv:2605.26087v2 Announce Type: replace-cross
Abstract: Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of...
By Matt L. Wiemann, Lindsay M. Smith, Peter Melchior, Siddharth Mishra-Sharma, Andrew Gordon Wilson, Pavel Izmailov, Carolina Cuesta-L\'azaro
The paper introduces a five-task diagnostic experiment that separates perceptual and reasoning failures in multimodal large language models on physics and geometry benchmarks. It finds that misinterpreting diagrams hurts performance even on text-only solvable problems, and that accuracy improves when models receive human-authored captions. The study shows that correcting captions can recover many errors, revealing distinct reasoning bottlenecks that differ by domain, while a heavily pretrained model still underperforms and often truncates reasoning traces.
By Raj Jaiswal, Sree Krishna Uppalapati, Dhruvkumar Patel, Ria Khatoniar, Tanuja Ganu, Rajiv Ratn Shah
Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual fidelity required to assess true low-level Newtonian understanding.
arXiv:2608. 14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically.
By Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber
arXiv:2606. 29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search.
By Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka, Varun Gandhi, Scott Niekum
arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.
By Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
arXiv:2512. 19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation.
By Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
The paper introduces a new benchmark for vision‑language models that tests their ability to decide whether to answer a physics question immediately or to request additional experimental evidence. Each problem presents one measurement image and four possible physical worlds defined by two masses and two values of another property; the model must either stop and answer or choose the cheapest experiment that resolves the question. Across six open models and 144 parameter sets, the models almost always repeat the same action even when the optimal choice changes, and only a single model gets both decisions correct on 5.9% of cases.
By Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang
arXiv:2605.18630v2 Announce Type: replace
Abstract: Large Language Models (LLMs) are increasingly deployed as scientific AI as- sistants, and a growing body of benchmarks evaluates their capabilities...
By Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya, Patrick Emami, Anurag Acharya, Sameera Horawalavithana, Shaowu Pan
OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.
By Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
arXiv:2609.13225v1 Announce Type: cross
Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fail...
By Sarthak Sattigeri