Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication.
OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.
By Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hin...
ReactBench is a benchmark designed to evaluate the structural reasoning abilities of multimodal large language models (MLLMs) using chemical reaction diagrams. The dataset contains 1,618 expert‑annotated question‑answer pairs that test reasoning across four hierarchical task dimensions, from simple endpoint counting to complex topological analysis. Evaluation of 24 MLLMs shows a performance gap of more than 30% between anchor‑based tasks and holistic structural reasoning tasks, indicating that current models struggle with reasoning over branching, converging, and cyclic structures.
By Qiang Xu, Shengyuan Bai, Yu Wang, He Cao, Leqing Chen, Yuanyuan Liu, Bin Feng, Zijing Liu, Yu Li
arXiv:2607. 12982v1 Announce Type: new Abstract: Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples.
By Ruoran Xu, Wending Gao, Qiufeng Wang
arXiv:2608. 08964v1 Announce Type: new Abstract: The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models (LLMs).
By Harish Kashyap, Kiran Byadarhaly, Sriram Chakaravarthy, Sanyukta Tuti, Aryan Mistry
PhysElite is a new bilingual multimodal benchmark designed to evaluate large language models on Olympiad-level physics problems. It contains 11,586 problems, each paired with visual diagrams, step-by-step bilingual Chinese‑English solution derivations, and the final answer. Benchmarking 18 models revealed that even the best reaches only 33.7% accuracy, and a step‑level analysis highlights where models falter in reasoning.
By Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang
arXiv:2608. 12262v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration.
By Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang
arXiv:2606. 13020v1 Announce Type: new Abstract: Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction.
By Pierre Beckmann, Marco Valentino, Andre Freitas
The paper introduces a five-task diagnostic experiment that separates perceptual and reasoning failures in multimodal large language models on physics and geometry benchmarks. It finds that misinterpreting diagrams hurts performance even on text-only solvable problems, and that accuracy improves when models receive human-authored captions. The study shows that correcting captions can recover many errors, revealing distinct reasoning bottlenecks that differ by domain, while a heavily pretrained model still underperforms and often truncates reasoning traces.
By Raj Jaiswal, Sree Krishna Uppalapati, Dhruvkumar Patel, Ria Khatoniar, Tanuja Ganu, Rajiv Ratn Shah
The paper introduces a framework that combines a Geometric Vision Parser and a Symbolic Solver to enable a Large Language Model to solve complex plane geometry problems. By translating diagrams into symbolic representations and performing formal deductions, the approach reduces hallucinations and produces interpretable, human-like solutions. Experiments on a new benchmark from 2025 Chinese Zhongkao exams show performance comparable to Gemini 2.5 Pro.
By Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou, Yan Cao, Xin Shen, Dongcai Lu, Yi Zhou
Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothesis generation requires an agent continuously coupled to the physical world.