arXiv AI By Raj Jaiswal, Sree Krishna Uppalapati, Dhruvkumar Patel, Ria Khatoniar, Tanuja Ganu, Rajiv Ratn Shah

Lost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning

Read the original on arXiv AI →

The paper introduces a five-task diagnostic experiment that separates perceptual and reasoning failures in multimodal large language models on physics and geometry benchmarks. It finds that misinterpreting diagrams hurts performance even on text-only solvable problems, and that accuracy improves when models receive human-authored captions. The study shows that correcting captions can recover many errors, revealing distinct reasoning bottlenecks that differ by domain, while a heavily pretrained model still underperforms and often truncates reasoning traces.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 10

From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning

The paper introduces a framework that combines a Geometric Vision Parser and a Symbolic Solver to enable a Large Language Model to solve complex plane geometry problems. By translating diagrams into symbolic representations and performing formal deductions, the approach reduces hallucinations and produces interpretable, human-like solutions. Experiments on a new benchmark from 2025 Chinese Zhongkao exams show performance comparable to Gemini 2.5 Pro.

By Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou, Yan Cao, Xin Shen, Dongcai Lu, Yi Zhou
arXiv Computation and Language
Aug 27

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.

By Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang