arXiv AI By Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang

PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

Read the original on arXiv AI →

PhysElite is a new bilingual multimodal benchmark designed to evaluate large language models on Olympiad-level physics problems. It contains 11,586 problems, each paired with visual diagrams, step-by-step bilingual Chinese‑English solution derivations, and the final answer. Benchmarking 18 models revealed that even the best reaches only 33.7% accuracy, and a step‑level analysis highlights where models falter in reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.

By Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
arXiv Machine Learning
Jul 30

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

arXiv:2604. 18584v2 Announce Type: replace-cross Abstract: Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity.

By Shaden Alshammari, Kevin Wen, Abrar Zainal, Mark Hamilton, Navid Safaei, Sultan Albarakati, William T. Freeman, Antonio Torralba