arXiv AI By Hsien Xin Peng, Anthony Kim, Alvin Li, Calvin Supasanya, Shivank Garg, Kevin Zhu

Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry

Read the original on arXiv AI →

The paper introduces a new benchmark for diagrammatic reasoning in olympiad geometry, comprising 954 self‑contained problems and a 297‑problem hard subset. Each problem is paired with a human‑authored, high‑fidelity diagram in Asymptote code and a suite of metrics for evaluating diagram construction. Experiments show that current foundation models excel at solving the problems but produce markedly less faithful diagrams, with an average compile success rate of only 36.14%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 5

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

arXiv:2607. 27670v2 Announce Type: replace-cross Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions.

By Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao
arXiv AI
Aug 18

Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry

Euclid-Omni is a unified neuro‑symbolic framework that integrates a formal geometry system with Large Language Models and Vision‑Language Models to solve both calculation and proving problems in Euclidean geometry up to Olympiad level. Its core component, Euclidea, automatically generates deductive reasoning steps and algebraic computations, while a data‑generation pipeline creates synthetic symbolic problems, diagrams, and natural‑language translations for training. Experiments show that VLMs trained on this synthetic data outperform on calculation tasks, and LLMs paired with Euclidea match state‑of‑the‑art proving systems using far less compute and data.

By Zhaoyu Li, Hangrui Bi, Youyuan Zhang, Wenjie Ma, Zenan Li, Zhaolei Zhang, Xujie Si, Kaiyu Yang
arXiv Computer Vision
Sep 14

ChitraMiti: Benchmarking Visual Grounding and Modality Reliance in Bengali Geometric Reasoning

ChitraMiti introduces a synthetic benchmark of 12,874 Bengali planar geometry problems with structured 15‑attribute descriptions, alongside a complementary set of 500 textbook diagrams. Using a three‑phase protocol that tests diagram‑only, diagram‑plus‑description, and description‑only inputs, the study finds that description‑only performance matches diagram‑plus‑description performance across several VLMs, yet models still struggle with cross‑modal verification. Fine‑tuning on ChitraMiti improves results on both datasets, though a gap remains compared to the best zero‑shot model.

By Khan Raiyan Ibne Reza, Sanjana Aktar Maria, Sumaiya Tabassum Nimi, Md Adnan Arefeen
arXiv AI
5d ago

PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

PhysElite is a new bilingual multimodal benchmark designed to evaluate large language models on Olympiad-level physics problems. It contains 11,586 problems, each paired with visual diagrams, step-by-step bilingual Chinese‑English solution derivations, and the final answer. Benchmarking 18 models revealed that even the best reaches only 33.7% accuracy, and a step‑level analysis highlights where models falter in reasoning.

By Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang