arXiv:2608. 08964v1 Announce Type: new Abstract: The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models (LLMs).
By Harish Kashyap, Kiran Byadarhaly, Sriram Chakaravarthy, Sanyukta Tuti, Aryan Mistry
arXiv:2607. 27670v2 Announce Type: replace-cross Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions.
By Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao
Euclid-Omni is a unified neuro‑symbolic framework that integrates a formal geometry system with Large Language Models and Vision‑Language Models to solve both calculation and proving problems in Euclidean geometry up to Olympiad level. Its core component, Euclidea, automatically generates deductive reasoning steps and algebraic computations, while a data‑generation pipeline creates synthetic symbolic problems, diagrams, and natural‑language translations for training. Experiments show that VLMs trained on this synthetic data outperform on calculation tasks, and LLMs paired with Euclidea match state‑of‑the‑art proving systems using far less compute and data.
By Zhaoyu Li, Hangrui Bi, Youyuan Zhang, Wenjie Ma, Zenan Li, Zhaolei Zhang, Xujie Si, Kaiyu Yang
ChitraMiti introduces a synthetic benchmark of 12,874 Bengali planar geometry problems with structured 15‑attribute descriptions, alongside a complementary set of 500 textbook diagrams. Using a three‑phase protocol that tests diagram‑only, diagram‑plus‑description, and description‑only inputs, the study finds that description‑only performance matches diagram‑plus‑description performance across several VLMs, yet models still struggle with cross‑modal verification. Fine‑tuning on ChitraMiti improves results on both datasets, though a gap remains compared to the best zero‑shot model.
By Khan Raiyan Ibne Reza, Sanjana Aktar Maria, Sumaiya Tabassum Nimi, Md Adnan Arefeen
arXiv:2608.28600v1 Announce Type: new
Abstract: Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their...
By Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo
PhysElite is a new bilingual multimodal benchmark designed to evaluate large language models on Olympiad-level physics problems. It contains 11,586 problems, each paired with visual diagrams, step-by-step bilingual Chinese‑English solution derivations, and the final answer. Benchmarking 18 models revealed that even the best reaches only 33.7% accuracy, and a step‑level analysis highlights where models falter in reasoning.
By Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang