arXiv AI

Finding Common Mistakes In Modelling With Mathematical Formalisms Using LLMs

arXiv Computation and Language
Aug 24

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

The paper introduces an epistemically and formally grounded ensemble (EFG) of large language model judges to evaluate autoformalization tasks in formal mathematics. It defines four criteria—logical preservation, mathematical consistency, formal quality, and formal validity—to provide a transparent, multi‑granular assessment. Experiments show that this ensemble outperforms coarse‑grained models, offering a scalable and interpretable proxy for evaluating formal mathematical reasoning.

By Lan Zhang, Marco Valentino, Jordan Meadows, Andre Freitas
arXiv Computation and Language
Aug 31

SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction

SMRC is a new method that aligns large language models with student reasoning for mathematical error correction. It treats student reasoning as a multi‑step decision problem and uses Monte Carlo Tree Search to find optimal correction paths, while a breadth‑first search guided by the model generates reward signals that are back‑propagated to supervise intermediate steps. The authors also introduce the MSEB benchmark of 158 high‑school math problems and a dual evaluation protocol focusing on solution accuracy and correct‑step retention, showing that SMRC outperforms existing methods on several datasets.

By Biaojie Zeng, Min Zhang, Juan Zhou, Fengrui Liu, Ruiyang Huang, Yu Song, Xin Lin
arXiv Computation and Language
Aug 27

MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

MathAdv is a diagnostic benchmark for formal theorem proving that covers 13 undergraduate- and graduate-level mathematics domains. It includes Lean 4 proofs and up to three auxiliary tasks—multiple-choice questions, fill-in-the-blank problems, and expert-crafted transformations—to probe knowledge, informal reasoning, and robustness to problem presentation. Evaluation of current theorem provers shows formalization is a major bottleneck, performance varies by domain, natural-language guidance can help or hinder models, and equivalent reformulations reveal significant robustness gaps.

By Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu, Jiaqi Wang, Chenghao Deng, Xiayimei Han, Vlasios Mastrantonis, Dmitrii Gudin, Shaopeng Zhu, Abdirisak Abdullahi Mohamed, Bilal Hamdi Aytekin, Jiewen Lang, Zezheng Song, Furong Huang
arXiv AI
Jul 3

Aria: An Agent For Retrieval and Iterative Auto-Formalization via Dependency Graph

arXiv:2510. 04520v2 Announce Type: replace Abstract: Accurate auto-formalization of theorem statements is essential for advancing automated discovery and verification of research-level mathematics, yet remains a major bottleneck for LLMs due to hallucinations, semantic mismatches, and their inability to synthesize new definitions.

By Hanyu Wang, Ruohan Xie, Yutong Wang, Guoxiong Gao, Xintao Yu, Bin Dong
arXiv AI
Jun 3

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems

arXiv:2606. 02588v1 Announce Type: cross Abstract: We present Lean-GAP (Lean-Graduate Agebra Problems), 430 formalized graduate-level algebra problems from the textbook Abstract Algebra by Dummit and Foote.

By Seewoo Lee, Byung-Hak Hwang, Hyojae Lim, Jihoon Hyun, Ilkyoo Choi, Yeachan Park, Jineon Baek, Hyukpyo Hong, Keewoo Lee, Jaeseong Heo, Hyungryul Baik, Chul-hee Lee, Kyu-Hwan Lee
arXiv AI
Sep 10

Tracing Mathematical Proficiency Through Problem-Solving Processes

The paper introduces Knowledge Tracing Leveraging Problem‑Solving Process (KT‑PSP), a method that incorporates students’ problem‑solving steps to model mathematical proficiency more comprehensively than traditional knowledge tracing. It presents the KT‑PSP‑25 dataset and a new framework, StatusKT, which uses a teacher‑student‑teacher LLM pipeline to extract proficiency indicators, generate responses, and evaluate mastery. Experiments show that StatusKT improves prediction accuracy and offers interpretable explanations by explicitly modeling proficiency.

By Jungyang Park, Suho Kang, Jaewoo Park, Jaehong Kim, Jaewoo Shin, Seonjoon Park, Youngjae Yu