arXiv:2606. 06526v1 Announce Type: new Abstract: Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step solutions, or complete proofs.
By Sherin Muckatira, Jesse Geneson, Slava Gerovitch, Pavel Etingof, Mikhail Gronas, Anna Rumshisky
The paper introduces a new human‑AI collaboration paradigm for mathematical discovery, shifting from selecting individual problems to exploring broad research directions. It presents the Find, Attempt, and Recommend (FAR) pipeline, which automatically searches a literature corpus, filters candidate conjectures, and surfaces promising resolutions for expert review. In a combinatorics pilot, FAR processed over 5,000 papers, identified thousands of open conjectures, and ultimately highlighted 77 items that led to new discoveries.
By Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
arXiv:2604. 03789v2 Announce Type: replace-cross Abstract: Recent advances in large language models have significantly improved their ability to perform mathematical reasoning, extending from elementary problem solving to increasingly capable performance on research-level problems.
By Haocheng Ju, Guoxiong Gao, Jiedong Jiang, Bin Wu, Zeming Sun, Shurui Liu, Leheng Chen, Yutong Wang, Yuefeng Wang, Zichen Wang, Wanyi He, Peihao Wu, Liang Xiao, Ruochuan Liu, Bryan Dai, Bin Dong
The paper introduces the AI Mathematician (AIM) framework, which leverages Large Reasoning Models (LRMs) to tackle frontier mathematical research. AIM addresses the complexity and procedural rigor of research problems through an exploration mechanism for longer solution paths and a pessimistic reasonable verification method for reliability. Early experiments show AIM can autonomously construct significant proof components and uncover non‑trivial insights across real‑world mathematical topics.
By Yuanhang Liu, Yanxing Huang, Yanqiao Wang, Peng Li, Yang Liu
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed.
arXiv:2607. 07779v1 Announce Type: cross Abstract: Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages.
By Eric Jiang, Xiao Liang, Yikai Zhang, Yingjia Wan, Mengting Li, Haikang Deng, Alexander K. Taylor, Justin Baker, Rushil Raghavan, Junyi Zhang, Ying Nian Wu, Andrea L. Bertozzi, Kai-Wei Chang, Raghu Meka, Matthew Sottile, Nanyun Peng, Amit Sahai, Terence Tao, Wei Wang