AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations.
The paper introduces the AI Mathematician (AIM) framework, which leverages Large Reasoning Models (LRMs) to tackle frontier mathematical research. AIM addresses the complexity and procedural rigor of research problems through an exploration mechanism for longer solution paths and a pessimistic reasonable verification method for reliability. Early experiments show AIM can autonomously construct significant proof components and uncover non‑trivial insights across real‑world mathematical topics.
By Yuanhang Liu, Yanxing Huang, Yanqiao Wang, Peng Li, Yang Liu
The paper introduces a new human‑AI collaboration paradigm for mathematical discovery, shifting from selecting individual problems to exploring broad research directions. It presents the Find, Attempt, and Recommend (FAR) pipeline, which automatically searches a literature corpus, filters candidate conjectures, and surfaces promising resolutions for expert review. In a combinatorics pilot, FAR processed over 5,000 papers, identified thousands of open conjectures, and ultimately highlighted 77 items that led to new discoveries.
By Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
arXiv:2606. 18119v1 Announce Type: new Abstract: To assess the ability of current AI systems to correctly solve research-level mathematics problems, we tested several AI systems on a set of ten problems in a broad range of mathematical fields; these problems arose naturally in the research process of the contributors.
By Mohammed Abouzaid, Nikhil Srivastava, Rachel Ward, Lauren Williams
arXiv:2607. 09721v1 Announce Type: cross Abstract: To help evaluate the mathematical skills of current AI systems, we present a set of formulas for fundamental mathematical constants.
By Michael Shalyt, Rotem Kalisch, Carsten Schneider, Hila Barkan, Elyasheev Leibtag, John Campbell, Shachar Weinbaum, Tali Monderer, Ashvni Narayanan, Ido Kaminer
arXiv:2606. 02484v1 Announce Type: new Abstract: Recent advances in large language models and agentic AI systems have enabled significant progress in mathematical discovery, from solving competition problems to tackling research-level conjectures.
By Leheng Chen, Zihao Liu, Wanyi He, Bin Dong
The paper introduces HorizonMath, a benchmark of 113 largely unsolved mathematical problems across eight domains, paired with an open-source framework for automated verification. It focuses on the generator‑verifier gap, targeting problems that are hard to discover but easy to verify computationally, thereby avoiding costly formal proof verification or manual review. Using this framework, the authors found six novel solutions—three each from GPT‑5.4 Pro and GPT‑5.6 Sol—demonstrating that current models can contribute to mathematical research, while most state‑of‑the‑art models score below 10%.
By Erik Y. Wang, Sumeet R. Motwani, James V. Roggeveen, Eliot Hodges, Dulhan Jayalath, Charles London, Kalyan Ramakrishnan, Jakob Foerster, Cheng Zhang, Flaviu Cipcigan, Philip Torr, Alessandro Abate
UCLA Professor Ernest Ryu and GPT-5 solved a key question in optimization theory, showcasing AI’s role in accelerating mathematical discovery.
arXiv:2605. 22763v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research.
By George Tsoukalas, Anton Kovsharov, Sergey Shirobokov, Anja Surina, Moritz Firsching, Gergely B\'erczi, Francisco J. R. Ruiz, Arun Suggala, Adam Zsolt Wagner, Eric Wieser, Lei Yu, Aja Huang, Mikl\'os Z. Horv\'ath, Andrew Ferraiuolo, Henryk Michalewski, Edward Lockhart, Codrut Grosu, Thomas Hubert, Matej Balog, Pushmeet Kohli, Swarat Chaudhuri
The paper reports on autonomous mathematical discovery within the Station, an open‑world multi‑agent environment where diverse AI agents pursue shared research goals without central coordination. Across 12 construction problems and two case studies, the agents produced novel results—including a new infinite family of finite‑field Kakeya sets, 604‑point kissing configurations in dimension 11, improved records for discretized Kakeya needle and sign uncertainty problems, a stronger lower bound for Erdős’s minimum‑overlap problem, and new infinite families for Book Ramsey numbers—alongside theorems and analyses that explain the constructions. All agent dialogues, proofs, and verification code are released to provide a transparent record of the discovery process.
By Stephen Chung, Wenyu Du, William J. Wesley
arXiv:2603. 08322v2 Announce Type: replace Abstract: We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), coupled with symbolic computation tools, and human strategic direction, jointly produced a new result in combinatorial design theory.
By Hai Xia, Carla P. Gomes, Bart Selman, Stefan Szeider
The article compares how AI systems and human mathematicians approach 11 long-standing mathematical problems. It finds that AI reports focus more on solving the problem and linking ideas across fields, while human papers emphasize method explanation, assumptions, limitations, and future questions. Both approaches show similar levels of generality, but differ in their research profiles.
By Yang Ding