FrontierMath Erd\H{o}s
arXiv:2609. 25050v1 Announce Type: new Abstract: We introduce FrontierMath Erd\H{o}s (FME), a benchmark of 68 Erd\H{o}s problems that are open as of August 2026.
arXiv:2606. 18119v1 Announce Type: new Abstract: To assess the ability of current AI systems to correctly solve research-level mathematics problems, we tested several AI systems on a set of ten problems in a broad range of mathematical fields; these problems arose naturally in the research process of the contributors.
arXiv:2609. 25050v1 Announce Type: new Abstract: We introduce FrontierMath Erd\H{o}s (FME), a benchmark of 68 Erd\H{o}s problems that are open as of August 2026.
arXiv:2608. 11195v1 Announce Type: new Abstract: AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively.
arXiv:2605. 22763v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research.
The article compares how AI systems and human mathematicians approach 11 long-standing mathematical problems. It finds that AI reports focus more on solving the problem and linking ideas across fields, while human papers emphasize method explanation, assumptions, limitations, and future questions. Both approaches show similar levels of generality, but differ in their research profiles.
We share our AI model’s proof attempts for the First Proof math challenge, testing research-grade reasoning on expert-level problems.
AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations.
arXiv:2606. 02484v1 Announce Type: new Abstract: Recent advances in large language models and agentic AI systems have enabled significant progress in mathematical discovery, from solving competition problems to tackling research-level conjectures.
arXiv:2607. 20525v1 Announce Type: new Abstract: OpenAI's recent disproof of the Erd\H{o}s unit distance conjecture marked a milestone for AI in mathematics.
arXiv:2607. 09721v1 Announce Type: cross Abstract: To help evaluate the mathematical skills of current AI systems, we present a set of formulas for fundamental mathematical constants.
The paper introduces the AI Mathematician (AIM) framework, which leverages Large Reasoning Models (LRMs) to tackle frontier mathematical research. AIM addresses the complexity and procedural rigor of research problems through an exploration mechanism for longer solution paths and a pessimistic reasonable verification method for reliability. Early experiments show AIM can autonomously construct significant proof components and uncover non‑trivial insights across real‑world mathematical topics.
arXiv:2607. 14582v1 Announce Type: new Abstract: Existing LLM-based theorem provers have achieved impressive results on formal mathematics benchmarks, yet they remain confined to acting as autonomous agents that prove a stated proposition.
arXiv:2607. 27705v1 Announce Type: cross Abstract: Large language models can contribute useful ideas to mathematical research, yet long-horizon proof attempts remain difficult to coordinate, evaluate, and reproduce.