arXiv AI By Mohammed Abouzaid, Nikhil Srivastava, Rachel Ward, Lauren Williams

First Proof Second Batch

Read the original on arXiv AI →

arXiv:2606. 18119v1 Announce Type: new Abstract: To assess the ability of current AI systems to correctly solve research-level mathematics problems, we tested several AI systems on a set of ten problems in a broad range of mathematical fields; these problems arose naturally in the research process of the contributors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

FrontierMath Erd\H{o}s

arXiv:2609. 25050v1 Announce Type: new Abstract: We introduce FrontierMath Erd\H{o}s (FME), a benchmark of 68 Erd\H{o}s problems that are open as of August 2026.

By Tom Adamczewski (Epoch AI), Thomas F. Bloom (University of Manchester)
arXiv AI
Jun 9

Advancing Mathematics Research with AI-Driven Formal Proof Search

arXiv:2605. 22763v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research.

By George Tsoukalas, Anton Kovsharov, Sergey Shirobokov, Anja Surina, Moritz Firsching, Gergely B\'erczi, Francisco J. R. Ruiz, Arun Suggala, Adam Zsolt Wagner, Eric Wieser, Lei Yu, Aja Huang, Mikl\'os Z. Horv\'ath, Andrew Ferraiuolo, Henryk Michalewski, Edward Lockhart, Codrut Grosu, Thomas Hubert, Matej Balog, Pushmeet Kohli, Swarat Chaudhuri
arXiv AI
Sep 17

AI and Human Approaches to Mathematical Problem Solving

The article compares how AI systems and human mathematicians approach 11 long-standing mathematical problems. It finds that AI reports focus more on solving the problem and linking ideas across fields, while human papers emphasize method explanation, assumptions, limitations, and future questions. Both approaches show similar levels of generality, but differ in their research profiles.

By Yang Ding