arXiv AI By Jiacheng Chen, Xinyu Zhang, Shunkai Zhang, Yanmohan Wang, Lin Li, Tiancheng Qin, Qin Wang, Zhengmao Zhu, Tianle Li, Jingyang Li, Zehan Li, Binyang Jiang, Jin Zhu, Han Ding, Fei Yu, Chenyu Du, Zijian Song, Jiayuan Song, Zhi Zhang, Yunan Huang, Weiyu Cheng, Pengyu Zhao, Yu Cheng

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Read the original on arXiv AI →

arXiv:2606. 13473v1 Announce Type: cross Abstract: We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

The paper presents a method for training Nemotron 3 Ultra to generate proofs for difficult Olympiad mathematics. By fine‑tuning two specialist checkpoints with supervised learning and reinforcement learning, the authors evaluate how checkpoint selection, verification, and refinement affect performance. The resulting open‑model pipeline, which operates entirely in natural language without external tools, achieved 30 out of 42 points at IMO 2026, meeting the gold‑medal threshold, and the authors release the checkpoints, training data, code, solutions, and a new benchmark of 200 problems.

By Ivan Moshkov, Stephen Ge, George Armstrong, Wei Du, Sadegh Mahdavi, Igor Gitman
arXiv Computation and Language
4d ago

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

AdvancedMathBench is a new benchmark suite that evaluates large language models on advanced mathematical proof generation and verification. It includes ProverBench, with 245 problems from undergraduate to doctoral qualifying‑exam levels, and VerifierBench, which tests models’ ability to judge proof validity using 888 expert‑annotated trajectories. The suite features an automatic verification pipeline trained on expert data, and results show that even state‑of‑the‑art models perform poorly, highlighting a gap between generation and verification skills.

By Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Zhouqi Hua, Wenyong Huang, Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu, Wenwei Zhang, Kai Chen
arXiv Computation and Language
Sep 17

ProofVerifier: A Scalable, Diversity-Driven Framework for Natural-Language Proof Verification

arXiv:2602.02377v3 Announce Type: replace Abstract: While large language models (LLMs) have achieved strong performance on mathematical problems with verifiable answers, many advanced problems are pr...

By Haotong Yang, Zitong Wang, Shijia Kang, Siqi Yang, Wenkai Yu, Xu Niu, Yike Sun, Yi Hu, Zhouchen Lin, Muhan Zhang