OpenAI Blog

Retro Contest: Results

The first run of our Retro Contest—exploring the development of algorithms that can generalize from previous experience—is now complete.

arXiv Machine Learning
Sep 3

Towards Solving the Gilbert-Pollak Conjecture via Large Language Models

The paper announces a new lower bound of 0.8559 for the Steiner ratio, improving on the previous 0.824 bound for the Gilbert‑Pollak Conjecture. It introduces an AI system that uses large language models to generate rule‑constrained geometric lemmas, which are then turned into executable verification functions that certify the bound. The approach relies on only thousands of LLM calls, highlighting the feasibility of LLM‑based methods for advanced mathematical research.

By Yisi Ke, Tianyu Huang, Yankai Shu, Di He, Jingchu Gai, Liwei Wang
arXiv AI
Jun 17

First Proof Second Batch

arXiv:2606. 18119v1 Announce Type: new Abstract: To assess the ability of current AI systems to correctly solve research-level mathematics problems, we tested several AI systems on a set of ten problems in a broad range of mathematical fields; these problems arose naturally in the research process of the contributors.

By Mohammed Abouzaid, Nikhil Srivastava, Rachel Ward, Lauren Williams
arXiv AI
Sep 4

Counterfactual Routing Using Integer Programming with Constraint Generation

The paper "Counterfactual Routing Using Integer Programming with Constraint Generation" presents a solution to the IJCAI 2025 Counterfactual Routing Competition. The authors model the problem as an integer program and iteratively add constraints until an exact solution is found. In evaluation on held‑out test instances, their method ranked fourth in solution quality and was the fastest, averaging 9.0 seconds versus 118.8 seconds for the next‑fastest submission.

By Dani\"el Vos, Sterre Lutz