OpenAI Blog

Advancing science and math with GPT-5.2

GPT-5. 2 is OpenAI’s strongest model yet for math and science, setting new state-of-the-art results on benchmarks like GPQA Diamond and FrontierMath.

Towards Data Science
Sep 8

How to Maximize GPT-6 Astra

The article titled "How to Maximize GPT-6 Astra" shares the author’s first impressions of OpenAI’s new frontier model. It discusses initial experiences and observations with GPT-6 Astra, offering insights into its capabilities and potential applications. The post was originally published on Towards Data Science.

By Eivind Kjosbakken
arXiv Computation and Language
Sep 23

FrontierMath Erd\H{o}s

arXiv:2609. 25050v1 Announce Type: new Abstract: We introduce FrontierMath Erd\H{o}s (FME), a benchmark of 68 Erd\H{o}s problems that are open as of August 2026.

By Tom Adamczewski (Epoch AI), Thomas F. Bloom (University of Manchester)
arXiv Machine Learning
Sep 11

Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification

The paper introduces HorizonMath, a benchmark of 113 largely unsolved mathematical problems across eight domains, paired with an open-source framework for automated verification. It focuses on the generator‑verifier gap, targeting problems that are hard to discover but easy to verify computationally, thereby avoiding costly formal proof verification or manual review. Using this framework, the authors found six novel solutions—three each from GPT‑5.4 Pro and GPT‑5.6 Sol—demonstrating that current models can contribute to mathematical research, while most state‑of‑the‑art models score below 10%.

By Erik Y. Wang, Sumeet R. Motwani, James V. Roggeveen, Eliot Hodges, Dulhan Jayalath, Charles London, Kalyan Ramakrishnan, Jakob Foerster, Cheng Zhang, Flaviu Cipcigan, Philip Torr, Alessandro Abate