OpenAI Blog

Evaluating AI’s ability to perform scientific research tasks

OpenAI introduces FrontierScience, a benchmark testing AI reasoning in physics, chemistry, and biology to measure progress toward real scientific research.

OpenAI Blog
Sep 6

Research acceleration: The view inside OpenAI

The article discusses how coding agents are transforming AI research within OpenAI. It presents early data on agent usage, experiment velocity, task complexity, and the resulting acceleration of research. The piece highlights the growing role of these agents in speeding up development and experimentation.

OpenAI Blog
Dec 11, 2025

Ten years

OpenAI reflects on ten years of progress, from early research breakthroughs to widely used AI systems that reshaped what’s possible. We share lessons from the past decade and why we remain optimistic about building AGI that benefits all of humanity.

arXiv AI
Sep 18

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.

By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister
arXiv AI
2d ago

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

arXiv:2610.00492v1 Announce Type: cross Abstract: When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, fi...

By Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
arXiv AI
Sep 16

The AI-Enabled Scientific Frontier

The paper "The AI-Enabled Scientific Frontier" analyzes 2,507 head‑to‑head comparisons of AI versus other scientific methods across 27 disciplines from 2000 to early 2025. It finds that AI often outperforms traditional statistics but at higher computational cost, while in about a quarter of cases AI is both more expensive and less effective. Compared to scientific computing, AI usually underperforms but at lower cost, though since 2020 its performance has improved, now surpassing computing in over half of the comparisons.

By Gabriel Manso, Emma Fu, Neil Thompson