OpenAI Blog

More on Dota 2

Our Dota 2 result shows that self-play can catapult the performance of machine learning systems from far below human level to superhuman, given sufficient compute. In the span of a month, our system went from barely matching a high-ranked player to beating the top pros and has continued to improve since then.

OpenAI Blog
Aug 11, 2017

Dota 2

We’ve created a bot which beats the world’s top professionals at 1v1 matches of Dota 2 under standard tournament rules. The bot learned the game from scratch by self-play, and does not use imitation learning or tree search.

OpenAI Blog
Oct 11, 2017

Competitive self-play

We’ve found that self-play allows simulated AIs to discover physical skills like tackling, ducking, faking, kicking, catching, and diving for the ball, without explicitly designing an environment with these skills in mind. Self-play ensures that the environment is always the right difficulty for an AI to improve.

arXiv Machine Learning
Aug 12

Scaling Self-Play with Self-Guidance

arXiv:2604. 20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conjecturer model creates problems for a Solver, and both improve together.

By Luke Bailey, Kaiyue Wen, Kefan Dong, Tatsunori Hashimoto, Tengyu Ma
OpenAI Blog
May 16, 2018

AI and compute

We’re releasing an analysis showing that since 2012, the amount of compute used in the largest AI training runs has been increasing exponentially with a 3. 4-month doubling time (by comparison, Moore’s Law had a 2-year doubling period)[^footnote-correction].

arXiv AI
3d ago

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

AREX-2 is a new approach that enhances the self‑improving ability of large language model agents by combining reflection—producing better solutions—and long‑horizon execution—maintaining effectiveness over many iterations. The method synthesizes improvement trajectories from machine‑learning and algorithmic programming tasks, providing verifiable feedback and sustained iteration. Trained on this data, an agent based on Qwen3.8‑27B achieves strong performance across multiple benchmarks and continues to improve as more iterative rounds are allowed.

By Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu
arXiv AI
6d ago

Self-Play Search Distillation for Large Language Model Reasoning

Self-Play Search Distillation (SPSD) is a framework that generates superhuman synthetic data by having MuZero-like networks play board games in executable environments. The search records are converted into structured reasoning chains that serve as environment‑grounded supervision for training large language models. When applied to Qwen3‑4B‑Base, SPSD improves performance on six mathematics benchmarks from 24.1 to 36.6 and raises the win rate on unseen games from 15% to 45%.

By Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro, Pavlos Vougiouklis, Jeff Z. Pan, Pasquale Minervini
arXiv AI
Jun 3

Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

arXiv:2510. 23216v4 Announce Type: replace Abstract: While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the game industry for crafting authentic AI behaviors.

By Alessandro Sestini, Joakim Bergdahl, Jean-Philippe Barrette-LaPierre, Florian Fuchs, Brady Chen, Fabio Zinno, Michael Jones, Linus Gissl\'en