OpenAI Blog

Evolved Policy Gradients

We’re releasing an experimental metalearning approach called Evolved Policy Gradients, a method that evolves the loss function of learning agents, which can enable fast training on novel tasks. Agents trained with EPG can succeed at basic tasks at test time that were outside their training regime, like learning to navigate to an object on a different side of the room from where it was placed during training.

arXiv Machine Learning
1d ago

Continual Reinforcement Learning with Neuroevolution

The paper investigates continual reinforcement learning using neuroevolution, comparing evolution strategies (ES) and genetic algorithms (GAs) across diverse environments and network sizes. ES consistently achieves a better balance between stability and plasticity, while GAs are more plastic but forget more. The authors attribute this to ES finding wider neighborhoods in weight space, with overlap between consecutive tasks correlating with the stability-plasticity trade‑off, and note that common RL plasticity issues do not transfer to neuroevolution.

By Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi
arXiv AI
Jun 18

LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents

arXiv:2606. 18388v1 Announce Type: cross Abstract: RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization parameters predominantly oscillate in response to shifting training dynamics.

By Haoyang Fang, Wei Zhu, Boran Han, Alex Zhang, Zhenyu Pan, Shuo Yang, Shuai Zhang, Jiading Gai, Peng Tang, Cuixiong Hu, Xuan Zhu, Huzefa Rangwala, George Karypis, Bernie Wang
arXiv Computation and Language
Aug 27

AEL: Evolving Agent Harness in Open-Ended Environments

The paper introduces Agent Evolving Learning (AEL), a two‑timescale framework that dynamically evolves an LLM agent’s memory‑retrieval harness in open‑ended environments. A fast Thompson‑Sampling bandit selects among retrieval policies each episode, while a slower LLM reflection diagnoses performance drops and injects new policies when the current set plateaus. AEL outperforms ten self‑improving and non‑LLM baselines on a sequential portfolio benchmark, boosting Sharpe ratio by 27% and achieving significant accuracy gains on a support‑ticket routing stream.

By Wujiang Xu, Jiaojiao Han, Minghao Guo, Kai Mei, Xi Zhu, Han Zhang, Dimitris N. Metaxas