OpenAI Blog

Evolution strategies as a scalable alternative to reinforcement learning

Read the original on OpenAI Blog →

We’ve discovered that evolution strategies (ES), an optimization technique that’s been known for decades, rivals the performance of standard reinforcement learning (RL) techniques on modern RL benchmarks (e. g.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

OpenAI Blog
Apr 18, 2018

Evolved Policy Gradients

We’re releasing an experimental metalearning approach called Evolved Policy Gradients, a method that evolves the loss function of learning agents, which can enable fast training on novel tasks. Agents trained with EPG can succeed at basic tasks at test time that were outside their training regime, like learning to navigate to an object on a different side of the room from where it was placed during training.

arXiv Machine Learning
1d ago

Continual Reinforcement Learning with Neuroevolution

The paper investigates continual reinforcement learning using neuroevolution, comparing evolution strategies (ES) and genetic algorithms (GAs) across diverse environments and network sizes. ES consistently achieves a better balance between stability and plasticity, while GAs are more plastic but forget more. The authors attribute this to ES finding wider neighborhoods in weight space, with overlap between consecutive tasks correlating with the stability-plasticity trade‑off, and note that common RL plasticity issues do not transfer to neuroevolution.

By Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi
arXiv Machine Learning
Jun 10

Discovering Interpretable Multi-Parameter Control Policies for Evolutionary Algorithms Using Deep Reinforcement Learning

arXiv:2606. 10129v1 Announce Type: new Abstract: While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis of parameter control remains largely restricted to single-parameter settings, owing to the difficulty of deriving effective, interpretable multi-parameter policies amenable to formal study.

By Tai Nguyen, Phong Le, Carola Doerr, Nguyen Dang