OpenAI Blog

Better exploration with parameter noise

Read the original on OpenAI Blog →

We’ve found that adding adaptive noise to the parameters of reinforcement learning algorithms frequently boosts performance. This exploration method is simple to implement and very rarely decreases performance, so it’s worth trying on any problem.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

arXiv Machine Learning
5d ago

Gradient-Momentum Coupling: A Parameter-Space Proxy for Learning Progress

The paper introduces Gradient‑Momentum Coupling (GMC), a method that quantifies learning progress by measuring how strongly a sample influences changes in the parameter space, using the normalized absolute product of its gradient and the momentum of previous gradients. GMC filters out noise by accumulating consistent directions of change while canceling random fluctuations, leading to a more uniform prioritization across tasks with varying noise levels and better ranking of learnable tasks by improvement speed. Experiments on MiniGrid MultiRoom tasks show that replacing prediction error with GMC in the Intrinsic Curiosity Module restores exploration capabilities that were lost to unpredictable observations.

By Samuel Blad, Martin L\"angkvist, Amy Loutfi
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas