arXiv Machine Learning

Survival Dynamics of Neural and Programmatic Policies in Evolutionary Reinforcement Learning

arXiv:2601. 04365v2 Announce Type: replace Abstract: In evolutionary reinforcement learning tasks (ERL), agent policies are often encoded as small artificial neural networks (NERL).

arXiv Machine Learning
Aug 28

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.

By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
arXiv Machine Learning
Sep 21

Survival Reinforcement Learning: Toward Scalable Self-Supervised RL

The paper introduces Survival Reinforcement Learning (SRL), an online classification-based method that extends the survival value learning framework to maximize an agent’s dwell time at target goals. SRL addresses limitations of contrastive reinforcement learning (CRL) in long-horizon, goal-conditioned tasks by avoiding the uniformity-tolerance dilemma and reducing undesirable bang-bang control behaviors. Across robotic benchmarks, SRL matches CRL on manipulation tasks and outperforms it by 2x to 8x on stable, long-horizon locomotion tasks, suggesting classification-based approaches are a promising direction for scaling reinforcement learning.

By Franki Nguimatsia-Tiofack, Fabian Schramm, Th\'eotime Le Hellard, Justin Carpentier
arXiv Machine Learning
1d ago

Continual Reinforcement Learning with Neuroevolution

The paper investigates continual reinforcement learning using neuroevolution, comparing evolution strategies (ES) and genetic algorithms (GAs) across diverse environments and network sizes. ES consistently achieves a better balance between stability and plasticity, while GAs are more plastic but forget more. The authors attribute this to ES finding wider neighborhoods in weight space, with overlap between consecutive tasks correlating with the stability-plasticity trade‑off, and note that common RL plasticity issues do not transfer to neuroevolution.

By Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi
arXiv Machine Learning
Jun 10

Discovering Interpretable Multi-Parameter Control Policies for Evolutionary Algorithms Using Deep Reinforcement Learning

arXiv:2606. 10129v1 Announce Type: new Abstract: While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis of parameter control remains largely restricted to single-parameter settings, owing to the difficulty of deriving effective, interpretable multi-parameter policies amenable to formal study.

By Tai Nguyen, Phong Le, Carola Doerr, Nguyen Dang
arXiv Machine Learning
Jun 25

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

arXiv:2601. 23075v2 Announce Type: replace Abstract: On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy, and policy updates must be conservative.

By Yuexin Bian, Jie Feng, Tao Wang, Yijiang Li, Sicun Gao, Yuanyuan Shi
arXiv AI
Jul 7

ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning

arXiv:2602. 21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks.

By Xiaoxuan Wang, Han Zhang, Haixin Wang, Yidan Shi, Ruoyan Li, Kaiqiao Han, Chenyi Tong, Haoran Deng, Renliang Sun, Alexander Taylor, Yanqiao Zhu, Jason Cong, Yizhou Sun, Wei Wang