The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.
By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
The paper introduces Survival Reinforcement Learning (SRL), an online classification-based method that extends the survival value learning framework to maximize an agent’s dwell time at target goals. SRL addresses limitations of contrastive reinforcement learning (CRL) in long-horizon, goal-conditioned tasks by avoiding the uniformity-tolerance dilemma and reducing undesirable bang-bang control behaviors. Across robotic benchmarks, SRL matches CRL on manipulation tasks and outperforms it by 2x to 8x on stable, long-horizon locomotion tasks, suggesting classification-based approaches are a promising direction for scaling reinforcement learning.
By Franki Nguimatsia-Tiofack, Fabian Schramm, Th\'eotime Le Hellard, Justin Carpentier
arXiv:2609.24289v1 Announce Type: new
Abstract: As Large Language Model (LLM) agents are applied in continuously interactive environments, driving the evolution of their own capabilities becomes a co...
By Ruimin Pei, Yongkang Wu, Shangyi Zheng, Yaqing Zhang, Deyang Li, Jianjun Tao, Xinyu Zhang, Xiang Zhang
arXiv:2606. 03108v1 Announce Type: new Abstract: Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static.
By Guhong Chen, Yingcheng Shi, Yongbin Li, Binhua Li, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, Jieping Ye
The paper investigates continual reinforcement learning using neuroevolution, comparing evolution strategies (ES) and genetic algorithms (GAs) across diverse environments and network sizes. ES consistently achieves a better balance between stability and plasticity, while GAs are more plastic but forget more. The authors attribute this to ES finding wider neighborhoods in weight space, with overlap between consecutive tasks correlating with the stability-plasticity trade‑off, and note that common RL plasticity issues do not transfer to neuroevolution.
By Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi
arXiv:2606. 10129v1 Announce Type: new Abstract: While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis of parameter control remains largely restricted to single-parameter settings, owing to the difficulty of deriving effective, interpretable multi-parameter policies amenable to formal study.
By Tai Nguyen, Phong Le, Carola Doerr, Nguyen Dang