arXiv AI By Mahshad Rastegarmoghaddam, Davoud Nikkhouy, Shima Samadzadeh

Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control

Read the original on arXiv AI →

arXiv:2608. 04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.

By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
arXiv Machine Learning
Sep 1

Uncertainty-Driven Replay Memory for Reinforcement Learning

The paper introduces Uncertainty-Driven Replay Memory (UDRM), a new experience replay buffer for reinforcement learning that prioritizes storing transitions with high uncertainty estimates. Unlike traditional buffers that rely on temporal difference error or transition distributions, UDRM updates its contents based on uncertainty derived from the RL model during training. Experiments show that this uncertainty-aware buffer leads to higher rewards during training compared to other uncertainty-aware RL frameworks.

By Sheeraja Rajakrishnan, Alexander G. Ororbia, Travis Desell, Daniel E. Krutz
arXiv Machine Learning
Aug 28

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.

By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu