arXiv Machine Learning

Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization

arXiv Machine Learning
Sep 14

A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning

The paper presents a unified framework for regularization-based robust reinforcement learning by deriving upper bounds on the performance gap between nominal and worst-case policies. These bounds are expressed as a regularization objective plus a KL-divergence penalty, explaining why KL penalties enhance robustness. The authors reformulate robust training as a constrained optimization problem, updating the Lagrange multiplier jointly with the policy to automatically tune regularization, and validate the approach with extensive adversarial evaluations on continuous control tasks.

By Amine Andam, Jamal Bentahar, Mustapha Hedabou
arXiv AI
Sep 21

Taming the Adversary: A Cost-to-Disturbance Ratio Approach to Adversarial Reinforcement Learning

The paper introduces CoDRA, a cost-to-disturbance ratio approach for adversarial reinforcement learning that balances controller performance and disturbance exposure without extra penalty terms. CoDRA uses a self‑normalized actor–critic update, scaling value terms by a stop‑gradient normalization constant derived from the current batch. Experiments on MuJoCo pendulum tasks show that CoDRA achieves the lowest cost across a range of forces and masses, outperforming other methods especially on the more challenging InvertedDoublePendulum environment.

By Taeho Lee, Donghwan Lee
Hugging Face Trending Papers
4d ago

Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

The paper introduces Q-learning Penalized Transformer (QPT), a training–inference consistent framework for safe offline reinforcement learning. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost while incorporating a Q-shaped penalty to balance safety, reward maximization, and behavior regularization. The method consistently outperforms strong baselines on 38 DSRL benchmark tasks and adapts robustly to varying constraint thresholds.