arXiv:2601. 22993v4 Announce Type: replace Abstract: We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems.
By Rohan Tangri, Jan-Peter Calliess
arXiv:2606. 04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep RL may demonstrate susceptibility to transition perturbations that result in unknown or unsafe behaviour.
By Mohit Prashant, Arvind Easwaran
The paper introduces Canary, a risk‑averse reinforcement learning method that optimizes Value‑at‑Risk (VaR) constraints. By applying Cantelli’s inequality, Canary derives a tractable, conservative, and smooth bound on the VaR constraint using only the first two moments of the cost return, yielding a stable constraint estimator even with tight violation thresholds. Extending the trust‑region framework of Constrained Policy Optimization (CPO), the authors provide worst‑case bounds for policy improvement and constraint violation, and empirically demonstrate that Canary reliably satisfies the VaR constraint in every tested environment.
By Rohan Tangri, Jan-Peter Calliess
arXiv:2606. 31653v1 Announce Type: cross Abstract: Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimising upper bounds on the worst-case loss over an allowed perturbation set.
By Matteo Melis, Jesus Martinez Del Rincon, Vishal Sharma
arXiv:2408. 09112v2 Announce Type: replace Abstract: Reinforcement learning policies parametrized by deep neural networks have achieved strong performance for continuous control, yet even small input perturbations may lead to unpredictable behavior.
By Manuel Wendl, Lukas Koller, Tobias Ladner, Matthias Althoff
BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.
By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
The paper introduces CoDRA, a cost-to-disturbance ratio approach for adversarial reinforcement learning that balances controller performance and disturbance exposure without extra penalty terms. CoDRA uses a self‑normalized actor–critic update, scaling value terms by a stop‑gradient normalization constant derived from the current batch. Experiments on MuJoCo pendulum tasks show that CoDRA achieves the lowest cost across a range of forces and masses, outperforming other methods especially on the more challenging InvertedDoublePendulum environment.
By Taeho Lee, Donghwan Lee
arXiv:2602. 03778v2 Announce Type: replace-cross Abstract: Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events.
By Aneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon, Erick Delage
arXiv:2608. 07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ.
By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
arXiv:2606. 19891v1 Announce Type: new Abstract: We study adversarial bandit optimization in which the loss functions may be non-convex and non-smooth.
By Zhuoyu Cheng, Kohei Hatano, Eiji Takimoto
arXiv:2605. 11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories.
By Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz
To improve the real-world applicability of reinforcement learning (RL), the field of adversarially robust RL studies how to train agents under adversarial environment perturbations. In this setting, a protagonist agent optimizes a policy under environmental perturbations from an adversary, resulting in a zero-sum Markov game.