The paper introduces Delta, a two‑phase framework for testing deep reinforcement learning agents. In the first phase, the agent under test is evaluated for catastrophic failures while collecting decision‑making data. The second phase trains a challenger agent from this data using offline RL; comparing the challenger’s rewards to the original agent reveals optimality bugs, and Delta successfully uncovered thousands of such issues across multiple environments.
By Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng, David Lo
The paper proposes EvalXRL, a benchmark that evaluates explainable reinforcement learning (XRL) methods by measuring how well their explanations help a large language model coding agent diagnose and fix bugs in RL agents. Unlike current metrics that focus on faithfulness or human ratings, EvalXRL uses a closed‑loop, scientific‑method style interaction where the coding agent repeatedly invokes XRL methods, refines hypotheses, and repairs the agent, scoring each method by the resulting RL reward. This approach enables a head‑to‑head comparison of multiple XRL techniques in realistic debugging scenarios.
By Ram Rachum, Yotam Amitai, B\'alint Gyevn\'ar, Reuth Mirsky, Cameron Allen
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of...
LEGO-RL is a framework that connects native coding-agent harnesses with scalable policy‑gradient training without altering the harnesses’ internal flow. It achieves faithful optimization through in‑process LLM proxying, reliable execution via sandbox orchestration, and observable training with automated validation and a Live UI. Experiments show LEGO‑RL improves the Qwen3.5‑35B‑A3B model’s performance on three native harnesses while preserving high rollout‑training probability correlation.
By Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai
The paper introduces new evaluation metrics for safe reinforcement learning that go beyond average safety guarantees by examining how often and how severely safety bounds are violated, consistency across tasks and bounds, and the relationship between training-time and final policy behavior. It also proposes a safety tier system for categorizing algorithms and presents empirical safety evaluations on multiple navigation tasks. The authors recommend reporting aggregate metrics, distributional data, and task‑specific results together, and provide an open‑source suite, SafeRLEval, to facilitate reliable safety assessment.
By Lindsay Spoor, Aske Plaat, Thomas Moerland
arXiv:2607. 01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks.
By Yunhao Feng, Ruixiao Lin, Ming Wen, Qinqin He, Yanming Guo, Yifan Ding, Yutao Wu, Jialuo Chen, Yunhao Chen, Xiaohu Du, Jianan Ma, Zixing Chen, Zhuoer Xu, Xingjun Ma, Xinhao Deng
arXiv:2608. 16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.
By Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
Teach-to-Crash is a closed‑loop testing framework that uses a dual‑LLM architecture to generate collision‑inducing scenarios for autonomous driving systems. A high‑reasoning Teacher LLM controls the search when collision metrics stagnate, while a low‑reasoning Student LLM produces simulator‑executable scenarios in JSON. In a CARLA case study, Teach‑to‑Crash achieved the highest collision hit rate (90.79 %), the shortest mean time‑to‑collision (18.31 s), and superior diversity and avoidability metrics compared to other methods.
By Zaid Ghazal, Khouloud Gaaloul, Bruce Maxim
arXiv:2604. 02022v4 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses.
By Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, Yanwei Fu, Yujiu Yang, Jing Shao, Xia Hu, Dongrui Liu
The paper introduces TRACE, a digital‑advertising diagnostic environment that uses simulated interventions to generate verifiable rewards for training reasoning agents. By injecting controlled interventions into a simulator, the hidden cause of anomalies becomes an oracle label, enabling agents to learn to identify root causes and affected segments through noisy, confounded evidence. Experiments show that reinforcement learning with these synthesized rewards outperforms large prompted baselines, achieving higher accuracy while using fewer tool calls.
By Rui Sun, Zhan Shi, Bing He
arXiv:2607. 07029v1 Announce Type: cross Abstract: Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks.
By Dennis Gross, Quentin Mazouni, Helge Spieker, Arnaud Gotlieb
arXiv:2608. 04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied.
By Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi, Sanggeon Yun, Hyunwoo Oh, SungHeon Jeong, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani