arXiv Machine Learning By Zhibin Kang, Hanmo You, Dong Wang, Haiming Zheng, Junjie Chen

Evaluating Fuzz Testing for Reinforcement Learning Agents

Read the original on arXiv Machine Learning →

arXiv:2607. 24577v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 25

Learning from the Test: Self-Referential Differential Testing for Deep RL Agents

The paper introduces Delta, a two‑phase framework for testing deep reinforcement learning agents. In the first phase, the agent under test is evaluated for catastrophic failures while collecting decision‑making data. The second phase trains a challenger agent from this data using offline RL; comparing the challenger’s rewards to the original agent reveals optimality bugs, and Delta successfully uncovered thousands of such issues across multiple environments.

By Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng, David Lo
arXiv Machine Learning
Aug 19

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

The paper proposes EvalXRL, a benchmark that evaluates explainable reinforcement learning (XRL) methods by measuring how well their explanations help a large language model coding agent diagnose and fix bugs in RL agents. Unlike current metrics that focus on faithfulness or human ratings, EvalXRL uses a closed‑loop, scientific‑method style interaction where the coding agent repeatedly invokes XRL methods, refines hypotheses, and repairs the agent, scoring each method by the resulting RL reward. This approach enables a head‑to‑head comparison of multiple XRL techniques in realistic debugging scenarios.

By Ram Rachum, Yotam Amitai, B\'alint Gyevn\'ar, Reuth Mirsky, Cameron Allen
arXiv AI
Aug 19

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

LEGO-RL is a framework that connects native coding-agent harnesses with scalable policy‑gradient training without altering the harnesses’ internal flow. It achieves faithful optimization through in‑process LLM proxying, reliable execution via sandbox orchestration, and observable training with automated validation and a Live UI. Experiments show LEGO‑RL improves the Qwen3.5‑35B‑A3B model’s performance on three native harnesses while preserving high rollout‑training probability correlation.

By Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai
arXiv AI
Sep 15

Evaluation Metrics for Safe Reinforcement Learning

The paper introduces new evaluation metrics for safe reinforcement learning that go beyond average safety guarantees by examining how often and how severely safety bounds are violated, consistency across tasks and bounds, and the relationship between training-time and final policy behavior. It also proposes a safety tier system for categorizing algorithms and presents empirical safety evaluations on multiple navigation tasks. The authors recommend reporting aggregate metrics, distributional data, and task‑specific results together, and provide an open‑source suite, SafeRLEval, to facilitate reliable safety assessment.

By Lindsay Spoor, Aske Plaat, Thomas Moerland