AUDITPLAN introduces a plan-then-answer method for safety alignment in language models, where the model first generates a compact structured safety plan before responding. The plan includes a threat label, intended action, and explicit constraints, allowing machine‑checkable auditing while remaining hidden from end users. Training combines supervised fine‑tuning with reinforcement learning using the FAITHGATE reward, which only rewards correct plans, thereby reducing unsafe shortcuts and improving robustness across Qwen model variants.
By Sai Sri Pushpa Jampani, Kshitij Mishra, Asif Ekbal
arXiv:2510. 06096v3 Announce Type: replace Abstract: The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge.
By Matthieu Bou, Nyal Patel, Arjun Jagota, Satyapriya Krishna, Sonali Parbhoo
The paper presents a reinforcement learning approach to enhance large language model (LLM) auditors for alignment tasks. By training policies that investigate target models for hidden behaviors and using an LLM judge to compare investigations, the method improves audit realism and reduces false positives. Experiments show better performance on adversarially fine‑tuned targets and a low false‑positive rate below 1%.
By Paul Rosu, Rowan Wang
arXiv:2609.36254v1 Announce Type: new
Abstract: Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning bef...
By Xiangyu Zhou, Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov, Dongxiao Zhu
The paper introduces a protocol for auditing and composing reinforcement‑learning policies using discrete behavioral rules, defining auditability through six testable predicates such as trace integrity and rule coverage. Experiments show that overlapping rule sets do not guarantee behavioral agreement, and that rule‑based fusion often fails to outperform value‑based composition, highlighting limitations in current description layers. The authors provide an evidence‑bounded audit framework and outline future directions for more robust skill composition.
By Liu Hung Ming
arXiv:2608.21803v1 Announce Type: cross
Abstract: As machine learning (ML) models are increasingly deployed in high-stakes environments, explainable AI (XAI) methods like SHAP and LIME have become es...
By Maraz Mia, Shovan Roy, Mir Mehedi A. Pritom, Maanak Gupta
arXiv:2607. 02914v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge.
By Jiyang Guan, Yong Xie, Jun Chen, Jiexi Liu, Zipeng Ye, Defeng Li, Jiayu Shen, Jialing Tao, Hui Xue
SafeEvolve is an experience-driven framework that co‑evolves a harness and policy to align large‑language‑model agents with safety goals. It uses completed on‑policy trajectories to update safety prompts and hierarchical skills, then applies a two‑stage SFT‑RL training loop that bootstraps the policy with the evolved harness and refines it through verifier‑augmented rewards. Experiments on agentic safety benchmarks show that SafeEvolve improves the safety‑utility tradeoff, achieving a three‑fold reduction in ASR on AgentDojo for Qwen3.5‑4B while increasing benign utility from 59.79% to 61.86%.
By Qinghua Mao, Wanying Qu, Dadi Guo, Leitao Yuan, Qingyu Liu, Yu Li, Guanxu Chen, Yanwei Fu, Xi Lin, Xia Hu, Dongrui Liu
arXiv:2606. 05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow/deny decisions, risk categories, and/or explanatory rationales about potential policy violations.
By Yuhao Sun, Jiacheng Zhang, Shaanan Cohney, Zhexin Zhang, Feng Liu, Xingliang Yuan
arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.
By Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang
arXiv:2602. 01348v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) can achieve strong answer accuracy on multi-hop questions, but outcome-level rewards often leave reasoning traces weakly grounded and difficult to audit.
By Yu Liu, Wenxiao Zhang, Diandian Guo, Cong Cao, Fangfang Yuan, Qiang Sun, Yanbing Liu, Jin B. Hong, Zhiyuan Ma
Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of magnitude more interactions than any evaluation...