Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithf...
The paper presents a reinforcement learning approach to enhance large language model (LLM) auditors for alignment tasks. By training policies that investigate target models for hidden behaviors and using an LLM judge to compare investigations, the method improves audit realism and reduces false positives. Experiments show better performance on adversarially fine‑tuned targets and a low false‑positive rate below 1%.
By Paul Rosu, Rowan Wang
arXiv:2510. 06096v3 Announce Type: replace Abstract: The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge.
By Matthieu Bou, Nyal Patel, Arjun Jagota, Satyapriya Krishna, Sonali Parbhoo
The paper introduces a protocol for auditing and composing reinforcement‑learning policies using discrete behavioral rules, defining auditability through six testable predicates such as trace integrity and rule coverage. Experiments show that overlapping rule sets do not guarantee behavioral agreement, and that rule‑based fusion often fails to outperform value‑based composition, highlighting limitations in current description layers. The authors provide an evidence‑bounded audit framework and outline future directions for more robust skill composition.
By Liu Hung Ming
arXiv:2608.21803v1 Announce Type: cross
Abstract: As machine learning (ML) models are increasingly deployed in high-stakes environments, explainable AI (XAI) methods like SHAP and LIME have become es...
By Maraz Mia, Shovan Roy, Mir Mehedi A. Pritom, Maanak Gupta
arXiv:2609.36254v1 Announce Type: new
Abstract: Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning bef...
By Xiangyu Zhou, Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov, Dongxiao Zhu