arXiv AI
Sep 18

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

AUDITPLAN introduces a plan-then-answer method for safety alignment in language models, where the model first generates a compact structured safety plan before responding. The plan includes a threat label, intended action, and explicit constraints, allowing machine‑checkable auditing while remaining hidden from end users. Training combines supervised fine‑tuning with reinforcement learning using the FAITHGATE reward, which only rewards correct plans, thereby reducing unsafe shortcuts and improving robustness across Qwen model variants.

By Sai Sri Pushpa Jampani, Kshitij Mishra, Asif Ekbal
arXiv Machine Learning
Aug 27

Training Alignment Auditors via Reinforcement Learning

The paper presents a reinforcement learning approach to enhance large language model (LLM) auditors for alignment tasks. By training policies that investigate target models for hidden behaviors and using an LLM judge to compare investigations, the method improves audit realism and reduces false positives. Experiments show better performance on adversarially fine‑tuned targets and a low false‑positive rate below 1%.

By Paul Rosu, Rowan Wang
arXiv AI
6d ago

Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning

The paper introduces a protocol for auditing and composing reinforcement‑learning policies using discrete behavioral rules, defining auditability through six testable predicates such as trace integrity and rule coverage. Experiments show that overlapping rule sets do not guarantee behavioral agreement, and that rule‑based fusion often fails to outperform value‑based composition, highlighting limitations in current description layers. The authors provide an evidence‑bounded audit framework and outline future directions for more robust skill composition.

By Liu Hung Ming