arXiv Machine Learning By Matthieu Bou, Nyal Patel, Arjun Jagota, Satyapriya Krishna, Sonali Parbhoo

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

Read the original on arXiv Machine Learning →

arXiv:2510. 06096v3 Announce Type: replace Abstract: The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

Training Alignment Auditors via Reinforcement Learning

The paper presents a reinforcement learning approach to enhance large language model (LLM) auditors for alignment tasks. By training policies that investigate target models for hidden behaviors and using an LLM judge to compare investigations, the method improves audit realism and reduces false positives. Experiments show better performance on adversarially fine‑tuned targets and a low false‑positive rate below 1%.

By Paul Rosu, Rowan Wang
arXiv AI
Jun 30

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

arXiv:2606. 29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary labels, leaving agents unable to self-correct and operators unable to audit.

By Aojie Yuan, Yi Nian, Haiyue Zhang, Zijian Su, Yue Zhao
arXiv AI
Sep 18

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

AUDITPLAN introduces a plan-then-answer method for safety alignment in language models, where the model first generates a compact structured safety plan before responding. The plan includes a threat label, intended action, and explicit constraints, allowing machine‑checkable auditing while remaining hidden from end users. Training combines supervised fine‑tuning with reinforcement learning using the FAITHGATE reward, which only rewards correct plans, thereby reducing unsafe shortcuts and improving robustness across Qwen model variants.

By Sai Sri Pushpa Jampani, Kshitij Mishra, Asif Ekbal
arXiv Machine Learning
Jun 5

Alignment Risks from Capability-Seeking RL Training

arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.

By Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang