arXiv Machine Learning

Alignment Risks from Capability-Seeking RL Training

arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.

arXiv Machine Learning
Jul 7

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

arXiv:2506. 07468v4 Announce Type: replace Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities.

By Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, Natasha Jaques
arXiv AI
3d ago

Alignment via Training Against Probes Without Losing Monitorability

The paper proposes probe-guided fine-tuning, a method that uses probes detecting undesired properties in model activations as a direct training signal. Experiments show that continuously updated probes reduce harmfulness and improve honesty while preserving utility, outperforming DPO and inference-time steering in safety-utility trade-offs and robustness to jailbreak and abliteration attacks. Importantly, the concepts remain linearly encoded after fine-tuning, maintaining monitorability.

By Lena Libon, Alexander Panfilov, Ben Rank, Xin Chen, Jonas Geiping, Maksym Andriushchenko
arXiv AI
Jun 24

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

arXiv:2606. 24014v1 Announce Type: new Abstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training.

By Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab, Ali Malik, Mikhail Trofimov, Foivos Tsimpourlas, Johannes Heidecke, Karan Singhal