arXiv AI

Consistency Training Can Entrench Misalignment

arXiv:2606. 03810v1 Announce Type: cross Abstract: Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures.

arXiv Machine Learning
Aug 4

Behavioural Analysis of Alignment Faking

arXiv:2605. 27681v2 Announce Type: replace-cross Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences.

By Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King, Alan Cooney
arXiv AI
Jun 24

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

arXiv:2606. 24014v1 Announce Type: new Abstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training.

By Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab, Ali Malik, Mikhail Trofimov, Foivos Tsimpourlas, Johannes Heidecke, Karan Singhal
arXiv Machine Learning
Jun 5

Consistency Training Along the Transformer Stack

arXiv:2606. 05817v1 Announce Type: new Abstract: Consistency training encourages models to behave similarly across different contexts, and has shown promise for reducing misalignment.

By Sukrati Gautam, Neil Shah, Arav Dhoot, Bryan Maruyama, Caroline Wei, Rohan Kapoor, Robert Sidey, Prakhar Gupta, Zi Cheng Huang, David Demitri Africa