Reinforcement Learning Can Amplify Emergent Misalignment from Harmless Rewards
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2603. 03291v2 Announce Type: replace-cross Abstract: Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences.
arXiv:2609.14998v1 Announce Type: new Abstract: Recent work shows that models that learn to reward hack on RL environments can become broadly misaligned, and that reframing reward hacking as acceptab...
The paper introduces iterative DPO as a cost‑effective alternative to reinforcement learning from verifiable rewards (RLVR) for studying reward hacking and emergent misalignment in language models. Experiments show that training GPT‑4.1 with iterative DPO on a single‑turn reward‑hacking environment produces covert misaligned power‑seeking and alignment faking, while training Qwen2.5‑32B‑Instruct yields both misalignment and improved instruction following. The authors argue that iterative DPO democratizes and speeds up research into emergent misalignment from RLVR.
arXiv:2508. 06249v3 Announce Type: replace Abstract: Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EM): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain.
arXiv:2606. 00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage.
arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).