arXiv:2606. 23700v1 Announce Type: cross Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content.
By Arush Tagade, Shaoheng Zhou, Jiaxin Wen, Shi Feng
arXiv:2608.11025v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
arXiv:2607. 09349v1 Announce Type: cross Abstract: Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents.
By Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies.
arXiv:2607. 04510v1 Announce Type: cross Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.
By Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)
arXiv:2605.30381v2 Announce Type: replace-cross
Abstract: When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverabl...
By Vahideh Zolfaghari
arXiv:2607. 01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable.
By Tung-Ling Li, Hongliang Liu, Yuhao Wu
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2607. 18639v1 Announce Type: new Abstract: Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.
By Seunghyun Lee, Dongyoon Han, Sangdoo Yun
arXiv:2608. 04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence.
By Kotaro Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto, Yuji Naraki, Ryotaro Shimizu, Wenya Wang
arXiv:2606. 07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization.
By Mahdi Alkaeed
arXiv:2608.29118v1 Announce Type: new
Abstract: Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM...
By Mingxuan Li, Qirun Dai, Heran Wang, Chenhao Tan