arXiv:2608.11025v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies.
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2601. 02896v3 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.
By Harshvardhan Saini, Yiming Tang, Dianbo Liu
The article surveys how large language models (LLMs) are being applied to mental health, outlining a three‑phase evolution: Phase I uses LLMs as passive information tools and pattern recognizers for assessment; Phase II employs them as empathetic conversationalists for stateless, in‑the‑moment interactions; Phase III aims to create longitudinal, personalized companions that act as stateful cognitive agents. It systematically reviews core technologies, agent architectures (Profile, Memory, Reasoning, Planning), and the datasets and benchmarks that support this progression, offering a coherent narrative and roadmap for future research. The survey also provides a curated resource list at https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.
By He Hu, Yucheng Zhou, Qianning Wang, Yingjian Zou, Chiyuan Ma, Juzheng Si, Jianzhuang Liu, Zitong Yu, Laizhong Cui, Fei Ma, Qi Tian
arXiv:2608.29118v1 Announce Type: new
Abstract: Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM...
By Mingxuan Li, Qirun Dai, Heran Wang, Chenhao Tan