arXiv:2607. 07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them.
By Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk, Irakli Shalibashvili, Cl\'ement Dumas, Konstantinos Voudouris, David Demitri Africa
arXiv:2607. 13162v1 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone.
By Winston Zeng, Ali Emami, Jinho Choi
arXiv:2608.11025v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies.
arXiv:2607. 24435v1 Announce Type: cross Abstract: Large language models may easily assign personality labels from text, but model interpretability remains an open problem.
By Brittany Harbison, Ashok K. Goel
arXiv:2609.37914v1 Announce Type: cross
Abstract: Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a pheno...
By Gon\c{c}alo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman
arXiv:2606. 23700v1 Announce Type: cross Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content.
By Arush Tagade, Shaoheng Zhou, Jiaxin Wen, Shi Feng
arXiv:2606. 05183v1 Announce Type: cross Abstract: Large language models are increasingly deployed as high-stakes advisors, yet standard alignment benchmarks treat sycophancy as a binary failure mode.
By Patrick Keough
arXiv:2609.15998v1 Announce Type: new
Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property...
By Tabia Tanzin Prama, Calla Glavin Beauregard, Christopher M. Danforth, Peter Sheridan Dodds
arXiv:2606. 09843v3 Announce Type: replace-cross Abstract: Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behave.
By Juan Manuel Contreras
arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.
By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
arXiv:2605.29791v2 Announce Type: replace
Abstract: While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions...
By Yutong Yang, Chenxi Miao, Weikang Li, Yunfang Wu