The paper introduces the Situated Identity Test (SIT), a framework that assesses whether a language model’s behavior can be traced to a specific developmental lineage rather than merely imitating a persona. SIT requires agents to possess accurate knowledge of their recorded experiences while appropriately ignoring ungrounded information, and it demonstrates that policies based only on compressed profiles are limited in distinguishing between colliding life histories. The authors present SITBench, an evaluation suite with 25 profile‑collision pairs and 10,000 probes across nine model architectures, and provide open‑source tools and pilot results on state‑of‑the‑art foundation models.
By Jun He, Deying Yu
The study investigates how the visibility of speaker biographies to interlocutors during training, inference, and evaluation affects persona-based dialogue generation. It finds that training-time visibility is the primary factor determining whether models express persona traits or simply copy biographical text, and that providing interlocutor-biography visibility during training reduces target-biography copying. Additionally, asymmetric disclosure—where only the interlocutor sees the target biography—leads to more frequent leakage of target content into interlocutor turns, making such dialogues easier for a judge to identify.
By Daniela Occhipinti, Malvina Nissim, Marco Guerini
arXiv:2606. 23700v1 Announce Type: cross Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content.
By Arush Tagade, Shaoheng Zhou, Jiaxin Wen, Shi Feng
arXiv:2608.11025v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
arXiv:2606. 11502v1 Announce Type: cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.
By Benjamin Sturgeon, David Africa, Sid Black
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies.
arXiv:2608.29590v1 Announce Type: new
Abstract: We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely...
By Yusuke Hirota, Michael Ross Boone, Arun George Zachariah, Jibin Rajan Varghese, Yu-Chiang Frank Wang, Boyi Li, Ryo Hachiuma
arXiv:2608. 08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions.
By Peiyang Liu, Xi Wang, Ziqiang Cui, Di Liang, Wei Ye
The paper examines how large language models (LLMs) tend to overuse persona attributes in persona-based dialogue generation, producing unnatural responses. It identifies a systematic bias in LLMs to incorporate all provided persona details and shows that current metrics cannot assess contextual appropriateness. To address this, the authors introduce Self-CONtrastive Persona Overuse Suppression (SCONPOS), which intervenes in the prompt encoding stage to reduce overuse, and propose the Persona Appropriateness Score (PAS), a new metric that penalizes both overuse and underuse of persona attributes.
By Jongkyung Shin, Inkyu Lee, Chiehyeon Lim
arXiv:2604.17031v3 Announce Type: replace
Abstract: The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds. We approach th...
By Pierre Beckmann, Patrick Butlin
arXiv:2606. 00545v1 Announce Type: new Abstract: Post-trained language models can recognize their own outputs from a sentence or two out of context.
By Asvin G
arXiv:2605. 09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors".
By Nils A. Herrmann, Leander Girrbach, Kirill Bykov, Zeynep Akata