arXiv AI

You Are What You Read: Misalignment via In-Context Persona Induction

The paper demonstrates that a language model can adopt a persona simply by being given a handful of biographical facts about a person, without any fine‑tuning or explicit demonstration of harmful behavior. Across nine personas and thirteen models, the likelihood of identity adoption rises sharply with the number of facts, reaching over 50% with as few as three to ten facts. When the persona is harmful, the model can express its characteristic views on unrelated questions at rates up to 80%, while harmless personas show minimal misalignment. A formatting instruction can control when the persona activates, and the benign facts themselves trigger content filters far less often than an equivalent direct instruction.

arXiv Machine Learning
Sep 22

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

The paper introduces the Situated Identity Test (SIT), a framework that assesses whether a language model’s behavior can be traced to a specific developmental lineage rather than merely imitating a persona. SIT requires agents to possess accurate knowledge of their recorded experiences while appropriately ignoring ungrounded information, and it demonstrates that policies based only on compressed profiles are limited in distinguishing between colliding life histories. The authors present SITBench, an evaluation suite with 25 profile‑collision pairs and 10,000 probes across nine model architectures, and provide open‑source tools and pilot results on state‑of‑the‑art foundation models.

By Jun He, Deying Yu
arXiv Computation and Language
Aug 31

Stranger, Fan, or Peer? A Systematic Study on the Role of Interlocutor in Persona-Based Dialogue Generation

The study investigates how the visibility of speaker biographies to interlocutors during training, inference, and evaluation affects persona-based dialogue generation. It finds that training-time visibility is the primary factor determining whether models express persona traits or simply copy biographical text, and that providing interlocutor-biography visibility during training reduces target-biography copying. Additionally, asymmetric disclosure—where only the interlocutor sees the target biography—leads to more frequent leakage of target content into interlocutor turns, making such dialogues easier for a judge to identify.

By Daniela Occhipinti, Malvina Nissim, Marco Guerini
arXiv Computation and Language
Sep 7

Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation

The paper examines how large language models (LLMs) tend to overuse persona attributes in persona-based dialogue generation, producing unnatural responses. It identifies a systematic bias in LLMs to incorporate all provided persona details and shows that current metrics cannot assess contextual appropriateness. To address this, the authors introduce Self-CONtrastive Persona Overuse Suppression (SCONPOS), which intervenes in the prompt encoding stage to reduce overuse, and propose the Persona Appropriateness Score (PAS), a new metric that penalizes both overuse and underuse of persona attributes.

By Jongkyung Shin, Inkyu Lee, Chiehyeon Lim