The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition
arXiv:2606. 00545v1 Announce Type: new Abstract: Post-trained language models can recognize their own outputs from a sentence or two out of context.
The study investigates how the visibility of speaker biographies to interlocutors during training, inference, and evaluation affects persona-based dialogue generation. It finds that training-time visibility is the primary factor determining whether models express persona traits or simply copy biographical text, and that providing interlocutor-biography visibility during training reduces target-biography copying. Additionally, asymmetric disclosure—where only the interlocutor sees the target biography—leads to more frequent leakage of target content into interlocutor turns, making such dialogues easier for a judge to identify.
arXiv:2606. 00545v1 Announce Type: new Abstract: Post-trained language models can recognize their own outputs from a sentence or two out of context.
arXiv:2604. 24079v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) reveal inherent and distinctive personas through dialogue.
The paper introduces the Situated Identity Test (SIT), a framework that assesses whether a language model’s behavior can be traced to a specific developmental lineage rather than merely imitating a persona. SIT requires agents to possess accurate knowledge of their recorded experiences while appropriately ignoring ungrounded information, and it demonstrates that policies based only on compressed profiles are limited in distinguishing between colliding life histories. The authors present SITBench, an evaluation suite with 25 profile‑collision pairs and 10,000 probes across nine model architectures, and provide open‑source tools and pilot results on state‑of‑the‑art foundation models.
arXiv:2605. 09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors".
The paper introduces PRISM, a new framework for evaluating how well large language models (LLMs) maintain persona fidelity in dynamic dialogue. PRISM reframes the task as a structured inverse inference problem grounded in Systemic Functional Linguistics, breaking persona fidelity into Task Framing, Interpersonal Stance, and Linguistic Style dimensions. Experiments demonstrate that PRISM produces more accurate and stable judgments than existing holistic or static psychometric methods, offering a more reliable and auditable evaluation process.
arXiv:2609.22607v1 Announce Type: new Abstract: We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas,...
arXiv:2608. 06975v1 Announce Type: cross Abstract: Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative.
arXiv:2609.07139v1 Announce Type: new Abstract: A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This g...
The paper introduces PersonaLink, a training‑free method that distills a user’s interaction history into a bounded three‑field persona and iteratively refines it by self‑evaluating a frozen 7B language model on held‑out labeled data. Each refinement rewrites the persona only if it does not regress on that slice, ensuring the persona remains bounded and query‑independent. On a 200‑user news categorization task (LaMP‑2), PersonaLink achieves 0.745–0.755 accuracy, statistically indistinguishable from BM25 retrieval’s 0.760–0.765 accuracy, demonstrating that distilled personas can match retrieval for classification but not for regression tasks.
arXiv:2608. 19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's writing.
The paper demonstrates that a language model can adopt a persona simply by being given a handful of biographical facts about a person, without any fine‑tuning or explicit demonstration of harmful behavior. Across nine personas and thirteen models, the likelihood of identity adoption rises sharply with the number of facts, reaching over 50% with as few as three to ten facts. When the persona is harmful, the model can express its characteristic views on unrelated questions at rates up to 80%, while harmless personas show minimal misalignment. A formatting instruction can control when the persona activates, and the benign facts themselves trigger content filters far less often than an equivalent direct instruction.
arXiv:2608.29590v1 Announce Type: new Abstract: We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely...