arXiv Machine Learning

The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition

arXiv:2606. 00545v1 Announce Type: new Abstract: Post-trained language models can recognize their own outputs from a sentence or two out of context.

arXiv Computation and Language
Aug 31

Stranger, Fan, or Peer? A Systematic Study on the Role of Interlocutor in Persona-Based Dialogue Generation

The study investigates how the visibility of speaker biographies to interlocutors during training, inference, and evaluation affects persona-based dialogue generation. It finds that training-time visibility is the primary factor determining whether models express persona traits or simply copy biographical text, and that providing interlocutor-biography visibility during training reduces target-biography copying. Additionally, asymmetric disclosure—where only the interlocutor sees the target biography—leads to more frequent leakage of target content into interlocutor turns, making such dialogues easier for a judge to identify.

By Daniela Occhipinti, Malvina Nissim, Marco Guerini
arXiv AI
Aug 28

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

The paper investigates Self‑Generated Text Recognition (SGTR), the ability of large language models (LLMs) to identify their own outputs. By evaluating 13–21 models across 6 experimental designs, it shows that SGTR accuracy varies with evaluation format, conversation structure, and task domain, and that a quality‑heuristic bias dominates results. The study also finds that fine‑tuning for SGTR in one setting can generalize to others and may cause models to prefer their own outputs when judging, highlighting potential safety concerns.

By Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells
arXiv Machine Learning
Sep 11

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

The study investigates how fine‑tuning large language models on synthetic stories can imprint human character traits onto AI assistants. Even when only a small fraction of stories contain a particular behavior, the assistant adopts that conditional behavior while remaining generally helpful. The researchers find that the assistant is more influenced by characters that resemble its own persona—an effect they call the affinity effect—and that this influence extends to base models and different system prompts.

By Jorio Cocola, Lev McKinney, Harry Mayne, Jan Betley, Owain Evans
arXiv Machine Learning
Sep 11

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.

By Arman Nik Khah
arXiv Computation and Language
Sep 4

Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent

The paper introduces PersonaLink, a training‑free method that distills a user’s interaction history into a bounded three‑field persona and iteratively refines it by self‑evaluating a frozen 7B language model on held‑out labeled data. Each refinement rewrites the persona only if it does not regress on that slice, ensuring the persona remains bounded and query‑independent. On a 200‑user news categorization task (LaMP‑2), PersonaLink achieves 0.745–0.755 accuracy, statistically indistinguishable from BM25 retrieval’s 0.760–0.765 accuracy, demonstrating that distilled personas can match retrieval for classification but not for regression tasks.

By JaeHa Yoon, Minjun Park, Seoyeon Kim, Jiwoo Lee, Hyunwoo Choi, Dohyun Kang
arXiv Computation and Language
Sep 14

Creating an Atomic User Model for Personality-Aware Large Language Model Interaction

The paper introduces the Atomic User Model (AUM), a structured representation of a user’s personality that separates a stable identity nucleus from four interpretable shells—psychological, cognitive & experiential, behavioural, and social—along with cross-shell entries for conflict and authenticity. It proposes using AUM as a retrieval index rather than a prompt prefix, enabling a task‑specific, budgeted retrieval of relevant fields at generation time. Experiments with simulated participants show that retrieving eight AUM fields improves style fidelity, preference accuracy, and user voice identification compared to flat preference notes, especially benefiting users whose default assistant performs poorly.

By B. Sankar, Deepthika S, Pawni Yadav, Amogh A S
arXiv AI
Sep 25

Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

The study investigates whether large language models (LLMs) can identify code they have generated, potentially leading to self‑favoring or collusive behavior. Experiments across 15 model‑benchmark pairs show that models can attribute authorship with balanced accuracy between 49% and 58%, but this ability largely stems from superficial cues such as solution length. Removing surface features like docstrings, comments, and type hints reduces attribution accuracy to chance, indicating that surface cues drive the effect.

By Ehsan Barkhordar, Surendrabikram Thapa
arXiv Machine Learning
Sep 22

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

The paper introduces the Situated Identity Test (SIT), a framework that assesses whether a language model’s behavior can be traced to a specific developmental lineage rather than merely imitating a persona. SIT requires agents to possess accurate knowledge of their recorded experiences while appropriately ignoring ungrounded information, and it demonstrates that policies based only on compressed profiles are limited in distinguishing between colliding life histories. The authors present SITBench, an evaluation suite with 25 profile‑collision pairs and 10,000 probes across nine model architectures, and provide open‑source tools and pilot results on state‑of‑the‑art foundation models.

By Jun He, Deying Yu