arXiv AI

Probing Persona-Dependent Preferences in Language Models

The study investigates how large language models (LLMs) encode persona-dependent preferences by training linear probes on residual-stream activations of Gemma‑3‑27B and Qwen‑3.5‑122B. It identifies a genuine preference vector that tracks the model’s task choices across various prompts and shows that steering along this vector can causally control pairwise choices. The research also finds that some preference information transfers between different personas, including a case where an evil persona’s preferences anti‑correlate with those of a helpful assistant.

arXiv Computation and Language
Sep 14

Creating an Atomic User Model for Personality-Aware Large Language Model Interaction

The paper introduces the Atomic User Model (AUM), a structured representation of a user’s personality that separates a stable identity nucleus from four interpretable shells—psychological, cognitive & experiential, behavioural, and social—along with cross-shell entries for conflict and authenticity. It proposes using AUM as a retrieval index rather than a prompt prefix, enabling a task‑specific, budgeted retrieval of relevant fields at generation time. Experiments with simulated participants show that retrieving eight AUM fields improves style fidelity, preference accuracy, and user voice identification compared to flat preference notes, especially benefiting users whose default assistant performs poorly.

By B. Sankar, Deepthika S, Pawni Yadav, Amogh A S
arXiv Computation and Language
3d ago

LLM Persona Unlearning

arXiv:2609.39882v1 Announce Type: new Abstract: Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-...

By Kemou Li, Zhuan Shi, Qizhou Wang, Fengpeng Li, Negar Rostamzadeh, Golnoosh Farnadi, Jiantao Zhou