arXiv Computation and Language

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

The paper introduces a diagnostic framework to disentangle the effects of character profile axes—Familiarity, Structure, and Disposition—on large language model role‑playing agents. Experiments on 211 personas and five LLMs show that Familiarity and Structure have little impact, whereas Disposition, particularly immoral traits, consistently degrades performance. The authors propose Field‑Aware Contrastive Decoding (FACD), a training‑free method that mitigates this performance gap without harming moral‑character performance.

arXiv AI
Aug 14

Synthetic Persona Pretraining: Alignment from Token Zero

arXiv:2608. 13482v1 Announce Type: cross Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical.

By Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, Roland Aydin, Robert West
arXiv AI
Jun 2

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

arXiv:2606. 01552v1 Announce Type: new Abstract: Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate surface-level fidelity and offer limited insight into decision making under role-alignment value conflicts.

By Huayi Lai, Shichao Song, Simin Niu, Hanyu Wang, Jiawei Yang, Zhouxing Wang, Zhiqiang Yin, Xun Liang
arXiv Computation and Language
Sep 23

PERSONAWEAVER: Controllable Diversity Beyond Conventional Archetypes in Procedural Character Generation

PERSONAWEAVER is a new approach to procedural character generation that separates world building from behavioral specification, using manually curated banks of moral positions and conversational reactions to diversify character behavior. By applying this method across ten realistic and fantastical settings and three large language models, the system produces broader moral and interactional response distributions, varied interpersonal language, response length, sentiment, and less archetypal world attribute combinations compared to prior work.

By Maan Qraitem, Kate Saenko, Bryan A. Plummer
arXiv Computation and Language
3d ago

LLM Persona Unlearning

arXiv:2609.39882v1 Announce Type: new Abstract: Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-...

By Kemou Li, Zhuan Shi, Qizhou Wang, Fengpeng Li, Negar Rostamzadeh, Golnoosh Farnadi, Jiantao Zhou
arXiv AI
Jul 10

Persona Cartography: Charting Language Model Personality Traits in Weight Space

arXiv:2607. 07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them.

By Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk, Irakli Shalibashvili, Cl\'ement Dumas, Konstantinos Voudouris, David Demitri Africa
arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu