arXiv AI

Localizing Persona Representations in LLMs

arXiv:2505. 24539v4 Announce Type: replace-cross Abstract: We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language models (LLMs).

arXiv Computation and Language
Sep 1

Political Ideology Shifts in Large Language Models

arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...

By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini
arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
arXiv AI
Jun 30

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

arXiv:2606. 28335v1 Announce Type: cross Abstract: We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space.

By Adib Sakhawat, Syed Rifat Raiyan, Tahsin Islam, Takia Farhin, Hasan Mahmud, Md Kamrul Hasan
arXiv AI
Sep 18

Xeno-Interpretability: Investigating the Alien Minds of LLMs

The paper introduces the concept of xeno-interpretability, which studies internal distinctions in large language models that lack corresponding human concepts. It distinguishes between human‑interpretable and xeno‑semantic spaces, showing that LLMs possess a far larger internal representational space than can be captured by finite human descriptions. The authors propose an empirical program to identify and characterize these xeno‑representations, noting their potential to influence model behavior in ways that are not fully visible through human‑readable communication.

By F. Pierucci, M. Bracale Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, P. Bisconti
arXiv AI
Jun 4

Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks

arXiv:2601. 22396v2 Announce Type: replace-cross Abstract: Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain.

By Candida M. Greco, Lucio La Cava, Andrea Tagarelli
arXiv Computation and Language
Sep 23

PERSONAWEAVER: Controllable Diversity Beyond Conventional Archetypes in Procedural Character Generation

PERSONAWEAVER is a new approach to procedural character generation that separates world building from behavioral specification, using manually curated banks of moral positions and conversational reactions to diversify character behavior. By applying this method across ten realistic and fantastical settings and three large language models, the system produces broader moral and interactional response distributions, varied interpersonal language, response length, sentiment, and less archetypal world attribute combinations compared to prior work.

By Maan Qraitem, Kate Saenko, Bryan A. Plummer