Auditing Alignment Controllability in LLMs via Political Axes
arXiv:2607. 23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass.
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered.
arXiv:2607. 23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass.
arXiv:2609.23039v1 Announce Type: new Abstract: LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their...
arXiv:2608.29198v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored assistance, measuring their political alignment...
arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...
As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that interactions with LLMs can influence users' political attitudes and choices, raising questions about how these models themselves evaluate political actors.
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2609.15849v1 Announce Type: cross Abstract: Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM per...
arXiv:2609.15207v1 Announce Type: new Abstract: Generative AI writing assistants and the Large Language Models (LLMs) that power them are increasingly part of how voters gather information before ele...
arXiv:2607. 20487v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to answer questions about political information, including in election-adjacent information settings where factual errors and ideological distortions are high-stakes.
arXiv:2606. 28335v1 Announce Type: cross Abstract: We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space.
The paper investigates how alignment training, specifically reinforcement learning from human feedback (RLHF), affects the internal partisan structure of a large language model. Using a mechanistic case study on Llama 3.1 8B, the authors find that alignment training does not erase the model’s partisan geometry but compresses its variance, producing consistently balanced, non‑partisan outputs. Sparse autoencoder analysis and feature‑level steering experiments reveal that policy‑encoding features become inactive in the aligned model, indicating a causal disconnect rather than structural removal of partisan knowledge.
arXiv:2606. 16127v1 Announce Type: cross Abstract: The worldwide surge of authoritarianism, combined with the increasing central role in users' everyday lives, raises the question of to what extent specific models exhibit or promote authoritarian attitudes and characteristics.