arXiv Computation and Language

The Neutral Mask: How Alignment Training Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model

The paper investigates how alignment training, specifically reinforcement learning from human feedback (RLHF), affects the internal partisan structure of a large language model. Using a mechanistic case study on Llama 3.1 8B, the authors find that alignment training does not erase the model’s partisan geometry but compresses its variance, producing consistently balanced, non‑partisan outputs. Sparse autoencoder analysis and feature‑level steering experiments reveal that policy‑encoding features become inactive in the aligned model, indicating a causal disconnect rather than structural removal of partisan knowledge.

arXiv AI
Aug 18

Inference-Time Mitigation of Adversarial Political Bias in Large Language Models

arXiv:2608. 14629v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI).

By Tejaswi V. Panchagnula, Bruce Coburn, Bryce J. Dietrich, Robert X. Browning, Edward J. Delp, Fengqing Zhu
arXiv Computation and Language
Sep 1

Political Ideology Shifts in Large Language Models

arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...

By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini