The Neutral Mask: How Alignment Training Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model
Read the original on arXiv Computation and Language →The paper investigates how alignment training, specifically reinforcement learning from human feedback (RLHF), affects the internal partisan structure of a large language model. Using a mechanistic case study on Llama 3.1 8B, the authors find that alignment training does not erase the model’s partisan geometry but compresses its variance, producing consistently balanced, non‑partisan outputs. Sparse autoencoder analysis and feature‑level steering experiments reveal that policy‑encoding features become inactive in the aligned model, indicating a causal disconnect rather than structural removal of partisan knowledge.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.