arXiv Computation and Language

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

Hugging Face Trending Papers
Aug 19

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.

arXiv AI
Sep 7

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli