arXiv AI By Kwan Soo Shin, In Seok Kang, Yunkyung Min, Munho Lee

A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss

Read the original on arXiv AI →

arXiv:2607. 09306v2 Announce Type: replace-cross Abstract: Whether a language model behaves as it claims is a judgement on which independent human raters cannot agree (Fleiss kappa = 0.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 19

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.