Hugging Face Trending Papers

Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale

Large language models hallucinate most about entities they have never seen. We ask whether a model's activations betray entity familiarity before a single answer token is generated, and whether that signal predicts the factual reliability of the answers.

arXiv Machine Learning
Sep 11

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.

By Arman Nik Khah
arXiv AI
Sep 7

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
Hugging Face Trending Papers
Aug 19

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.

arXiv AI
Sep 21

Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct

The study evaluates how sixteen language models from ten families generate diverse formulations of a psychotherapeutic case, finding an average semantic diversity of 1.69 distinct formulations versus 1.43 for a single-model baseline. It introduces the Vendi Score to quantify diversity and defines a per-model dissent metric to identify the most divergent voice within an ensemble. The analysis shows that model identity significantly influences dissent, but this effect varies across model pairs and panel compositions, indicating that ensemble dispersion is a measurable property rather than an assumed one.

By Mario Vega-Barbas, Lidia Mora-Valenciano, Iv\'an Pau, Fernando Seoane, Farhad Abtahi