Hugging Face Trending Papers

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items.

arXiv AI
6d ago

Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More

The paper investigates whether the reasoning process in large language models mitigates or exacerbates bias. Using a within-model ablation on three high-stakes datasets (Adult, COMPAS, Credit) across three 32‑B models, the authors find that reasoning resolves some counterfactual fairness flips but creates roughly five times as many new flips at high confidence. They introduce two dynamic tools—Counterfactual Depth Probability Gap and Bias Transition Matrix—to trace how bias propagates and amplifies during reasoning depth and to explain the asymmetric dual effect.

By Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, Nitesh Chawla
arXiv AI
Jul 21

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.

By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin
arXiv Computation and Language
Sep 18

Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

The study investigates how limited reading capacity and claim wording influence consensus outcomes in language‑model networks. By modeling message capacity as the number of messages an agent reads, the authors show that when agents read fewer than about 6.4 messages on average, a wrong consensus becomes unreachable. However, the wording of a claim—its inherent threshold—can override this effect, leading to incorrect consensus even when most agents start correct.

By Makoto Fukushima
Hugging Face Trending Papers
Jul 20

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model.

Hugging Face Trending Papers
Aug 19

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.