arXiv AI

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

arXiv:2606. 26502v1 Announce Type: new Abstract: Large reasoning models (LRMs) take longer on harder problems, just as humans do.

arXiv AI
6d ago

Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More

The paper investigates whether the reasoning process in large language models mitigates or exacerbates bias. Using a within-model ablation on three high-stakes datasets (Adult, COMPAS, Credit) across three 32‑B models, the authors find that reasoning resolves some counterfactual fairness flips but creates roughly five times as many new flips at high confidence. They introduce two dynamic tools—Counterfactual Depth Probability Gap and Bias Transition Matrix—to trace how bias propagates and amplifies during reasoning depth and to explain the asymmetric dual effect.

By Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, Nitesh Chawla
arXiv AI
Jul 21

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.

By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin
arXiv AI
Sep 15

From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

The paper compares human group discussions with large language model (LLM) deliberation traces on various reasoning tasks, finding that both humans and LLMs exhibit an assembly bonus asymmetry where discussion benefits the average member more than the best initial member. While LLM groups mirror some outcome-level patterns of human deliberation, they differ in process-level behaviors: they tend to follow majorities, surface less unique information, and converge earlier. Interventions inspired by human group‑decision research yield modest outcome improvements but do not eliminate coordination bottlenecks.

By Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch
arXiv Computation and Language
Sep 18

Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

The study investigates how limited reading capacity and claim wording influence consensus outcomes in language‑model networks. By modeling message capacity as the number of messages an agent reads, the authors show that when agents read fewer than about 6.4 messages on average, a wrong consensus becomes unreachable. However, the wording of a claim—its inherent threshold—can override this effect, leading to incorrect consensus even when most agents start correct.

By Makoto Fukushima
arXiv Computation and Language
Aug 24

Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed

The paper demonstrates that a prompt’s influence is not inherent to the prompt itself but depends on the model, as prompts optimized for one model degrade on another and rankings shift under neutral reformatting. By examining a task‑free structural readout—specifically the fixed‑point behavior of a short‑window argmax map—the authors show that nine tokens of conditioning can move the fixed‑point fraction across most of its range, altering structural classes and model rankings, while instruction tuning has no effect. Attempts to explain this phenomenon through prefix length, content type, bidirectionality, or attention‑sink dominance all fail, indicating that the prompt‑model pair is the fundamental unit of explanation. whyItMatters":"The study reveals that prompt effectiveness is model‑specific and that simple structural readouts can capture this interaction, challenging assumptions about prompt generality and guiding future prompt‑engineering efforts."

By Nicol\'as Vera Z\'u\~niga
Hugging Face Trending Papers
Aug 19

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.