arXiv AI By Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, Nitesh Chawla

Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More

Read the original on arXiv AI →

The paper investigates whether the reasoning process in large language models mitigates or exacerbates bias. Using a within-model ablation on three high-stakes datasets (Adult, COMPAS, Credit) across three 32‑B models, the authors find that reasoning resolves some counterfactual fairness flips but creates roughly five times as many new flips at high confidence. They introduce two dynamic tools—Counterfactual Depth Probability Gap and Bias Transition Matrix—to trace how bias propagates and amplifies during reasoning depth and to explain the asymmetric dual effect.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 21

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.

By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin