Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More
Read the original on arXiv AI →The paper investigates whether the reasoning process in large language models mitigates or exacerbates bias. Using a within-model ablation on three high-stakes datasets (Adult, COMPAS, Credit) across three 32‑B models, the authors find that reasoning resolves some counterfactual fairness flips but creates roughly five times as many new flips at high confidence. They introduce two dynamic tools—Counterfactual Depth Probability Gap and Bias Transition Matrix—to trace how bias propagates and amplifies during reasoning depth and to explain the asymmetric dual effect.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.