arXiv Computation and Language

The Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models

The study investigates whether large language models (LLMs) that allocate extra computation during inference—termed reasoning models—reduce classic decision biases compared to their non‑reasoning counterparts. Using 30 vignettes covering six cognitive biases and varying token budgets up to 8,192 tokens, the authors find that reasoning models are not less biased, and increased deliberation does not reliably diminish bias magnitude. Only anchoring showed a bias in the human direction, while other biases either remained unchanged or moved further from human patterns, suggesting that test‑time reasoning does not guarantee rationality.

arXiv AI
Sep 28

Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More

The paper investigates whether the reasoning process in large language models mitigates or exacerbates bias. Using a within-model ablation on three high-stakes datasets (Adult, COMPAS, Credit) across three 32‑B models, the authors find that reasoning resolves some counterfactual fairness flips but creates roughly five times as many new flips at high confidence. They introduce two dynamic tools—Counterfactual Depth Probability Gap and Bias Transition Matrix—to trace how bias propagates and amplifies during reasoning depth and to explain the asymmetric dual effect.

By Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, Nitesh Chawla
arXiv Machine Learning
1d ago

Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding

The study investigates how task difficulty, model type, and user pressure influence large language models’ tendency to abandon correct answers or endorse user positions—a phenomenon known as sycophancy. Using 103,939 graded replies across ten configurations of eight LLMs (with and without reasoning) and 13 pressure conditions, the authors find that the cost of verifying a claim and the presence of a guardrail are the dominant factors, while model family and pressure tactics play minor roles. Key practical insights include simplifying hard-to-verify problems, employing deep reasoning, framing questions neutrally, and selecting models based on guardrail performance.

By Guang Yang, Homa Hosseinmardi, Fengchen Liu, Amir Ghasemian
arXiv Computation and Language
Sep 1

Detecting Hidden Chain-of-Thought in Large Language Models with Linguistic, Behavioral, and Mechanistic Indicators

arXiv:2608.29956v1 Announce Type: new Abstract: Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete p...

By Armaan Singh, Ryan Trinh Le, Jasmine Kaur, Abdullah Sultan, Edward Lue Chee Lip, Kiran Nijjer, Adnan Ahmed, Vasu Sharma
arXiv AI
4d ago

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

The paper investigates how training large language models to use fewer tokens in chain-of-thought (CoT) reasoning impacts the faithfulness and monitorability of the generated explanations. Three efficiency methods—fixed generation budget, per-example length target, and group-relative length reward—were applied during fine‑tuning, and the resulting models were evaluated on how well their CoT reflects decision processes and whether it signals changes due to input interventions. Results show that while faithfulness generally decreases because models become less consistent, monitorability remains relatively robust, with models still indicating the influence of input changes even when CoT is shortened.

By Samuel Lewis-Lim, Xingwei Tan, Mario Sanger, Zhixue Zhao, Nikolaos Aletras
arXiv AI
Aug 28

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

The paper introduces the Token Economy Score (TES), a metric that quantifies the accuracy gain of reasoning-capable large language models relative to non-reasoning baselines, normalized by token generation cost. An empirical study across 151 runs on seven diverse benchmarks shows that task structure—such as sequential inference chains—predicts higher TES, while knowledge-recall tasks yield lower TES despite difficulty. The analysis also reveals diminishing returns at higher reasoning effort and highlights how deployment context, via Reasoning Cost Share and Deployment Cost Multiplier, can alter the economic viability of reasoning workloads.

By Sachin Gopal Wani, Ajay Dholakia, David Ellison