arXiv:2608.29464v1 Announce Type: cross
Abstract: Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness...
By Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
The paper investigates whether the reasoning process in large language models mitigates or exacerbates bias. Using a within-model ablation on three high-stakes datasets (Adult, COMPAS, Credit) across three 32‑B models, the authors find that reasoning resolves some counterfactual fairness flips but creates roughly five times as many new flips at high confidence. They introduce two dynamic tools—Counterfactual Depth Probability Gap and Bias Transition Matrix—to trace how bias propagates and amplifies during reasoning depth and to explain the asymmetric dual effect.
By Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, Nitesh Chawla
The study investigates how task difficulty, model type, and user pressure influence large language models’ tendency to abandon correct answers or endorse user positions—a phenomenon known as sycophancy. Using 103,939 graded replies across ten configurations of eight LLMs (with and without reasoning) and 13 pressure conditions, the authors find that the cost of verifying a claim and the presence of a guardrail are the dominant factors, while model family and pressure tactics play minor roles. Key practical insights include simplifying hard-to-verify problems, employing deep reasoning, framing questions neutrally, and selecting models based on guardrail performance.
By Guang Yang, Homa Hosseinmardi, Fengchen Liu, Amir Ghasemian
arXiv:2606. 26502v1 Announce Type: new Abstract: Large reasoning models (LRMs) take longer on harder problems, just as humans do.
By Han-yu Wang
arXiv:2606. 11211v1 Announce Type: cross Abstract: The ability of large language models (LLMs) to express calibrated uncertainty is important for safe deployment.
By Prakul Sunil Hiremath, Harshit R. Hiremath
arXiv:2608.29956v1 Announce Type: new
Abstract: Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete p...
By Armaan Singh, Ryan Trinh Le, Jasmine Kaur, Abdullah Sultan, Edward Lue Chee Lip, Kiran Nijjer, Adnan Ahmed, Vasu Sharma