arXiv Computation and Language

Two Regimes of Chain-of-Thought Unfaithfulness: Metric-Based Detection Fails Where Models Are Wrong

arXiv Machine Learning
Jun 16

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

arXiv:2606. 15127v1 Announce Type: new Abstract: Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students intermediate steps, decision-support systems may require human oversight, and audit workflows may inspect traces for misleading or biased input.

By Xian Sun, Wei Gao, Yingshuo Wang, Lingdong Kong, Yanhang Li, Zhichao Fan, Zexin Zhuang, Wenlong Dong, Zhiyuan Zheng, Hrishikesh Paranjape, Abhishek Mandal, Johnny R. Zhang
arXiv AI
1d ago

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

The study evaluates whether the chain-of-thought (CoT) rationales produced by medical language models truly influence their answers. Using a 30‑operator perturbation audit that modifies both the question and the CoT (e.g., severity reversal, negation flip, demographic swap, evidence ablation), the authors found that 72.9% of edits did not change the model’s answer—a high Chain‑Decoupling Rate (CDR). Across 14 large language models and four medical QA benchmarks, the CoT text had little impact on accuracy, and removing CoT prompting did not reduce performance. "whyItMatters":"The findings suggest that current medical CoT outputs may be more documentation than genuine reasoning, highlighting the need for better faithfulness checks in clinical AI systems."

By Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long