arXiv:2608.01017v2 Announce Type: replace-cross
Abstract: Large language models can answer a medical question correctly and still abandon that answer when a user pushes back. We study this failure as...
By Kaike Ping, Buse \c{C}ar{\i}k, Caleb Wohn, Xiaohan Ding, Tongshuai Wang, Eugenia Rho
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
By Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
arXiv:2606. 07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization.
By Mahdi Alkaeed
arXiv:2607. 27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear.
By Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta, Anik Pal Chowdhury
MedFabric is a new benchmark for detecting word‑level medical fabrications, comprising 646 fabricated statements each paired with a ground‑truth passage that shares the same LLM authorship and nearly identical wording. The study shows that current detectors perform poorly—expert clinicians achieve only 53.3% macro‑F1 and no detector family surpasses 60% without gold evidence—highlighting that detection hinges on evidence correctness rather than subtlety of fabrication. The authors demonstrate that a retrieval‑confidence gate can substantially improve performance, raising macro‑F1 from 61% to 74%.
By Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng
arXiv:2601. 09853v3 Announce Type: replace-cross Abstract: Real-world health questions from patients often unintentionally embed false assumptions or premises.
By Sraavya Sambara, Yuan Pu, Ayman Ali, Vishala Mishra, Lionel Wong, Monica Agrawal
arXiv:2603. 09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigate against.
By Brandon C. Colelough, Davis Bartels, Dina Demner-Fushman
arXiv:2606. 07951v1 Announce Type: cross Abstract: Humans increasingly turn to Language Models (LMs) in ways that shape beliefs and drive decisions, including discussing, rewriting, and summarizing information from scientific articles, news, and medical reports.
By Catarina G Belem, Shang Wu, Hongyu Yao, Mark Steyvers, Sameer Singh, Padhraic Smyth
arXiv:2606. 29034v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly summarize clinical evidence, where a claim's weight depends on how strongly it is supported.
By Soroosh Tayebi Arasteh
arXiv:2503. 10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.
By Krishna Subedi
arXiv:2608. 14630v1 Announce Type: cross Abstract: Human decision-making is often shaped by a range of well-documented cognitive biases.
By Zirui Cheng, Joey Chan, Simo Du, Chenhao Tan, Yue Guo, Hao Peng
The study evaluates whether the chain-of-thought (CoT) rationales produced by medical language models truly influence their answers. Using a 30‑operator perturbation audit that modifies both the question and the CoT (e.g., severity reversal, negation flip, demographic swap, evidence ablation), the authors found that 72.9% of edits did not change the model’s answer—a high Chain‑Decoupling Rate (CDR). Across 14 large language models and four medical QA benchmarks, the CoT text had little impact on accuracy, and removing CoT prompting did not reduce performance.
"whyItMatters":"The findings suggest that current medical CoT outputs may be more documentation than genuine reasoning, highlighting the need for better faithfulness checks in clinical AI systems."
By Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long