The study examines how language models (LMs) alter the expressed certainty of statements when rewriting text, a process termed certainty distortion. Using an LM‑based metric aligned with human judgments, the authors find that up to 75% of LM outputs exhibit such distortion, with most models more likely to inflate certainty than reduce it. Repeated paraphrasing can amplify this effect, especially in medical contexts, and while prompt interventions help, they do not fully eliminate the bias.
By Catarina G Belem, Shang Wu, Hongyu Yao, Mark Steyvers, Sameer Singh, Padhraic Smyth
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
By Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
arXiv:2606. 07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization.
By Mahdi Alkaeed
arXiv:2509. 08604v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pre-training or fine-tuning on medical data to enhance domain-specific accuracy and safety.
By Anran Li, Lingfei Qian, Mengmeng Du, Yu Yin, Yan Hu, Zihao Sun, Yihang Fu, Hyunjae Kim, Erica Stutz, Xuguang Ai, Qianqian Xie, Rui Zhu, Jimin Huang, Yifan Yang, Siru Liu, Yih-Chung Tham, Lucila Ohno-Machado, Hyunghoon Cho, Zhiyong Lu, Hua Xu, Qingyu Chen
The paper introduces a prompt-response concept model that links the amount of task-relevant information in a prompt to the uncertainty of responses generated by large language models (LLMs). It identifies four sources of response uncertainty—prompt underspecification, model quality, task variability, and semantic redundancy—and demonstrates that uncertainty decreases as prompt informativeness or model quality increases, analogous to epistemic uncertainty in probabilistic models. Experiments on real-world datasets confirm the theoretical predictions and validate the model.
By Ze Yu Zhang, Arun Verma, Finale Doshi-Velez, Bryan Kian Hsiang Low
arXiv:2608. 14630v1 Announce Type: cross Abstract: Human decision-making is often shaped by a range of well-documented cognitive biases.
By Zirui Cheng, Joey Chan, Simo Du, Chenhao Tan, Yue Guo, Hao Peng
arXiv:2606. 00467v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions.
By Etienne Casanova, Rafal Kocielnik, R. Michael Alvarez
arXiv:2608.17809v2 Announce Type: replace
Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevita...
By Quang Minh Nguyen, Luis Frentzen Salim
arXiv:2607. 20462v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking.
By Melanie Rieff, Robin Staab, Thibaud Gloaguen, Stefan Hegselmann, Martin Vechev
arXiv:2603. 24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment.
By Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy
MIRA is a bilingual benchmark that evaluates whether large language models (LLMs) provide consistent medical information across different user phrasings, languages, and health literacy levels. It contains 4,320 prompts derived from 60 medically reviewed low‑risk health questions and reveals that models tend to omit key information and offer fewer concrete next steps when responding to low health‑literacy signals, a phenomenon termed Differential Information Dilution (DID). A knowledge‑guided mitigation prompt can reduce this dilution for most models, notably improving Claude and Qwen.
By Mengyu Xu, Qiaoxin Yang, Qianqian Wang, Xiwei Dai, Weiyi Wu, Chongyang Gao
The paper investigates whether language models’ self-reported confidence is meaningful without additional training. By evaluating three training‑free signals—direct verbalization, post‑hoc probability estimates, and agreement across multiple generations—on 100 TriviaQA questions, the authors find that direct verbalization alone achieves high AUROC scores (0.956 and 0.937) for correctness prediction, while agreement-based methods perform noticeably worse. Re‑eliciting confidence for the same answers shows modest score shifts and occasional decision flips, and an audit of biography claims reveals only a small confidence gap between supported and contradicted statements.
By Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li