Hugging Face Trending Papers

Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs

Large language models (LLMs) frequently encounter inputs that disagree with their prior outputs, through user pushback, retrieved documents, or web search results. While the way they resolve such conflicts -- a process we frame as cognitive dissonance resolution -- has been characterized behaviorally, its connection to internal model uncertainty is not well understood.

arXiv AI
Sep 17

From 'May' to 'Is': Certainty Distortion in Language Model Rewriting

The study examines how language models (LMs) alter the expressed certainty of statements when rewriting text, a process termed certainty distortion. Using an LM‑based metric aligned with human judgments, the authors find that up to 75% of LM outputs exhibit such distortion, with most models more likely to inflate certainty than reduce it. Repeated paraphrasing can amplify this effect, especially in medical contexts, and while prompt interventions help, they do not fully eliminate the bias.

By Catarina G Belem, Shang Wu, Hongyu Yao, Mark Steyvers, Sameer Singh, Padhraic Smyth
arXiv AI
Jul 3

Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

arXiv:2607. 01951v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat from established consensus when a user signals doubt -- drifting toward a false balance that treats settled science as one view among several.

By Minjong Cheon
arXiv Machine Learning
Sep 3

Prompting the Unknown: Understanding Response Uncertainty in Large Language Models

The paper introduces a prompt-response concept model that links the amount of task-relevant information in a prompt to the uncertainty of responses generated by large language models (LLMs). It identifies four sources of response uncertainty—prompt underspecification, model quality, task variability, and semantic redundancy—and demonstrates that uncertainty decreases as prompt informativeness or model quality increases, analogous to epistemic uncertainty in probabilistic models. Experiments on real-world datasets confirm the theoretical predictions and validate the model.

By Ze Yu Zhang, Arun Verma, Finale Doshi-Velez, Bryan Kian Hsiang Low
arXiv AI
Jul 23

Rethinking Uncertainty Evaluation in Large Language Models

arXiv:2607. 19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function.

By Krish Matta, Atharv Naphade, Andy Zou
arXiv AI
Aug 28

Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study

The study compares communication styles of large language models (LLMs) and humans in explaining COVID‑19 misinformation, using a dataset of 1,498 fact‑checking claims and 99 blinded reader evaluations. LLM‑generated explanations scored lower on persuasive strategies, certainty, and alignment with social values, yet over 60% of participants preferred LLM content for clarity, completeness, and persuasiveness. The findings suggest that reader preference may not align with traditional measures of communication quality, highlighting both the promise and limits of LLMs in health communication.

By Jiawei Zhou, Kritika Venkatachalam, Minje Choi, Koustuv Saha, Munmun De Choudhury