arXiv AI

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

arXiv:2508. 08992v4 Announce Type: replace Abstract: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.

arXiv Machine Learning
Sep 3

Prompting the Unknown: Understanding Response Uncertainty in Large Language Models

The paper introduces a prompt-response concept model that links the amount of task-relevant information in a prompt to the uncertainty of responses generated by large language models (LLMs). It identifies four sources of response uncertainty—prompt underspecification, model quality, task variability, and semantic redundancy—and demonstrates that uncertainty decreases as prompt informativeness or model quality increases, analogous to epistemic uncertainty in probabilistic models. Experiments on real-world datasets confirm the theoretical predictions and validate the model.

By Ze Yu Zhang, Arun Verma, Finale Doshi-Velez, Bryan Kian Hsiang Low
arXiv AI
Sep 10

Beliefs and Behavior in Language Models

arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...

By Alex Smolin, Bryan Wilder
arXiv Machine Learning
1d ago

"very likely" Means "uncertain"? How LLMs Diverge from Humans in Linguistic Uncertainty Quantification

The paper examines how large language models (LLMs) express uncertainty compared to humans, noting that humans use verbal markers like "possible" or "likely" to convey metacognitive awareness. By curating a corpus of human uncertainty markers and benchmarking LLMs against it, the authors find that LLMs encode these markers with numerical levels that differ substantially from human usage. They introduce METHODNAME, an optimization-based algorithm that learns an optimal uncertainty profile over verbal markers directly from LLM outputs, enabling a direct comparison of confidence semantics and revealing systematic disparities in verbal expressions.

By Jinhao Duan, Zicheng Liu, Zijie Liu, Kaidi Xu, Tianlong Chen
arXiv AI
Sep 10

When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

The paper introduces a decision‑theoretic framework that elicits both probability judgments and decisions from large language models (LLMs) to test whether their reported beliefs are consistent with their actions. It shows that this framework yields empirically testable conditions without assuming a specific utility function. In clinical diagnosis simulations, the authors find that while LLMs’ reported beliefs are not perfect reflections of the information in their decisions, the discrepancies are small for the strongest models.

By Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder
Hugging Face Trending Papers
Jun 22

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.

arXiv AI
Jun 17

Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.

By Manon Reusens, Sofie Goethals, Toon Calders, David Martens
arXiv AI
Jul 7

The Anatomy of Uncertainty in LLMs

arXiv:2603. 24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment.

By Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy
arXiv AI
Jul 23

Rethinking Uncertainty Evaluation in Large Language Models

arXiv:2607. 19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function.

By Krish Matta, Atharv Naphade, Andy Zou