The paper introduces a decision‑theoretic framework that elicits both probability judgments and decisions from large language models (LLMs) to test whether their reported beliefs are consistent with their actions. It shows that this framework yields empirically testable conditions without assuming a specific utility function. In clinical diagnosis simulations, the authors find that while LLMs’ reported beliefs are not perfect reflections of the information in their decisions, the discrepancies are small for the strongest models.
By Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder
arXiv:2606. 19353v1 Announce Type: cross Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations.
By Jinseok Chung, Minkyoung Song, Hyunji Jung, Namhoon Lee
arXiv:2606. 03245v1 Announce Type: cross Abstract: Concepts of calibration formalize the compatibility between probabilistic predictions and the respective outcomes.
By Johannes Resin, Lu Yang, Tilmann Gneiting
The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.
By Erin George, Deanna Needell, Berk Ustun
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
By Osvaldo Simeone
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
By Arthur Hoarau
arXiv:2602. 21889v2 Announce Type: replace-cross Abstract: Predictions from ML models support human decision making in several fields, including high-stakes ones such as healthcare and the judiciary.
By Otto Nyberg, Fausto Carcassi, Davide Tugnoli, Giovanni Cin\`a
arXiv:2505. 19033v2 Announce Type: replace-cross Abstract: Conformal prediction (CP) is a widely used frequentist framework to quantify uncertainty by constructing prediction sets with user-specified marginal coverage guarantees.
By Alireza Javanmardi, Soroush H. Zargarbashi, Santo M. A. R. Thies, Willem Waegeman, Aleksandar Bojchevski, Eyke H\"ullermeier
arXiv:2606. 00102v1 Announce Type: new Abstract: Over the centuries, probability theory has grown from the calculus of games of chance into a central framework for reasoning under uncertainty.
By Jean-Louis Le Mou\"el, Vincent Courtillot, Dominique Gibert, Vladimir Kossobokov, Jean-Baptiste Boul\'e, Pierpaolo Zuddas, Fernando Lopes, Pa\"ikan Marccagi, Alexis Maineult
arXiv:2607. 15196v1 Announce Type: cross Abstract: We present a novel viewpoint for uncertainty quantification.
By Raghad Alamri, Michele Caprio, Gavin Brown
arXiv:2507. 06722v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) internally represent and process their predictions is central to detecting uncertainty and preventing hallucinations.
By Sunwoo Kim, Haneul Yoo, Alice Oh
The paper investigates how Large Language Models can be used to approximate domain expert priors for Bayesian Networks by extracting probabilistic knowledge about real‑world events. Experiments on eighty publicly available networks across domains such as healthcare and finance show that LLM‑derived conditional probabilities outperform random, uniform, and next‑token baselines. The authors also demonstrate that these LLM‑generated priors can refine data‑driven distributions, especially when data is scarce, and provide the first comprehensive baseline for evaluating LLM performance in probabilistic knowledge extraction.
By Aliakbar Nafar, Kristen Brent Venable, Zijun Cui, Parisa Kordjamshidi