arXiv AI

Referential Uncertainty in Human--AI Collaboration

The study investigates how humans and AI collaborate on a puzzle task, focusing on referential uncertainty—when a description could refer to multiple objects. It finds that eliciting a belief distribution over candidate pieces yields better calibration and discrimination than raw action probabilities, and that precise descriptions or well‑targeted hedges significantly reduce the acceptance of wrong placements. However, the AI rarely externalizes uncertainty, and poorly targeted hedges can be counterproductive.

Hugging Face Trending Papers
6d ago

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.

arXiv AI
Aug 28

How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

The study evaluates how large language models (LLMs) interpret verbal probability expressions by mapping words to numbers and testing consistency across 19 models. Results show that LLMs largely mirror human benchmarks—preserving word order, recovering key anchor points, and reflecting the high variance of the term "possible"—but they exhibit a systematic upward bias for negative expressions like "unlikely" and "improbable." Explanation elicitation reduces within‑model variance but increases divergence between models, while a bidirectional roundtrip test reveals that leading models maintain coherent internal representations.

By Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic
arXiv AI
Jul 7

ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observability

arXiv:2607. 02686v1 Announce Type: new Abstract: Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance from small language models (SLMs) that carry broad reasoning priors.

By Juarez Monteiro, Nathan Gavenski, Guilherme Lima, Francisco Galuppo, Odinaldo Rodrigues, Adriano Veloso