arXiv Machine Learning

Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

arXiv:2607. 26845v1 Announce Type: new Abstract: Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions.

Hugging Face Trending Papers
Jul 29

Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty.

arXiv AI
Sep 21

How do LLMs Compute Verbal Confidence

arXiv:2603.17839v4 Announce Type: replace-cross Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...

By Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Veli\v{c}kovi\'c
arXiv Computation and Language
4d ago

Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice

The paper investigates whether input ablations on predictive models can reliably reveal the importance of information for explaining human sequential choice behavior. Using two synthetic bandit tasks with known generating policies, the authors compare GRUs, Transformers, a fine‑tuned LLaMA, and cognitive models under varied reward contributions. They find that while neural models can predict choices well, their responses to ablations often diverge from the true generating process, indicating that predictive accuracy alone does not guarantee faithful model ablations.

By Hanbo Xie
arXiv AI
Aug 28

Assessing mentalization in humans and large language models

The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.

By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
arXiv AI
Jul 7

ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observability

arXiv:2607. 02686v1 Announce Type: new Abstract: Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance from small language models (SLMs) that carry broad reasoning priors.

By Juarez Monteiro, Nathan Gavenski, Guilherme Lima, Francisco Galuppo, Odinaldo Rodrigues, Adriano Veloso