arXiv AI By Sadia Kamal, Arefa Patwary, Anthony Marchiafava, Atriya Sen, Sagnik Ray Choudhury

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

Read the original on arXiv AI →

arXiv:2607. 05554v1 Announce Type: cross Abstract: Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Beliefs and Behavior in Language Models

arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...

By Alex Smolin, Bryan Wilder
arXiv AI
Jul 31

Ask don't tell: Reducing sycophancy in large language models

arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.

By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
arXiv Machine Learning
Sep 3

Prompting the Unknown: Understanding Response Uncertainty in Large Language Models

The paper introduces a prompt-response concept model that links the amount of task-relevant information in a prompt to the uncertainty of responses generated by large language models (LLMs). It identifies four sources of response uncertainty—prompt underspecification, model quality, task variability, and semantic redundancy—and demonstrates that uncertainty decreases as prompt informativeness or model quality increases, analogous to epistemic uncertainty in probabilistic models. Experiments on real-world datasets confirm the theoretical predictions and validate the model.

By Ze Yu Zhang, Arun Verma, Finale Doshi-Velez, Bryan Kian Hsiang Low