arXiv Machine Learning

Can Revealed Preferences Clarify LLM Alignment and Steering?

The paper proposes an empirical pipeline to estimate the preferences that a large language model (LLM) implicitly optimizes by combining the model’s probability distribution over unknowns with its chosen action, and fitting a discrete choice model to recover the underlying cost function. This revealed-preference framework enables rigorous assessment of whether LLMs act consistently toward a goal, can articulate objectives that align with their decision policy, and can be steered by prompting to follow a user-specified cost function. Experiments across four medical diagnosis domains and various frontier and open-source models show that while many LLMs exhibit internal coherence, they still struggle to accurately report or adopt preferences when guided by users.

arXiv AI
Sep 10

When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

The paper introduces a decision‑theoretic framework that elicits both probability judgments and decisions from large language models (LLMs) to test whether their reported beliefs are consistent with their actions. It shows that this framework yields empirically testable conditions without assuming a specific utility function. In clinical diagnosis simulations, the authors find that while LLMs’ reported beliefs are not perfect reflections of the information in their decisions, the discrepancies are small for the strongest models.

By Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder
arXiv AI
Jul 28

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

arXiv:2508. 08992v4 Announce Type: replace Abstract: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.

By Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Peixuan Han, Weiqi Wang, Yangqiu Song
arXiv Computation and Language
Aug 27

Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making

The study presents a benchmark of 208 rare‑disease clinical vignettes to evaluate how large language models (LLMs) handle ethically charged decision‑making. Across 11 state‑of‑the‑art LLMs, the models consistently favored justice—specifically equal resource allocation—over other bioethical principles such as beneficence, non‑maleficence, and autonomy. The authors also found that the framing of authority (committee vs. clinician vs. patient) influences which ethical principle the models prioritize, suggesting that institutional pressures may shape LLM decision support in rare‑disease care.

By Minda Zhao, Xu Han, Rishabh Goel, Maya Dagan, Noa Dagan, Adithya Madduri, Payal Chandak, Shilpa Nadimpalli Kobren, Isaac S. Kohane
arXiv Computer Vision
Aug 24

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v2 Announce Type: replace Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing mo...

By Uma Ranjan, Kunal Tilaganji, Aditya Koul, Anurag Mahipal, Dashpreet Singh, Hriday Rana, Manan Jain, Sidharth Gupta, Ajo Babu George, Vineeth Balasubramanian, Nagarajan Natarajan, Amit Sharma
arXiv Machine Learning
Aug 27

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

The paper introduces P4-DT, a personalized patient preference predictor that uses dilemma training to elicit context‑dependent decision reasoning. In a study of 12 patient‑surrogate pairs, P4‑DT achieved 81.7% accuracy in predicting patient treatment choices, outperforming unassisted surrogates (55.0%) and surrogates aided by a simpler P4 model (61.7%). The authors show that incorporating contextual scenarios and open‑ended text into prompts improves accuracy by 15 percentage points over static value ratings.

By Natasha Ureyang, Sebastian Porsdam Mann, Yuxin Liu, Zuriel Hassirim, Melanie Almonte, Wenhao Chen, Joyce Ng, Thant Nay Lin, Aung Thiha, Gerald CH Koh, Brian David Earp, Pin Sym Foong
arXiv AI
Jun 10

Superficial Beliefs in LLM Decision-Making

arXiv:2606. 11016v1 Announce Type: new Abstract: We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure.

By Gabriel Freedman, Francesca Toni