arXiv Machine Learning
Aug 27

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

The paper introduces P4-DT, a personalized patient preference predictor that uses dilemma training to elicit context‑dependent decision reasoning. In a study of 12 patient‑surrogate pairs, P4‑DT achieved 81.7% accuracy in predicting patient treatment choices, outperforming unassisted surrogates (55.0%) and surrogates aided by a simpler P4 model (61.7%). The authors show that incorporating contextual scenarios and open‑ended text into prompts improves accuracy by 15 percentage points over static value ratings.

By Natasha Ureyang, Sebastian Porsdam Mann, Yuxin Liu, Zuriel Hassirim, Melanie Almonte, Wenhao Chen, Joyce Ng, Thant Nay Lin, Aung Thiha, Gerald CH Koh, Brian David Earp, Pin Sym Foong
arXiv AI
Jul 9

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

arXiv:2601. 16529v4 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines.

By Dongshen Peng, Yi Wang, Austin Schoeffler, Sun-ha Hong, Brian Suffoletto, David Kim, Carl Preiksaitis, Christian Rose
arXiv Computation and Language
Aug 28

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models proposes a new framework to improve medical reasoning in LLMs. It introduces two key conditions—Causal Sufficiency and Proximal Learnability—to curate high-quality training trajectories, using agreement-based self-verification and dynamic entropy bounds. Experiments on medical multimodal and text-only benchmarks show that CARE outperforms competitors, reducing incorrect reasoning and enhancing training stability.

By Yucheng Zhou, Peng Luo, Qianning Wang, Chengzhong Xu, Jianbing Shen
arXiv AI
Jul 15

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

arXiv:2512. 01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.

By David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh