arXiv Machine Learning

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

The paper introduces P4-DT, a personalized patient preference predictor that uses dilemma training to elicit context‑dependent decision reasoning. In a study of 12 patient‑surrogate pairs, P4‑DT achieved 81.7% accuracy in predicting patient treatment choices, outperforming unassisted surrogates (55.0%) and surrogates aided by a simpler P4 model (61.7%). The authors show that incorporating contextual scenarios and open‑ended text into prompts improves accuracy by 15 percentage points over static value ratings.

arXiv AI
Jul 9

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

arXiv:2601. 16529v4 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines.

By Dongshen Peng, Yi Wang, Austin Schoeffler, Sun-ha Hong, Brian Suffoletto, David Kim, Carl Preiksaitis, Christian Rose
arXiv Computation and Language
Aug 28

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models proposes a new framework to improve medical reasoning in LLMs. It introduces two key conditions—Causal Sufficiency and Proximal Learnability—to curate high-quality training trajectories, using agreement-based self-verification and dynamic entropy bounds. Experiments on medical multimodal and text-only benchmarks show that CARE outperforms competitors, reducing incorrect reasoning and enhancing training stability.

By Yucheng Zhou, Peng Luo, Qianning Wang, Chengzhong Xu, Jianbing Shen
arXiv AI
Jul 15

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

arXiv:2512. 01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.

By David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
arXiv Computation and Language
Sep 7

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

MedProb is a lightweight probing framework that predicts multiple-choice medical visual question answering (Med‑VQA) answers directly from frozen vision‑language model (VLM) representations, avoiding free‑text generation. On datasets such as PATH‑VQA, SLAKE, and VQA‑RAD, MedProb extracts more answer‑relevant signal than prompting and outperforms both medical VLMs and agentic systems. The approach also narrows the performance gap between small and large models, shows that medical adaptation does not consistently improve linear decodability, and reveals positional biases in both prompting and generation.

By Erfan Nourbakhsh, Ke Yang, Anthony Rios
arXiv AI
Jun 30

An AI agent for treatment reasoning over a biomedical tool universe

arXiv:2606. 28692v1 Announce Type: new Abstract: Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy.

By Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik
arXiv AI
Aug 10

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

arXiv:2608. 07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy.

By Valentin Li\'{e}vin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang
arXiv AI
Jul 7

Interactive Learning for LLM Reasoning

arXiv:2509. 26306v5 Announce Type: replace Abstract: Existing multi-agent learning approaches have developed interactive training environments to explicitly promote collaboration among multiple Large Language Models (LLMs), thereby constructing stronger multi-agent systems (MAS).

By Hehai Lin, Shilei Cao, Sudong Wang, Haotian Wu, Minzhi Li, Linyi Yang, Juepeng Zheng, Chengwei Qin