arXiv Computation and Language

Incidental information contaminates patient notes and disrupts clinical reasoning in large language models

arXiv Computation and Language
Sep 22

Knowledge Graph-Augmented Ambient AI for Clinical Note Generation

arXiv:2609.22239v1 Announce Type: new Abstract: Ambient AI is increasingly adopted in healthcare to automatically generate clinical notes from patient-clinician conversations, with the potential to s...

By Jakir Hossain, Yi-Fei Zhao, Hongjian Wang, Minmei Shih, Katie Leigh Mullen, Ahmad P. Tafti, Leming Zhou, Manoj Purohit, William Hogan, Jay Zeng, Elizabeth Skidmore, Yanshan Wang
arXiv AI
6d ago

Scaling Clinical Judgment to Evaluate Medical AI

The paper introduces PrecepTron, a 32‑billion‑parameter language model fine‑tuned with low‑rank adaptation to evaluate clinical reasoning in large language models (LLMs) at a physician level. It also releases GRAND‑ROUNDS, a benchmark of 9,217 scored responses from 160 clinicians across seven studies. Using PrecepTron, the authors replicate key findings from major medical AI studies and explore new questions about LLM diagnostic accuracy, demonstrating that fine‑tuned models can provide consistent, scalable physician‑level scoring.

By Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai
arXiv AI
Sep 3

Untangling the Mechanisms of Misleading Context in Medical Question Answering

The paper investigates how misleading context—specifically fabricated evidence and bare assertions—affects large language models’ medical question‑answering performance. Experiments on MedMisBench show that models are more prone to adopt answers based on assertions than fabricated evidence, and that these misleading cues are often disclosed in reasoning traces but rarely in final responses. A monitor that reads open reasoning traces can detect most corrupted decisions, whereas monitoring only responses is less effective.

By Robin Linzmayer, No\'emie Elhadad
arXiv AI
Aug 20

Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation

The paper investigates how the inference setup of large language models (LLMs) influences their behavior in a medical resource‑allocation scenario. By comparing paired‑context and independent‑inference experiments, the authors show that adding a single contrasting patient sentence can shift the model’s probability assignments in opposite directions across most tested models. Additional experiments varying scenario attributes further demonstrate that patient information can have context‑dependent effects on LLM outputs.

By Spencer Gibson, Tyler Crosse, Magnus Saebo, Achyutha Menon, Eyon Jang, Diogo Cruz
arXiv AI
Jul 16

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

arXiv:2607. 13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.

By Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler
arXiv AI
Jun 2

Understanding Stigmatizing Language in Clinical Documentation: A Paired Comparison of Ambient AI Drafts and Clinician Finalized Notes

arXiv:2606. 00019v1 Announce Type: cross Abstract: Ambient artificial intelligence (AI) documentation tools are increasingly deployed to reduce clinician documentation burden, but their implications for biased language in clinical notes remain unclear.

By Yiliang Zhou, Yawen Guo, Sairam Sutari, Jasmine Dhillon, Alexandra L. Beck, Emilie Chow, Steven Tam, Danielle Perret, Deepti Pandita, Gelareh Sadigh, Archana J. McEligot, Kai Zheng