arXiv AI

MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

arXiv:2601. 09853v3 Announce Type: replace-cross Abstract: Real-world health questions from patients often unintentionally embed false assumptions or premises.

arXiv Machine Learning
Jul 17

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

arXiv:2508. 00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against.

By Jiazhen Pan (Cherise), Bailiang Jian (Cherise), Paul Hager (Cherise), Yundi Zhang (Cherise), Che Liu (Cherise), Friederike Jungmann (Cherise), Hongwei Bran Li (Cherise), Julian Canisius (Cherise), Chenyu You (Cherise), Junde Wu (Cherise), Jiayuan Zhu (Cherise), Fenglin Liu (Cherise), Yuyuan Liu (Cherise), Niklas Bubeck (Cherise), Moritz Knolle (Cherise), Chen (Cherise), Chen (Cherise), Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert
arXiv Computation and Language
Sep 1

CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance

The paper introduces MIMIC-DOS, a dataset derived from MIMIC-IV that focuses on ICU cases where patient symptoms and medical signs are discordant. It presents CARE, a privacy‑compliant multi‑stage agentic reasoning framework that uses a proprietary LLM to generate structured categories and transitions, while a local LLM performs evidence acquisition and decision‑making. In retrospective evaluations on MIMIC‑DOS, CARE outperforms other LLMs and agentic workflows, demonstrating stronger handling of conflicting clinical evidence while preserving patient privacy.

By Haochen Liu, Weien Li, Rui Song, Zeyu Li, Chun Jason Xue, Xiao-Yang Liu, Sam Nallaperuma-Herzberg, Xue Liu, Ye Yuan
arXiv AI
Sep 3

Untangling the Mechanisms of Misleading Context in Medical Question Answering

The paper investigates how misleading context—specifically fabricated evidence and bare assertions—affects large language models’ medical question‑answering performance. Experiments on MedMisBench show that models are more prone to adopt answers based on assertions than fabricated evidence, and that these misleading cues are often disclosed in reasoning traces but rarely in final responses. A monitor that reads open reasoning traces can detect most corrupted decisions, whereas monitoring only responses is less effective.

By Robin Linzmayer, No\'emie Elhadad
arXiv Computation and Language
Sep 3

HarmReduction: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

The paper introduces HarmReduction, a benchmark for evaluating large language models (LLMs) on their ability to provide accurate and safe harm reduction information to people who use drugs (PWUD). The benchmark, HR-Basic, contains 2,160 question‑answer‑evidence pairs covering safety boundary checks, quantitative value provision, and polysubstance risk inference. Experiments show that even state‑of‑the‑art LLMs struggle with accuracy and can pose severe safety risks, underscoring the need for a dedicated evaluation framework.

By Kaixuan Wang, Chenxin Diao, Jason T. Jacques, Zhongliang Guo, Shuai Zhao
arXiv AI
Jul 21

Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry

arXiv:2505. 02722v2 Announce Type: replace Abstract: Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited.

By Junu Kim, Chaeeun Shim, Sungjin Park, Su Yeon Lee, Gee Young Suh, Chae-Man Lim, Seong Jin Choi, Song Mi Moon, Kyoung-Ho Song, Eu Suk Kim, Hong Bin Kim, Sejoong Kim, Chami Im, Dong-Wan Kang, Yong Soo Kim, Hee-Joon Bae, Sung Yoon Lim, Han-Gil Jeong, Edward Choi
arXiv Machine Learning
Sep 10

HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care

arXiv:2609.06976v1 Announce Type: new Abstract: As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients...

By Yuchen Niu, Yanan Ma, Srinivasan Nandakumar, Maolin Chen, Viktor Schlegel, Kexin Wei, Ling Cheng, Anna Bird, Anil Anthony Bharath, Siew-Kei Lam
arXiv Computation and Language
Sep 1

MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability

MedConceal is a new benchmark for evaluating medical dialogue systems on hidden‑concern reasoning under partial observability. It features 300 curated cases and 600 clinician‑LLM interactions, using an interactive patient simulator that hides latent concerns and tracks their revelation and resolution through theory‑grounded communication signals. The benchmark assesses both confirmation (surfacing hidden concerns) and intervention (addressing the primary concern), revealing that current models excel on different metrics while human clinicians still outperform them on intervention success.

By Yikun Han, Joey Chan, Jingyuan Chen, Mengting Ai, Simo Du, Yue Guo