arXiv AI By Chang Hong, Minghao Wu, Qingying Xiao, Yuchi Wang, Xiang Wan, Guangjun Yu, Benyou Wang, Yan Hu

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

Read the original on arXiv AI →

arXiv:2508. 05132v3 Announce Type: replace-cross Abstract: As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 18

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

arXiv:2608. 15424v1 Announce Type: cross Abstract: The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making.

By Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey
arXiv AI
Sep 4

Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications

The paper reviews how Large Language Models (LLMs) are being adapted for medical reasoning, moving beyond single-step answers to systems that can systematically, transparently, and verifiably reason. It introduces a taxonomy of enhancement techniques, split into training-time methods such as supervised fine‑tuning and reinforcement learning, and test-time methods like prompt engineering and multi‑agent systems. The review examines their application across text, image, and code modalities in key clinical areas—diagnosis, education, and treatment planning—and tracks the shift in evaluation benchmarks from simple accuracy to more nuanced assessments of reasoning quality and visual interpretability.

By Zizhan Ma, Wenxuan Wang, Meidan Ding, Shiyi Zheng, Shengyuan Liu, Jie Liu, Jiaming Ji, Linlin Shen, Yixuan Yuan, Wenting Chen
arXiv Computation and Language
Sep 1

CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance

The paper introduces MIMIC-DOS, a dataset derived from MIMIC-IV that focuses on ICU cases where patient symptoms and medical signs are discordant. It presents CARE, a privacy‑compliant multi‑stage agentic reasoning framework that uses a proprietary LLM to generate structured categories and transitions, while a local LLM performs evidence acquisition and decision‑making. In retrospective evaluations on MIMIC‑DOS, CARE outperforms other LLMs and agentic workflows, demonstrating stronger handling of conflicting clinical evidence while preserving patient privacy.

By Haochen Liu, Weien Li, Rui Song, Zeyu Li, Chun Jason Xue, Xiao-Yang Liu, Sam Nallaperuma-Herzberg, Xue Liu, Ye Yuan