arXiv AI

Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction

The paper introduces CARing, a framework that improves next‑visit diagnosis prediction by representing diagnoses with compositional Semantic IDs (SIDs) and optimizing multi‑label coverage through reinforcement learning. CARing encodes ontology‑enriched disease semantics into compact SIDs, aligns them with natural language and EHR contexts, and employs a coverage reward to encourage diverse diagnosis predictions. On MIMIC‑III and MIMIC‑IV datasets, CARing outperforms all EHR‑trained baselines in weighted F1 and achieves the highest top‑k recall, including R@30 scores above 46% in reasoning mode.

arXiv AI
Aug 25

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

arXiv:2608.22622v1 Announce Type: cross Abstract: Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficul...

By Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun, Hruday Tej Akkaladevi, Peiyu Lu, Jordan Rosen, Sumit Kapoor, Sasank Desaraju, Grace R. Thompson, Jacob Purcell, Michael Petrauskis, Philip KW. Hong, Meghan Brennan, Sarah Chrabaszcz, Tierra Smith, Ronnie Ren, Michel S. Kabbash, Ceyhun Haziroglu, Rushi Patel, Gabriel Gomez, Charlotte Chaiklin, Randy Leung, Kenneth N. John, Whitman Wiggins, Philip Kayser, Vincent Bird, Maria Bruzzone, Tyler J. Loftus, Azra Bihorac, Parisa Rashidi
arXiv Machine Learning
Jun 2

OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction

arXiv:2510. 17532v2 Announce Type: replace-cross Abstract: Predicting cancer treatment outcomes requires models that are both accurate and interpretable, particularly in the presence of heterogeneous clinical data.

By Raghu Vamshi Hemadri, Geetha Krishna Guruju, Kristi Topollai, Anna Ewa Choromanska
arXiv Machine Learning
Sep 14

Reinforcement Learning over Patient Trajectories for Clinical Reasoning in EHR Foundation Models

The paper proposes a reinforcement learning (RL) fine‑tuning framework for electronic health record (EHR) foundation models, treating them as generative policies over patient trajectories. By framing clinical prediction tasks such as hospital readmission as event‑conditioned, time‑windowed reasoning problems and designing time‑aware, rollout‑sensitive rewards, the authors show that RL fine‑tuning consistently outperforms pre‑trained backbones and strong baselines. The approach enables smaller models to surpass larger pre‑trained models in data‑limited settings, induces positive transfer across tasks, and produces trajectories with stronger structural and semantic alignment to ground truth, improving downstream utility.

By Yuxin Xiao, Sheng Zhang, Chandan Singh, Tristan Naumann, Hoifung Poon, Jianfeng Gao, Xiaodong Liu
arXiv AI
Jul 2

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation

arXiv:2607. 00147v1 Announce Type: new Abstract: Rare disease differential diagnosis is a critical yet arduous clinical task, requiring physicians to identify precise phenotypes from complex, unstructured patient symptoms and execute intricate reasoning within a vast search space.

By Deyang Jiang, Haoran Wu, Ziyi Wang, Yiming Rong, Yunlong Zhao, Ye Jin, Bo Xu
arXiv Computation and Language
Aug 28

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models proposes a new framework to improve medical reasoning in LLMs. It introduces two key conditions—Causal Sufficiency and Proximal Learnability—to curate high-quality training trajectories, using agreement-based self-verification and dynamic entropy bounds. Experiments on medical multimodal and text-only benchmarks show that CARE outperforms competitors, reducing incorrect reasoning and enhancing training stability.

By Yucheng Zhou, Peng Luo, Qianning Wang, Chengzhong Xu, Jianbing Shen
arXiv AI
Aug 19

Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

The paper introduces ICD-Deepresearch, a workflow that combines foundation models for electronic health records (EHR) and language models with medical search and ICD dictionaries to forecast future ICD codes for upcoming clinical encounters. It evaluates candidate code transitions by linking patient evidence, external clinical relations, and exact code semantics within a fixed top‑K budget, using SparseEHR for initial priors, GPT‑5 for complementary forecasts, and a final selection step that validates, deduplicates, and ranks candidates. The method achieves patient‑averaged precision/recall of 24.60/35.09% on MIMIC‑III and 25.14/48.32% on MIMIC‑IV, with physicians rating 51–68% of its retrieved documents as useful, outperforming standalone GPT‑5 web search and Medical Deep Research.

By Junda Wang, Meysam Ghaffari, Akshat Choube, Mohsen Sharifi Renani, Hong Yu, Carlos Morato
arXiv AI
Aug 18

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.

By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
arXiv Computation and Language
Sep 23

Learning Diagnostic Reasoning for Decision Support in Toxicology

The paper introduces DeToxR, a reinforcement‑learning‑enhanced large language model designed to support decision making in acute toxicology cases. It fuses unstructured narratives from paramedics and patients with structured vital‑sign data to predict co‑ingested substances across 14 classes. In preliminary validation, DeToxR outperforms baseline models, achieving higher micro‑F1 and recall scores for poison identification.

By Nico Oberl\"ander, David Bani-Harouni, Tobias Zellner, Nassir Navab, Florian Eyer, Matthias Keicher
arXiv AI
Jun 3

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

arXiv:2606. 02802v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model structured longitudinal electronic health records (EHRs).

By Bo-Hong Wang, Baicheng Peng, Ruilin Wang, Jun Bai, Ziyang Song, Yue Li