Hugging Face Trending Papers

Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)

Background: Disease severity is a multidimensional construct difficult to capture with rule-based approaches in Electronic Healthcare Records (EHR). Agentic large language model (LLM) systems could synthesise clinical evidence and reason over EHRs, but remain unevaluated for this task.

arXiv AI
Aug 28

Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset

The study evaluates a standalone large language model (LLM) versus a four‑step agentic pipeline for generating explanations of ICU mortality predictions on the eICU Demo dataset. XGBoost achieved an AUROC of 0.855 and an AUPRC of 0.332. In a 38‑case explanation subset, the standalone LLM produced one explanation with outcome leakage, while the agentic pipeline produced none; among 14 overlapping SHAP cases, the standalone LLM had higher SHAP alignment and direction consistency, whereas the agentic pipeline showed better guideline grounding, value specificity, and plausibility.

By Di Zhu, Chen Xie, Haoyun Zhang, Zihan Wei, Ziwei Wang, Jiazhao Shi, Ziyu Wang, Qiyang Xie
arXiv AI
Jun 16

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

arXiv:2603. 01131v3 Announce Type: replace-cross Abstract: Clinical diagnosis is a gradual process of evidence integration, in which physicians move from symptoms and medical history to examinations, competing hypotheses, disease relations, and treatment decisions.

By Yuqi Zhan, Xinyue Wu, Tianyu Lin, Yutong Bao, Xiaoyu Wang, Weihao Cheng, Huangwei Chen, Feiwei Qin, Zhu Zhu
arXiv Computation and Language
Sep 1

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

arXiv:2608.31128v1 Announce Type: new Abstract: Large language models (LLMs) offer promising clinical decision support but remain vulnerable to hallucinated facts, unsupported recommendations, and ci...

By Yung Wei Shueh, Zhi-Jie Chen, Chia-Hsuan Hsu, Hsin-Ling Hsu, Donghua Zhang, Chenwei Wu, Jun-En Ding, Tongze Zhang, Shihao Yang, Pengfei Hu, Fang-Ming Hung, Feng Liu
arXiv Computation and Language
Aug 27

Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation

The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.

By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
arXiv Computation and Language
Sep 22

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

The study evaluates large language models (LLMs) on sequential emergency department triage, where acuity labels are predicted from progressively longer nurse‑patient conversations. Six LLMs were tested at five checkpoints on simulated and physician‑authored dialogues, showing a decline from moderate‑to‑substantial agreement on full records to only fair‑to‑moderate agreement at each checkpoint. The models consistently anchor on chief complaint exchanges and fail to integrate later evidence, yielding low agreement with clinicians (QWK 0.295 vs. 0.887‑0.929) and concentrating predictions on ESI‑2 and ESI‑3. whyItMatters":"The findings reveal that LLMs, despite strong offline performance, cannot reliably handle the sequential nature of real‑time triage, highlighting a critical gap for safe deployment in emergency settings."

By Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv AI
Jun 30

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

arXiv:2606. 28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and substantial expert effort.

By Yen-Hsun Huang (Department of Education, Taipei Veterans General Hospital, Taipei, Taiwan), Yu-Shiou Lin (Department of Psychiatry, Taipei Veterans General Hospital, Taipei, Taiwan)
arXiv AI
Sep 2

AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.

By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv AI
Jul 29

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.

By Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf