arXiv AI

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

arXiv:2607. 26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules.

arXiv AI
Jul 10

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

arXiv:2607. 08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning.

By Fan Ma, Mauro Giuffr\`e, Donald Wright, Kent McCann, Mark Iscoe, Lingfei Qian, Mingyang Jiang, Chi Wing Ng, Na Hong, Huan He, Cathy Shyr, Qingyu Chen, Lee Schwamm, Lucila Ohno-Machado, Hua Xu
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv AI
Sep 18

Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support

The paper introduces Oph‑Guid‑RAG, a multimodal retrieval‑augmented generation system tailored for ophthalmology clinical question answering. It treats each guideline page as an independent evidence unit, retrieving page images to preserve tables, flowcharts, and layout, and employs a controllable retrieval framework with routing and filtering to reduce noise. Evaluated on HealthBench, the system outperforms GPT‑5.2 and GPT‑5.4 on hard cases, achieving significant gains in overall score and accuracy, and ablation studies confirm the importance of reranking, routing, and retrieval design.

By Shuying Chen, Sen Cui, Zhong Cao
arXiv AI
Sep 15

VeriDx: Earning the Right to Diagnose with Disease-Centric Verification

VeriDx is a disease‑centric verification framework that links free‑form diagnostic reasoning to structured disease profiles, tracking whether each hypothesis satisfies, remains unresolved, or violates its clinical obligations. It exposes failures such as missing critical tests, unresolved differentials, ignored contradictions, unsupported claims, and premature closure. Applied to complex respiratory diagnosis, VeriDx reveals that many diagnostic errors stem from broken commitments made earlier in the reasoning process.

By Zhong Cao, Shuying Chen
arXiv AI
Aug 25

MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM

The paper introduces MACD, a Multi-Agent Clinical Diagnosis framework that enables large language models to self‑learn clinical knowledge through a multi‑agent pipeline of summarization, refinement, and application. MACD is extended into a human‑AI collaborative workflow where multiple diagnostician agents consult iteratively, guided by a judge agent and human oversight. Evaluation on the MIMIC‑MACD cohort shows significant gains in diagnostic accuracy—an average 11.6 percentage‑point improvement over authoritative knowledge for open‑weight LLMs and an 18.3‑percentage‑point boost over physician‑only diagnosis in text‑only vignettes.

By Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
arXiv AI
Aug 13

Teaching agentic AI to learn expert reasoning for rare disease diagnosis

arXiv:2606. 16149v3 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.

By Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr
arXiv Computation and Language
Aug 27

Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation

The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.

By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz