arXiv AI

STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation

STRIVE is a new framework for longitudinal radiology report generation that separates clinical reasoning into distinct Diagnosis, Attribute, and Temporal Change agents, each producing explicit evidence. The Temporal Change agent is refined with a Progression-Aware GRPO reward that differentiates direction-preserving errors from reversals. Verification occurs twice: a Consistency Gate aligns agent outputs before report generation, and a Validation Agent ensures the final report is supported by the aggregated evidence. On the Longitudinal-MIMIC dataset, STRIVE achieves superior clinical efficacy and more than doubles Longitudinal Change Concordance compared to the strongest baseline.

arXiv AI
Aug 5

HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize

arXiv:2601. 03321v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision.

By Kun Zhao, Guodong Liu, Hui Ji, Siyuan Dai, Pan Wang, Jifeng Song, Chenghua Lin, Liang Zhan, Haoteng Tang
arXiv AI
Jun 16

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

arXiv:2603. 01131v3 Announce Type: replace-cross Abstract: Clinical diagnosis is a gradual process of evidence integration, in which physicians move from symptoms and medical history to examinations, competing hypotheses, disease relations, and treatment decisions.

By Yuqi Zhan, Xinyue Wu, Tianyu Lin, Yutong Bao, Xiaoyu Wang, Weihao Cheng, Huangwei Chen, Feiwei Qin, Zhu Zhu
arXiv Computation and Language
Sep 7

CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

CT‑ΔBench is a new benchmark designed to evaluate vision‑language models on longitudinal 3D medical imaging difference reporting. It provides patient‑level split data, change‑aware metrics, and physician‑validated references to assess clinically meaningful interval changes between two CT scans. The paper also introduces DeltaMed, a baseline model that directly reasons over paired CT scans, and compares it to an indirect two‑stage approach that first generates single‑timepoint reports before differencing.

By Kegeng Tang, Jingbo Wang, Shaogang Ren, Zihao Wang
arXiv AI
Aug 26

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

EviDx is a new framework for evidence-aware active diagnosis that pairs patient-specific diagnostic environments with a clinical scaffold and an observer-guided runtime harness. The framework constructs interactive environments from raw clinical cases, organizes role-specialized agents and evidence tools, and regulates diagnostic termination by tracking uncertainty and evidence coverage. Experiments demonstrate that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.

By Lihang Zeng, Shaoting Zhang, Xiaofan Zhang