arXiv AI By Junyeong Maeng, Eunsong Kang, Heung-Il Suk

STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation

Read the original on arXiv AI →

STRIVE is a new framework for longitudinal radiology report generation that separates clinical reasoning into distinct Diagnosis, Attribute, and Temporal Change agents, each producing explicit evidence. The Temporal Change agent is refined with a Progression-Aware GRPO reward that differentiates direction-preserving errors from reversals. Verification occurs twice: a Consistency Gate aligns agent outputs before report generation, and a Validation Agent ensures the final report is supported by the aggregated evidence. On the Longitudinal-MIMIC dataset, STRIVE achieves superior clinical efficacy and more than doubles Longitudinal Change Concordance compared to the strongest baseline.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 5

HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize

arXiv:2601. 03321v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision.

By Kun Zhao, Guodong Liu, Hui Ji, Siyuan Dai, Pan Wang, Jifeng Song, Chenghua Lin, Liang Zhan, Haoteng Tang
arXiv AI
Jun 16

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

arXiv:2603. 01131v3 Announce Type: replace-cross Abstract: Clinical diagnosis is a gradual process of evidence integration, in which physicians move from symptoms and medical history to examinations, competing hypotheses, disease relations, and treatment decisions.

By Yuqi Zhan, Xinyue Wu, Tianyu Lin, Yutong Bao, Xiaoyu Wang, Weihao Cheng, Huangwei Chen, Feiwei Qin, Zhu Zhu
arXiv Computation and Language
Sep 7

CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

CT‑ΔBench is a new benchmark designed to evaluate vision‑language models on longitudinal 3D medical imaging difference reporting. It provides patient‑level split data, change‑aware metrics, and physician‑validated references to assess clinically meaningful interval changes between two CT scans. The paper also introduces DeltaMed, a baseline model that directly reasons over paired CT scans, and compares it to an indirect two‑stage approach that first generates single‑timepoint reports before differencing.

By Kegeng Tang, Jingbo Wang, Shaogang Ren, Zihao Wang