arXiv AI

HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize

arXiv:2601. 03321v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision.

arXiv Computation and Language
Sep 23

Calibrated Confidence Expression for Radiology Report Generation

The paper introduces ConRad, a reinforcement learning framework that fine‑tunes large vision‑language models to generate calibrated verbalized confidence estimates for radiology reports. ConRad offers both a single report‑level confidence score and a sentence‑level variant, trained with the GRPO algorithm and logarithmic scoring rewards to encourage truthful self‑assessment. Experiments show significant calibration improvements over existing methods, and clinical evaluation indicates that report‑level scores align well with clinicians’ judgments, enabling targeted review of low‑confidence statements.

By David Bani-Harouni, Chantal Pellegrini, Julian L\"uers, Su Hwan Kim, Markus Baalmann, Benedikt Wiestler, Rickmer Braren, Nassir Navab, Matthias Keicher
Hugging Face Trending Papers
Aug 20

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements.

arXiv AI
Jun 24

E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

arXiv:2606. 23888v1 Announce Type: cross Abstract: While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucinations and a lack of grounding in 3D CT data.

By Sijing Li, Zhongwei Qiu, Zhuoya Wang, Boxiang Yun, Zhenyu Yi, Jianwei Xu, Wenqiao Zhang, Yingda Xia, Ling Zhang
arXiv Computation and Language
Aug 28

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models proposes a new framework to improve medical reasoning in LLMs. It introduces two key conditions—Causal Sufficiency and Proximal Learnability—to curate high-quality training trajectories, using agreement-based self-verification and dynamic entropy bounds. Experiments on medical multimodal and text-only benchmarks show that CARE outperforms competitors, reducing incorrect reasoning and enhancing training stability.

By Yucheng Zhou, Peng Luo, Qianning Wang, Chengzhong Xu, Jianbing Shen
arXiv AI
Aug 26

STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation

STRIVE is a new framework for longitudinal radiology report generation that separates clinical reasoning into distinct Diagnosis, Attribute, and Temporal Change agents, each producing explicit evidence. The Temporal Change agent is refined with a Progression-Aware GRPO reward that differentiates direction-preserving errors from reversals. Verification occurs twice: a Consistency Gate aligns agent outputs before report generation, and a Validation Agent ensures the final report is supported by the aggregated evidence. On the Longitudinal-MIMIC dataset, STRIVE achieves superior clinical efficacy and more than doubles Longitudinal Change Concordance compared to the strongest baseline.

By Junyeong Maeng, Eunsong Kang, Heung-Il Suk
arXiv AI
Aug 5

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.

By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu