arXiv AI By Kun Zhao, Guodong Liu, Hui Ji, Siyuan Dai, Pan Wang, Jifeng Song, Chenghua Lin, Liang Zhan, Haoteng Tang

HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize

Read the original on arXiv AI →

arXiv:2601. 03321v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

Calibrated Confidence Expression for Radiology Report Generation

The paper introduces ConRad, a reinforcement learning framework that fine‑tunes large vision‑language models to generate calibrated verbalized confidence estimates for radiology reports. ConRad offers both a single report‑level confidence score and a sentence‑level variant, trained with the GRPO algorithm and logarithmic scoring rewards to encourage truthful self‑assessment. Experiments show significant calibration improvements over existing methods, and clinical evaluation indicates that report‑level scores align well with clinicians’ judgments, enabling targeted review of low‑confidence statements.

By David Bani-Harouni, Chantal Pellegrini, Julian L\"uers, Su Hwan Kim, Markus Baalmann, Benedikt Wiestler, Rickmer Braren, Nassir Navab, Matthias Keicher
Hugging Face Trending Papers
Aug 20

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements.

arXiv AI
Jun 24

E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

arXiv:2606. 23888v1 Announce Type: cross Abstract: While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucinations and a lack of grounding in 3D CT data.

By Sijing Li, Zhongwei Qiu, Zhuoya Wang, Boxiang Yun, Zhenyu Yi, Jianwei Xu, Wenqiao Zhang, Yingda Xia, Ling Zhang