arXiv:2609.01470v1 Announce Type: new
Abstract: As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge. Large language...
By Charles Corbi\`ere, L\'eo Machado, Aubin Charley, Baptiste Callard, Pierre Manceron, Corentin Dancette
arXiv:2606. 17062v1 Announce Type: cross Abstract: Radiology report evaluation must distinguish clinical compatibility from surface similarity, because negation, laterality, or normal-abnormal polarity can reverse a finding.
By Zhenhong Yang, Zhuoyun Liu, Jintao Fei, Wen Tang, Shichao Quan, Jun Zhao, Jun Xu
arXiv:2609.27607v1 Announce Type: cross
Abstract: An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its pre...
By Jiaju Huang, Hao Yang, Xinyu Ma, Xinglong Liang, Kunyan Cai, Junqiang Ma, Shaobin Chen, Yue Sun, Tao Tan
arXiv:2607. 05880v1 Announce Type: cross Abstract: Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone.
By Suneeta Mall, Vladimir Nekrasov, Ashnil Kumar, Sajith Karunasena, Aiden Nibali, Alix Bird, Mateo Diaz Shine, Jarrel Seah
arXiv:2609.22281v1 Announce Type: new
Abstract: Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured...
By Benjamin Renoust, Pierre Baudot, Tiffany Foriel, Yousra Haddou, Charles Voyton, Pierre-Henri Siot, Ezequiel Geremia, Danny Francis, Jean-Christophe Brisset, Val\'erie Bourd\`es, Sylvain Bodard, Benoit Huet
arXiv:2608. 10505v1 Announce Type: new Abstract: Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content.
By Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz
arXiv:2607. 25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases.
By Mateusz Koz{\l}owski
Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions.
The study examines how differences in radiologists’ reporting styles—such as terminology, shorthand, formatting, and detail—affect the evaluation of AI-generated chest X‑ray reports. By quantifying the sensitivity of common metrics to these variations, the authors show that changes in reference reports can shift model rankings. They introduce a taxonomy of reporting variations and a rewriting method, ReRef, that preserves clinical meaning while altering style, and release a validated dataset of paired reference reports to aid future research.
By Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
The study evaluates whether a fine‑tuned open‑weight model (Gemma‑3‑12B) can match the performance of GPT‑4o in extracting multi‑label intracranial hemorrhage acuity from non‑contrast head‑CT reports. Using a 2×2 design that varied adaptation strategy (classification head vs. instruction fine‑tuning) and training‑data source (distilled real GPT‑4o labels vs. synthetic GPT‑4o‑generated reports), the distilled instruction‑tuned model achieved macro‑F1 scores comparable to GPT‑4o and surpassed the untuned base model. The key finding is that the source of training data—distilled real reports—was more important than the fine‑tuning method, and that the entire fine‑tuning and inference process fits on a single 24 GB consumer GPU.
By Aawez Mansuri, Kush Mehta, Mohammadreza Chavoshi, Jahanzaib Malik, Theodorus Dapamede, Frank Li, Rohan Isaac, Beatrice Brown-Mulry, Chiratidzo Rudado Sanyika, YoungSeok Jeon, Judy W. Gichoya, Ali Emami, Hari Trivedi
arXiv:2607. 14205v1 Announce Type: new Abstract: Federated learning (FL) enables multi-institutional training on clinical text without sharing raw data, but gradient inversion can reconstruct sensitive information from shared model updates.
By Santhosh Parampottupadam, Andres Martinez, Dimitrios Bounias, Sinem Sav, Klaus Maier-Hein, Ralf Floca
arXiv:2608. 00147v1 Announce Type: cross Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility.
By Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf, Lena Schmitzer, Lea Schumann, Jannik Kahmann, Friedrich Puttkammer, Johannes Moll, Jannik L\"ubberstedt, Zeineb Ben Chaaben, Anirudh Narayanan, Cosmin I. Bercea, Sebastian Ziegelmayer, Marcus R. Makowski, Daniel Rueckert, Lisa C. Adams, Keno K. Bressem