arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.
By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
arXiv:2605.07785v3 Announce Type: replace
Abstract: Concept Bottleneck Models (CBMs) in medical imaging aim to improve model interpretability by predicting intermediate clinical concepts before final...
By Amy Rafferty, Rishi Ramaesh, Ajitha Rajan
arXiv:2608. 09774v1 Announce Type: cross Abstract: Thoracic pathologies rarely occur in isolation, yet standard multi-label classifiers rely on shared global descriptors, discarding \emph{where} findings lie and \emph{how} they co-occur.
By Akash Gogineni, Nagur Shareef Shaik, Aasrith Mandava, Adnan Masood, Dong Hye Ye
Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.
By Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha
Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
The CXR‑LT 2026 Challenge introduces a multi‑center, long‑tailed chest X‑ray classification benchmark with over 145,000 radiologist‑annotated images from PadChest and NIH datasets. It defines two core tasks: robust multi‑label classification on 30 known classes and open‑world generalization to 6 unseen rare disease classes. The paper outlines data collection, annotation, solution strategies, and evaluates performance across head‑vs‑tail, calibration, and cross‑center gaps, noting that vision‑language models improve in‑distribution and zero‑shot performance but rare‑finding detection under multi‑center shift remains difficult.
By Hexin Dong, Yi Lin, Pengyu Zhou, Fengnian Zhao, Alan Clint Legasto, Juno Cho, Dohui Kim, Justin Namuk Kim, Mingeon Kim, Sunwoo Kwak, Gabriel Moy\`a-Alcover, Ky Trung Nguyen, Thanh-Huy Nguyen, Ha-Hieu Pham, Huy-Hieu Pham, Huy Le Pham, Nikhileswara Rao Sulake, Aina Tur-Serrano, Ruichi Zhang, Ang Zu, Adam E. Flanders, Zhiyong Lu, Ronald M. Summers, Mingquan Lin, Hao Chen, Yuzhe Yang, George Shih, Yifan Peng