arXiv Machine Learning

Who Gets Missed in the Tail? Thresholded Subgroup Underdiagnosis in Long-Tailed Chest X-ray Classification

arXiv:2607. 07717v1 Announce Type: new Abstract: In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroups.

arXiv Machine Learning
Aug 27

FRAME: separating sampling variation from representational cause in medical imaging fairness

The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.

By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
arXiv AI
Sep 25

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Med-AR introduces two autoregressive vision‑language models, Med‑AR‑8B and Med‑AR‑2B, pretrained on structured radiology reports, abnormality‑focused text, and region annotations to address long‑tailed chest X‑ray classification. The models outperform existing contrastive, self‑supervised, and supervised encoders—including Med‑CLIP, CheXFound, EVA‑Base, ARK, and BioViL‑T—across PadChest, MIMIC‑CXR, and CheXpert, achieving higher mean AUROC and AUPRC for head, medium, and tail findings and lower excess area under the risk‑coverage curve. Med‑AR also demonstrates improved selective‑prediction performance, with Med‑AR‑8B raising tail‑label mean AUPRC on MIMIC‑CXR from 0.1033 to 0.1441 and Med‑AR‑2B delivering the strongest discrimination on PadChest.

By Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha
arXiv Computer Vision
Sep 4

Subgroup performance analysis of adaptation strategies for chest X-ray foundation models

The study examines how three parameter‑efficient adaptation methods—linear heads on the raw CLS token, an MLP, and an attention‑pooling module—affect pathology classification accuracy and subgroup fairness when applied to a frozen Rad‑DINO chest X‑ray encoder. Using the MIMIC‑CXR dataset, the authors evaluate eight pathologies across race, sex, and imaging‑view subgroups, finding that attention pooling yields the best overall performance and encodes protected attributes most strongly, yet higher performance does not consistently reduce subgroup disparities. The results show that attribute encoding strength and layer choice do not reliably predict fairness outcomes, indicating that fairness must be assessed directly for each task.

By Dhruv Gupta, Emma A. M. Stanley, Fabio De Sousa Ribeiro, Sujal R. Desai, Ben Glocker
arXiv AI
Aug 12

RadFusion: Towards Threshold-Controllable Radiology Report Generation

arXiv:2608. 10505v1 Announce Type: new Abstract: Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content.

By Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz
Hugging Face Trending Papers
Aug 11

RadFusion: Towards Threshold-Controllable Radiology Report Generation

Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions.