arXiv AI

CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders

arXiv Machine Learning
Jul 30

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.

By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
arXiv Computer Vision
4d ago

MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models

MVC-Bench is a new benchmark designed to evaluate the calibration of vision‑language models (VLMs) and medical VLMs (Medical‑VLMs) for medical image classification. It tests calibration across robustness to modality, backbone, and domain shift; effectiveness of calibration strategies and prompt‑tuning methods; and stability under prompt‑template and random‑seed variations. The benchmark includes eight backbones, three medical modalities (fundus imaging, histopathology, chest X‑ray), and compares post‑hoc, train‑time, and zero‑shot calibration approaches, reporting accuracy, Expected Calibration Error (ECE), Maximum Calibration Error (MCE), and Adaptive Calibration Error (ACE) over 1,638 experiments, while also proposing a Multi‑Class Margin (MCM) regularization technique that improves ECE in most settings.

By Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan, Muhammad Akhtar Munir Sujair Ibrahim, Mohamed Rafeek Mareer Ahamed, Yutong Xie, Imran Razzak, Muhammad Haris Khan
arXiv AI
Jul 23

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

arXiv:2607. 20274v1 Announce Type: cross Abstract: Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure.

By Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather, Daniel Truhn
arXiv Computer Vision
5d ago

Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?

The study evaluates 15 frozen hematology foundation-model embeddings across four single‑cell acquisition domains, finding that while in‑domain accuracy is near‑saturated (macro‑F1 0.98–0.997), cross‑dataset performance drops dramatically (34–72%) and model rankings shift. Probe‑dependent rank transfer is observed, with 1‑NN retrieval more stable than linear heads, yet neither reliably predicts target robustness. Calibration deteriorates off‑domain (ECE rises from 0.004 to 0.35), and exposure to internal cohorts confounds shift analysis; a training‑free pseudo‑label‑balanced feature normalization (CBR) modestly improves target‑prior robustness and calibration. whyItMatters":"The findings highlight that frozen hematology foundation models, though accurate in‑domain, may fail under realistic scanner, site, and class‑prior shifts, underscoring the need for comprehensive audits of accuracy, calibration, exposure, and robustness before clinical deployment."

By Jai Kumar Sharma, Peeyush Tapadiya
arXiv Computer Vision
5d ago

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

The paper introduces a label‑free method called AURCC for selecting the best foundational model for medical image classification when the target domain lacks labels. AURCC uses a pseudo‑label discrepancy computed by the SUDO framework to score models without fine‑tuning. Experiments on chest X‑ray data across three inter‑hospital shifts show that AURCC closely matches the true model ranking, outperforming simple source‑accuracy baselines especially when source data are limited.

By Juan I\~naki Larrea, Lucas Mansilla, Enzo Ferrante
Hugging Face Trending Papers
5d ago

MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models

MVC-Bench is a calibration-focused benchmark for medical vision‑language models, evaluating how well these models express confidence across different modalities, backbones, and domain shifts. It tests robustness to modality, backbone, and domain changes, the effectiveness of calibration and prompt‑tuning strategies, and stability under prompt‑template and random‑seed variations. The benchmark includes 1638 experiments, reporting accuracy and Expected Calibration Error (ECE) along with other calibration metrics, and introduces a simple train‑time calibration method, Multi‑Class Margin (MCM) regularization, that achieves the lowest ECE in most settings.

arXiv AI
Aug 18

Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

arXiv:2608. 16198v1 Announce Type: cross Abstract: Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing.

By Fabian Gr\"oger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly
arXiv AI
Jun 8

DaX: Learning General Pathology Representations Across Scales

arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.

By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu