The study evaluates 15 frozen hematology foundation-model embeddings across four single‑cell acquisition domains, finding that while in‑domain accuracy is near‑saturated (macro‑F1 0.98–0.997), cross‑dataset performance drops dramatically (34–72%) and model rankings shift. Probe‑dependent rank transfer is observed, with 1‑NN retrieval more stable than linear heads, yet neither reliably predicts target robustness. Calibration deteriorates off‑domain (ECE rises from 0.004 to 0.35), and exposure to internal cohorts confounds shift analysis; a training‑free pseudo‑label‑balanced feature normalization (CBR) modestly improves target‑prior robustness and calibration.
whyItMatters":"The findings highlight that frozen hematology foundation models, though accurate in‑domain, may fail under realistic scanner, site, and class‑prior shifts, underscoring the need for comprehensive audits of accuracy, calibration, exposure, and robustness before clinical deployment."
By Jai Kumar Sharma, Peeyush Tapadiya
arXiv:2608.22059v1 Announce Type: cross
Abstract: Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited....
By Xingtao Lin, Hangqi Ren, Caiwan Sun, You Chen
arXiv:2608. 10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios.
By Carlos Zamora, Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos
arXiv:2607. 24519v2 Announce Type: replace Abstract: Pretrained EEG foundation models are proposed for clinical decoding, but whether reported gains transfer across populations or survive negative controls is unclear.
By Marzieh Zare
arXiv:2607. 00385v2 Announce Type: replace-cross Abstract: Automated malaria diagnosis from blood smear microscopy is a critical global health AI challenge; expert scarcity remains the primary diagnostic bottleneck.
By Kaysarul Anas Apurba, Md Hasibul Hasan, Mohammed Ali, Tanzilur Rahman
arXiv:2607. 25497v1 Announce Type: cross Abstract: Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centres.
By Cl\'ement Grisi, Jeroen van der Laak, Geert Litjens
MVC-Bench is a new benchmark designed to evaluate the calibration of vision‑language models (VLMs) and medical VLMs (Medical‑VLMs) for medical image classification. It tests calibration across robustness to modality, backbone, and domain shift; effectiveness of calibration strategies and prompt‑tuning methods; and stability under prompt‑template and random‑seed variations. The benchmark includes eight backbones, three medical modalities (fundus imaging, histopathology, chest X‑ray), and compares post‑hoc, train‑time, and zero‑shot calibration approaches, reporting accuracy, Expected Calibration Error (ECE), Maximum Calibration Error (MCE), and Adaptive Calibration Error (ACE) over 1,638 experiments, while also proposing a Multi‑Class Margin (MCM) regularization technique that improves ECE in most settings.
By Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan, Muhammad Akhtar Munir Sujair Ibrahim, Mohamed Rafeek Mareer Ahamed, Yutong Xie, Imran Razzak, Muhammad Haris Khan
arXiv:2606. 04365v1 Announce Type: cross Abstract: Radiology reports describe kidney lesions by type, size, enhancement, and attenuation, yet existing 3D methods predict only at the patient or organ level.
By Renjie Liang, Zhengkang Fan, Jinqian Pan, Chenkun Sun, Jiang Bian, Russell Terry, Jie Xu
arXiv:2407. 13632v2 Announce Type: replace-cross Abstract: Deploying deep learning-based imaging tools across various clinical sites poses significant challenges due to inherent domain shifts and regulatory hurdles associated with site-specific fine-tuning.
By Abhijeet Parida, Antonia Alomar, Zhifan Jiang, Pooneh Roshanitabrizi, Austin Tapp, Maria Ledesma-Carbayo, Ziyue Xu, Syed Muhammed Anwar, Marius George Linguraru, Holger R. Roth
MVC-Bench is a calibration-focused benchmark for medical vision‑language models, evaluating how well these models express confidence across different modalities, backbones, and domain shifts. It tests robustness to modality, backbone, and domain changes, the effectiveness of calibration and prompt‑tuning strategies, and stability under prompt‑template and random‑seed variations. The benchmark includes 1638 experiments, reporting accuracy and Expected Calibration Error (ECE) along with other calibration metrics, and introduces a simple train‑time calibration method, Multi‑Class Margin (MCM) regularization, that achieves the lowest ECE in most settings.
The paper introduces SWIFT, a Swin V2‑based model pretrained on 10,444 3D CT volumes and fine‑tuned for rectal cancer segmentation on T2‑weighted MRI. Four configurations—full fine‑tuning (SWIFT), decoder compression (SWIFTe), low‑rank adaptation (SWIFTe‑LoRA), and a LoRA‑decoder ensemble (SWIFTe‑LDE4)—were evaluated on 247 cases, showing that SWIFTe reduces parameters by 70.1% while improving tumor detection and radiomic agreement. The study also demonstrates a trade‑off between detection and boundary agreement, and highlights that SWIFTe‑LDE4 achieves the lowest calibration errors after temperature scaling.
By Aneesh Rangnekar, Jorge Tapias Gomez, Joseph O Deasy, Harini Veeraraghavan
arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.
By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu