arXiv Computer Vision By Jai Kumar Sharma, Peeyush Tapadiya

Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?

Read the original on arXiv Computer Vision →

The study evaluates 15 frozen hematology foundation-model embeddings across four single‑cell acquisition domains, finding that while in‑domain accuracy is near‑saturated (macro‑F1 0.98–0.997), cross‑dataset performance drops dramatically (34–72%) and model rankings shift. Probe‑dependent rank transfer is observed, with 1‑NN retrieval more stable than linear heads, yet neither reliably predicts target robustness. Calibration deteriorates off‑domain (ECE rises from 0.004 to 0.35), and exposure to internal cohorts confounds shift analysis; a training‑free pseudo‑label‑balanced feature normalization (CBR) modestly improves target‑prior robustness and calibration. whyItMatters":"The findings highlight that frozen hematology foundation models, though accurate in‑domain, may fail under realistic scanner, site, and class‑prior shifts, underscoring the need for comprehensive audits of accuracy, calibration, exposure, and robustness before clinical deployment."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Aug 12

Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

arXiv:2608. 10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios.

By Carlos Zamora, Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos
arXiv Computer Vision
4d ago

MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models

MVC-Bench is a new benchmark designed to evaluate the calibration of vision‑language models (VLMs) and medical VLMs (Medical‑VLMs) for medical image classification. It tests calibration across robustness to modality, backbone, and domain shift; effectiveness of calibration strategies and prompt‑tuning methods; and stability under prompt‑template and random‑seed variations. The benchmark includes eight backbones, three medical modalities (fundus imaging, histopathology, chest X‑ray), and compares post‑hoc, train‑time, and zero‑shot calibration approaches, reporting accuracy, Expected Calibration Error (ECE), Maximum Calibration Error (MCE), and Adaptive Calibration Error (ACE) over 1,638 experiments, while also proposing a Multi‑Class Margin (MCM) regularization technique that improves ECE in most settings.

By Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan, Muhammad Akhtar Munir Sujair Ibrahim, Mohamed Rafeek Mareer Ahamed, Yutong Xie, Imran Razzak, Muhammad Haris Khan