Hugging Face Trending Papers

Layer Selection in VLMs for Zero-Shot OOD Detection via Multi-Resolution Entropy Estimation

The paper investigates zero‑shot out‑of‑distribution (OOD) detection in medical imaging using vision‑language models (VLMs). It shows that intermediate layers, rather than only the final layer, provide valuable OOD signals and that the best layer depends on the imaging modality. To overcome instability in entropy‑based layer selection, the authors introduce a multi‑resolution entropy estimation that aggregates histogram statistics across scales, achieving consistent improvements over state‑of‑the‑art methods on the MIDOG and OASIS benchmarks.

Hugging Face Trending Papers
Jul 23

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncertain source-domain predictions.

arXiv Computer Vision
Aug 26

Example-based Robust Abnormality Detection with Minimal Annotations using Exemplar Med-DETR

arXiv:2608.24281v1 Announce Type: new Abstract: Reducing annotation requirements remains a key challenge in developing robust medical object detectors. To address this, Vision-Language (VL) object de...

By Sheethal Bhat, Bogdan Georgescu, Awais Mansoor, Mathias Zinnen, Pranjal Sahu, Florin C. Ghesu, Sasa Grbic, Andreas Maier
Hugging Face Trending Papers
Jul 27

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence.

arXiv AI
Jun 8

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

arXiv:2606. 06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy.

By Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
arXiv Machine Learning
Jul 27

Autoregressive EHR Foundation Models with Multimodal Inputs

arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.

By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
arXiv AI
2d ago

ProtoDCS: Towards Robust and Efficient Open-Set Test-Time Adaptation for Vision-Language Models

ProtoDCS introduces a robust open‑set test‑time adaptation framework for vision‑language models, addressing the challenge of simultaneously handling covariate‑shifted in‑distribution (csID) and out‑of‑distribution (csOOD) data. It replaces brittle thresholding with a double‑check separation using a probabilistic Gaussian Mixture Model and employs an evidence‑driven adaptation strategy that updates prototypes efficiently, reducing overconfidence and computational cost. Experiments on CIFAR‑10/100‑C and Tiny‑ImageNet‑C show state‑of‑the‑art performance, improving both known‑class accuracy and OOD detection metrics.

By Wei Luo, Yangfan Ou, Jin Deng, Zeshuai Deng, Xiquan Yan, Zhiquan Wen, Mingkui Tan