arXiv Machine Learning

Evaluating the trustworthiness of the Fr\'echet Inception Distance with stochastic embedding representations

arXiv:2601. 21979v2 Announce Type: replace Abstract: Feature embeddings acquired from pretrained models are widely used in medical applications of deep learning to assess the characteristics of datasets; e.

arXiv Computer Vision
Sep 25

Interpretable Similarity of Synthetic Image Utility

The paper introduces Interpretable Utility Similarity (IUS), a novel metric for quantifying how closely a synthetic image set matches a real image set in terms of usefulness for deep‑learning clinical decision support systems. IUS is interpretable, leveraging generalized neural additive models to explain why one synthetic dataset may outperform another based on clinically relevant image features. Experiments on color medical imaging modalities—endoscopic, dermoscopic, and fundus—show that selecting synthetic images with high IUS can boost classification performance by up to 54.6%, and the method also generalizes to grayscale X‑ray and ultrasound data.

By Panagiota Gatoula, George Dimas, Dimitris K. Iakovidis
arXiv Machine Learning
Aug 5

Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation

arXiv:2608. 03990v1 Announce Type: new Abstract: Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications.

By Seyed Kahaki, Shijie Li, Weijie Chen, Nicholas Petrick
arXiv Computer Vision
Aug 31

Image Augmentation as Test Generation for Deep Learning-Based Image Retrieval Systems

The paper reviews 50 image augmentation and generation techniques, categorizing them into ten groups, and conducts a large‑scale empirical study to assess their effectiveness as test generators for embedding‑based image retrieval systems. Using Amazon Titan and OpenCLIP embeddings, the authors evaluate the techniques across four dimensions—embedding‑space similarity, embedding uncertainty, semantic realism, and retrieval failure rate—on CIFAR‑10, ImageNet‑1K, and an industrial dataset. Results show that weather simulation and SaSPA yield the highest uncertainty and failure rates while maintaining realistic visuals, whereas GAN‑based methods produce low realism due to synthetic artifacts.

By Yehan De Silva, Anirudh Sridhar, Armin Lotfy, Nafiseh Kahani, Yvan Labiche, Ziyu Wang, Frank Ouyang, Clare Carty, Azalia Shamsaei
Hugging Face Trending Papers
Jul 28

Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer

Background: Early prediction of distant metastasis (DM) risk in head and neck cancer (HNC) can enable timely interventions that may improve treatment outcomes. Many current machine learning methods rely on prior knowledge of the region of interest such as tumor segmentations, which require expert knowledge, is time-consuming and introduces user-dependent variability.

arXiv Computer Vision
Aug 27

What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift

The paper investigates how medical vision‑language models (VLMs) behave when faced with distribution shifts such as changes in acquisition domain, supervision, or evaluation protocol. Using datasets like NIH ChestXray14, CheXpert, PadChest, and OpenI, the authors isolate cross‑dataset visual transfer, evaluate multimodal alignment, and quantify source‑proxy leakage in frozen embeddings. They find that self‑supervised visual initialization improves transfer, adversarial adaptation is only marginally helpful, and that multimodal retrieval performance drops under external stress tests while source‑proxy information remains recoverable, highlighting hidden failure modes in medical VLMs.

By Ayoub Louaye Bouaziz, Lokmane Chebouba, Yassine Himeur
arXiv Computer Vision
Sep 22

What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization

arXiv:2609.24691v1 Announce Type: new Abstract: Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the late...

By Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa, Julian McGinnis, Jiancheng Yang, Daniel Rueckert, Jiazhen Pan
arXiv AI
Aug 17

CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

arXiv:2608. 13939v1 Announce Type: cross Abstract: Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5).

By Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li