arXiv AI

DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features

arXiv:2608. 04652v1 Announce Type: cross Abstract: Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease grading, where labels encode a progression of severity.

arXiv Computer Vision
Aug 25

Ordinal Diffusion Models for Color Fundus Images

The paper introduces an ordinal latent diffusion model for generating color fundus images that incorporates the ordered structure of diabetic retinopathy (DR) severity, using a scalar disease representation instead of categorical conditioning. Evaluations on the EyePACS dataset show improved visual realism, with reduced Fréchet inception distance for most stages and a higher quadratic weighted κ from 0.79 to 0.87. Interpolation experiments demonstrate the model captures a continuous spectrum of disease progression derived from coarse, ordered labels.

By Gustav Schmidt, Philipp Berens, Sarah M\"uller
arXiv Computer Vision
Sep 14

Early Intervention for VFM-based Multimodal Medical Image Classification

The paper introduces an Early Intervention (EI) framework for multimodal medical image classification that addresses two key challenges: limited exploitation of complementary multimodal information and scarcity of labeled data for Vision Foundation Models (VFMs). EI treats one modality as the target and uses high‑level semantic tokens from other modalities as intervention tokens to guide the target’s embedding early in the process. The authors also propose Mixture of varied‑rank LoRAs (MoR) for efficient VFM adaptation, and demonstrate the method’s effectiveness on retinal, skin, and knee medical image datasets.

By Qijie Wei, Hailan Lin, Xirong Li
arXiv AI
Jun 8

DaX: Learning General Pathology Representations Across Scales

arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.

By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv AI
Jun 24

HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction

arXiv:2603. 19957v2 Announce Type: replace-cross Abstract: Pathology reports are structured, multi-granular documents encoding diagnostic conclusions, histological grades, and ancillary test results across one or more anatomical sites; yet existing pathology vision-language models (VLMs) reduce this output to a flat label or free-form text.

By Ruicheng Yuan, Zhenxuan Zhang, Anbang Wang, Liwei Hu, Xiangqian Hua, Yaya Peng, Jiawei Luo, Guang Yang