arXiv AI
Aug 11

Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation

arXiv:2605. 30716v2 Announce Type: replace-cross Abstract: Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolution, long visual-token sequences, and the complexity of case-level reasoning, where a single case may contain multiple WSIs with heterogeneous tissues and ambiguous findings.

By Zhiyuan Yang, Jiahao Cheng, Vincent Quoc-Huy Trinh, Mahdi S. Hosseini
arXiv Computer Vision
Aug 25

LanGuSTE: Language-Guided Coarse-to-Fine Patch Selection for Efficient Whole Slide Image Analysis

LanGuSTE is a patch‑selection framework for whole slide image analysis that uses vision‑language models and large language model knowledge. It introduces Cross‑Scale Visual Prompt Tuning to align low‑resolution and high‑resolution patches, and a coarse‑to‑fine selection module that encodes only informative high‑resolution patches. Experiments show LanGuSTE cuts overall processing time to about one‑third of the baseline while matching or surpassing diagnostic performance of exhaustive and state‑of‑the‑art methods.

By Yonghan Shin, Gangsu Kim, Won-Ki Jeong
Hugging Face Trending Papers
Jul 26

PathSelect: Sequential Token Selection for Whole Slide Pathology

Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominantly rely on spatial sampling or training-free pruning, which risk diluting weak but informative signals, leading to the loss of critical diagnostic evidence due to the spatially diffuse nature of pathological cues.

arXiv AI
Jun 26

MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Understanding in Vision-Language Models

arXiv:2603. 11625v2 Announce Type: replace-cross Abstract: While specialized Medical Vision-Language Models (VLMs) have achieved remarkable success in interpreting 2D and 3D medical modalities, their deployment for 3D volumetric data remains constrained by significant computational inefficiencies.

By Shengyuan Liu, Zanting Ye, Yunrui Lin, Chen Hu, Wanting Geng, Xu Han, Bulat Ibragimov, Yefeng Zheng, Yixuan Yuan
arXiv Computer Vision
4d ago

Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images

The paper introduces SlideTIM, a transductive few‑shot classification method tailored for whole‑slide images (WSIs). SlideTIM extends the LC‑TIM approach by adding a spatial‑latent regularizer and a class‑distribution prior, ensuring that spatially and semantically similar patches receive consistent predictions and that predicted class proportions are calibrated. Experiments on four histology datasets show that SlideTIM outperforms existing TIM variants, boosting macro‑F1 scores by up to 8.1 percentage points over the best baseline and 19.4 percentage points over zero‑shot predictions at one shot.

By Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq, Christophe De Vleeschouwer