arXiv Computer Vision

Gated Spatial Redundancy Projection for Pathology Transformer Attentions

arXiv AI
Jun 8

DaX: Learning General Pathology Representations Across Scales

arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.

By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv AI
Jun 2

Aligning Cellular Sheaves with Classifier Attention for Interpretable Weakly-Supervised Pathology Localization

arXiv:2606. 00092v1 Announce Type: cross Abstract: Weakly-supervised classification of whole-slide images with attention-based multiple instance learning (ABMIL) on top of foundation features now reaches near-saturation on Camelyon16 slide-level performance, but the corresponding attention maps are an imperfect localization signal: in clinical interpretation, a model that classifies correctly without firing on the actual lesion is hard to trust.

By Devansh Lalwani, Swapnil Bhat, Maulik Shah
arXiv Computer Vision
Sep 15

Weakly Supervised Spatial Grounding for Discriminative Attention-Based Ultrasound-Histopathology Alignment in Prostate Cancer Grading

arXiv:2609.15150v1 Announce Type: new Abstract: Unpaired cross-modal distillation transfers grade structure from histopathology into a micro-ultrasound (micro-US) encoder by aligning a pooled needle-...

By Obed Korshie Dzikunu, Emma Willis, Mohammad Mahdi Abootorabi, Mohamed Harmanani, Zhuoxin Guo, Ferdinand Luger, Adam Kinnaird, Brian Wodlinger, Parvin Mousavi, Purang Abolmaesumi
arXiv Machine Learning
Aug 26

RACR-MIL: Rank-aware contextual reasoning for weakly supervised grading of squamous cell carcinoma using whole slide images

RACR-MIL is a weakly‑supervised method for grading squamous cell carcinoma (SCC) from whole‑slide images, using an attention‑based multiple‑instance learning framework. It introduces a hybrid WSI graph to capture local tissue context and non‑local phenotypic dependencies, and applies rank‑ordering constraints on attention to prioritize higher‑grade tumor regions, mirroring pathologists’ diagnostic reasoning. The approach achieves state‑of‑the‑art performance, improving SCC grading accuracy by 3–9% over existing methods and up to 10% in tumor localization, and a pilot study showed pathologists reported increased grading efficiency in 60% of cases.

By Anirudh Choudhary, Mosbah Aouad, Krishnakant Saboo, Angelina Hwang, Jacob Kechter, Blake Bordeaux, Puneet Bhullar, David DiCaudo, Steven Nelson, Nneka Comfere, Emma Johnson, Olayemi Sokumbi, Jason Sluzevich, Leah Swanson, Dennis Murphree, Aaron Mangold, Ravishankar Iyer
arXiv Computer Vision
Sep 7

MultiAttenGastro: Multi-Dimensional Attention Augmentation for Gastrointestinal Endoscopy Classification

MultiAttenGastro is a plug‑and‑play attention framework that adds parallel 1‑D channel, 2‑D spatial, and 3‑D contextual heads to existing CNN and transformer backbones for gastrointestinal endoscopy classification. Across eight backbones and five public GI datasets, the framework improves performance on large‑gap datasets such as Kvasir‑Capsule but shows no benefit on small‑gap benchmarks like Kvasir‑v2, with mixed results elsewhere. Analysis using Centered Kernel Alignment indicates that the gains are linked to representational redundancy: low inter‑head redundancy under large domain gaps yields consistent improvements, while high redundancy under small gaps leads to losses.

By Sadhana Devarajan, Praveen Kumar Chandaliya, Dhruvin Jashvant Kumar Shah, Kishor Upla, Kiran Raja
arXiv Computer Vision
6d ago

Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images

The paper introduces SlideTIM, a transductive few‑shot classification method tailored for whole‑slide images (WSIs). SlideTIM extends the LC‑TIM approach by adding a spatial‑latent regularizer and a class‑distribution prior, ensuring that spatially and semantically similar patches receive consistent predictions and that predicted class proportions are calibrated. Experiments on four histology datasets show that SlideTIM outperforms existing TIM variants, boosting macro‑F1 scores by up to 8.1 percentage points over the best baseline and 19.4 percentage points over zero‑shot predictions at one shot.

By Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq, Christophe De Vleeschouwer
arXiv Computer Vision
6d ago

MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization

MDSkin-Net is a multi‑task skin lesion analysis framework that integrates Pattern Analysis priors into a hybrid CNN‑Transformer architecture. It introduces a Pattern Analysis‑Guided Attention Module (PAGAM) with improved Efficient Channel Attention, Multi‑Scale Spatial Attention, and Biased Asymmetry Attention, along with a multi‑scale spatial alignment regularization that uses segmentation masks as soft supervision. Trained only on the ISIC 2017 training split, the model achieves high segmentation and classification performance on multiple datasets, demonstrating strong zero‑shot generalization across different cohorts.

By Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas, Nikolaos Papanikolopoulos
arXiv Computer Vision
Aug 25

LanGuSTE: Language-Guided Coarse-to-Fine Patch Selection for Efficient Whole Slide Image Analysis

LanGuSTE is a patch‑selection framework for whole slide image analysis that uses vision‑language models and large language model knowledge. It introduces Cross‑Scale Visual Prompt Tuning to align low‑resolution and high‑resolution patches, and a coarse‑to‑fine selection module that encodes only informative high‑resolution patches. Experiments show LanGuSTE cuts overall processing time to about one‑third of the baseline while matching or surpassing diagnostic performance of exhaustive and state‑of‑the‑art methods.

By Yonghan Shin, Gangsu Kim, Won-Ki Jeong