arXiv Computer Vision

Semantic-Aware Subgraph State Space Model for WSI Classification in Histopathology

The paper introduces the Semantic-Aware Subgraph State Space Model (SASG-SSM) for classifying whole slide images (WSIs) in histopathology. It groups spatially connected patches into semantic-aware subgraphs that preserve internal spatial organization, then uses a graph neural network encoder combined with a Mamba-based state space encoder to capture both local and global contextual information. Experiments on four WSI subtyping datasets show consistent performance gains over state‑of‑the‑art methods, with additional robustness in small‑cohort and few‑shot scenarios.

Hugging Face Trending Papers
Sep 3

Semantic-Aware Subgraph State Space Model for WSI Classification in Histopathology

The paper introduces the Semantic-Aware Subgraph State Space Model (SASG-SSM) for classifying whole slide images in histopathology. It groups spatially connected patches into semantic subgraphs that preserve internal spatial organization, then uses a graph neural network encoder combined with a Mamba-based state space encoder to integrate local and global contextual information. Experiments on four WSI subtyping datasets show consistent performance gains over state‑of‑the‑art methods, with additional robustness in small‑cohort and few‑shot scenarios.

arXiv AI
Sep 10

Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation

SCOUT is a concept‑grounded multimodal transformer that generates whole‑slide pathology reports by integrating local histological patterns, whole‑slide context, and expert‑curated diagnostic concepts. It uses evolving visual representations and recursively updated slide‑ and concept‑conditioned representations, with separate attention pathways during decoding that are fused adaptively for each token. Evaluated on TCGA‑BRCA, HistAI, and REG‑2025, SCOUT outperformed existing methods, improving BLEU, METEOR, and ROUGE‑L scores and raising the Clinical Report Quality Score on REG‑2025.

By Suryakant Singh, Saarthak Kapse, Joel Saltz, Prateek Prasanna
arXiv Computer Vision
Sep 25

Context-aware Skin Cancer Epithelial Cell Classification with Scalable Graph Transformers

The paper introduces scalable Graph Transformers for classifying healthy versus tumor epithelial cells in whole-slide images of cutaneous squamous cell carcinoma. By constructing a full‑WSI cell graph and incorporating morphological, texture, and neighboring cell class features, the proposed SGFormer and DIFFormer models outperform traditional image‑based methods, achieving balanced accuracies above 85% on single‑WSI tests and 83.6% on multi‑WSI evaluations. The study demonstrates that preserving tissue‑level context through graph representations improves classification of morphologically similar cell types.

By Lucas Sanc\'er\'e, No\'emie Moreau, Katarzyna Bozek
arXiv AI
Sep 2

SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning

SlideBank is a training‑free framework that turns each whole‑slide image into a persistent, concept‑indexed evidence bank. It performs coarse‑to‑fine exploration to locate informative regions and multi‑scale views, converts them into explicit morphological observations, and anchors pathology signals to the supporting patches and slide coordinates. During inference, questions are routed to relevant signals and evidence scales, and a confidence‑based cross‑level consensus integrates global, regional, and patch evidence, achieving high accuracy on WSI‑VQA and SlideBench‑BCNB while enabling consistent re‑phrasing and reduced inference cost.

By Beidi Zhao, Gexin Huang, Ciro Zhang, Anqi Li, Yusheng Tan, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li
arXiv AI
Sep 15

End-to-End Cell Detection via Instance-aware Graph Modeling

The paper introduces an end‑to‑end framework for detecting and classifying cells in pathology images by jointly modeling visual features and instance‑level interactions. It employs a dynamic graph construction module that builds cell graphs from learnable queries and an instance‑aware graph network that filters and reorganizes features, integrating appearance and relational evidence. Experiments on multiple staining protocols show the method surpasses existing approaches in both detection and classification accuracy.

By Ruochen Liu, Yalin Zheng, Jingxin Liu, Jianfeng Zhang, Shoujun Huang, Dexing Kong, Haofeng Li, Wei Lou
arXiv Computer Vision
Sep 16

Hyper-RED: Scalable Event Pre-training via Semantic Hypergraph Distillation

Hyper-RED introduces a scalable image-to-event pretraining framework that transfers high‑order semantic structures via hypergraphs, avoiding rigid pixel‑wise alignment. By constructing image, event, and cross‑modal hypergraphs and applying a hypergraph relational distillation loss, the method preserves local relational consistency and event‑specific characteristics while inheriting image‑derived semantic organization. Experiments across five event datasets show consistent scaling from ViT‑S to ViT‑L and state‑of‑the‑art performance.

By Meisen Wang, Zhiqiang Tian, Wei Bao, Chengjie Wang, Shaoyi Du, Siqi Li
arXiv Computer Vision
4d ago

Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images

The paper introduces SlideTIM, a transductive few‑shot classification method tailored for whole‑slide images (WSIs). SlideTIM extends the LC‑TIM approach by adding a spatial‑latent regularizer and a class‑distribution prior, ensuring that spatially and semantically similar patches receive consistent predictions and that predicted class proportions are calibrated. Experiments on four histology datasets show that SlideTIM outperforms existing TIM variants, boosting macro‑F1 scores by up to 8.1 percentage points over the best baseline and 19.4 percentage points over zero‑shot predictions at one shot.

By Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq, Christophe De Vleeschouwer