arXiv AI

Aligning Cellular Sheaves with Classifier Attention for Interpretable Weakly-Supervised Pathology Localization

arXiv:2606. 00092v1 Announce Type: cross Abstract: Weakly-supervised classification of whole-slide images with attention-based multiple instance learning (ABMIL) on top of foundation features now reaches near-saturation on Camelyon16 slide-level performance, but the corresponding attention maps are an imperfect localization signal: in clinical interpretation, a model that classifies correctly without firing on the actual lesion is hard to trust.

arXiv Machine Learning
Aug 26

RACR-MIL: Rank-aware contextual reasoning for weakly supervised grading of squamous cell carcinoma using whole slide images

RACR-MIL is a weakly‑supervised method for grading squamous cell carcinoma (SCC) from whole‑slide images, using an attention‑based multiple‑instance learning framework. It introduces a hybrid WSI graph to capture local tissue context and non‑local phenotypic dependencies, and applies rank‑ordering constraints on attention to prioritize higher‑grade tumor regions, mirroring pathologists’ diagnostic reasoning. The approach achieves state‑of‑the‑art performance, improving SCC grading accuracy by 3–9% over existing methods and up to 10% in tumor localization, and a pilot study showed pathologists reported increased grading efficiency in 60% of cases.

By Anirudh Choudhary, Mosbah Aouad, Krishnakant Saboo, Angelina Hwang, Jacob Kechter, Blake Bordeaux, Puneet Bhullar, David DiCaudo, Steven Nelson, Nneka Comfere, Emma Johnson, Olayemi Sokumbi, Jason Sluzevich, Leah Swanson, Dennis Murphree, Aaron Mangold, Ravishankar Iyer
arXiv Computer Vision
6d ago

MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization

MDSkin-Net is a multi‑task skin lesion analysis framework that integrates Pattern Analysis priors into a hybrid CNN‑Transformer architecture. It introduces a Pattern Analysis‑Guided Attention Module (PAGAM) with improved Efficient Channel Attention, Multi‑Scale Spatial Attention, and Biased Asymmetry Attention, along with a multi‑scale spatial alignment regularization that uses segmentation masks as soft supervision. Trained only on the ISIC 2017 training split, the model achieves high segmentation and classification performance on multiple datasets, demonstrating strong zero‑shot generalization across different cohorts.

By Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas, Nikolaos Papanikolopoulos
arXiv Computer Vision
3d ago

SheafStain: Sheaf-Theoretic Schr\"odinger Bridge for Spatially and Biologically Coherent Virtual Staining

SheafStain introduces a sheaf-theoretic Schr"odinger Bridge framework to improve virtual staining of gigapixel whole-slide images. By treating Vision Foundation Model features as sheaf-like sections, it integrates class and patch tokens to enforce spatial and biological coherence, mitigating patch-boundary artifacts. The method is evaluated on HER2, ER, PR, and Ki‑67 stains, outperforming six prior approaches on stitched 1024 × 1024 outputs.

By Hyeongyeol Lim, Hongjun Yoon, Eunjin Jang, Daeky Jeong, Won June Cho, Hwamin Lee
arXiv Computer Vision
Sep 15

Weakly Supervised Spatial Grounding for Discriminative Attention-Based Ultrasound-Histopathology Alignment in Prostate Cancer Grading

arXiv:2609.15150v1 Announce Type: new Abstract: Unpaired cross-modal distillation transfers grade structure from histopathology into a micro-ultrasound (micro-US) encoder by aligning a pooled needle-...

By Obed Korshie Dzikunu, Emma Willis, Mohammad Mahdi Abootorabi, Mohamed Harmanani, Zhuoxin Guo, Ferdinand Luger, Adam Kinnaird, Brian Wodlinger, Parvin Mousavi, Purang Abolmaesumi
arXiv Computer Vision
Sep 7

An Attention-Guided Global and Local Fusion Framework for Lesion-Focused Image Classification

The paper introduces an attention‑guided fusion framework that combines global and lesion‑focused local information for image classification. Using a three‑branch architecture built on DenseNet‑121, the model generates attention maps with Grad‑CAM, refines local features with CBAM, and adaptively fuses the two representations. Experiments on synthetic and real datasets, including skin, guava leaf, and grape leaf images, show that the fusion branch outperforms individual branches, achieving up to 97.75% accuracy on skin lesions and 99.64% on guava leaves.

By Mst Shafia Tasnima, Md Samaun Elaheea, Tanjim Taharat Aurpab, Md Musfique Anwar
arXiv Computer Vision
Sep 7

Compositional Reward Models for Conditional Medical Image Generation

The paper introduces PRISM, a Compositional Reward Model framework that decomposes image quality into multiple verifier‑grounded stages for conditional medical image generation. By assigning distinct rewards for fine‑to‑coarse properties—such as intensity, texture, structural alignment, and semantic fidelity—and combining them via a Hierarchical Constrained Propagation mechanism, PRISM addresses shortcomings of single‑scalar reward approaches. Experiments on PanNuke, CeDeM, and ISIC datasets show that data generated with PRISM improves downstream model performance, achieving higher mDice, lower MRE, and increased F1 scores compared to baseline methods.

By Aayush Kumar Tyagi, Prathosh A. P., Mausam
arXiv Computer Vision
Sep 3

Seeing Beyond the Lesion: Disease Recognition from Reactive CNS Tissue

The study evaluates whether disease can be identified from reactive, non‑lesional brain tissue in intracranial biopsies. Using four foundation‑model encoders within an attention‑based multiple‑instance learning framework on 245 whole‑slide images, the authors find that disease labels remain predictive even after controlling for slide size and sampling bias, and that performance is similar across all encoders. Signed instance‑contribution maps and expert review confirm that predictive signals localize to reactive parenchyma rather than artifacts such as blood. "whyItMatters":"The findings demonstrate that weakly supervised models can recover disease signals from tissue traditionally considered non‑diagnostic, highlighting the need for provenance‑only baselines in computational pathology benchmarks."

By Jan Schnorrenberg, Jan Ernsting, Enrico K\"ullenberg, Tim Hahn, Benjamin Risse, Christian Thomas