RACR-MIL is a weakly‑supervised method for grading squamous cell carcinoma (SCC) from whole‑slide images, using an attention‑based multiple‑instance learning framework. It introduces a hybrid WSI graph to capture local tissue context and non‑local phenotypic dependencies, and applies rank‑ordering constraints on attention to prioritize higher‑grade tumor regions, mirroring pathologists’ diagnostic reasoning. The approach achieves state‑of‑the‑art performance, improving SCC grading accuracy by 3–9% over existing methods and up to 10% in tumor localization, and a pilot study showed pathologists reported increased grading efficiency in 60% of cases.
By Anirudh Choudhary, Mosbah Aouad, Krishnakant Saboo, Angelina Hwang, Jacob Kechter, Blake Bordeaux, Puneet Bhullar, David DiCaudo, Steven Nelson, Nneka Comfere, Emma Johnson, Olayemi Sokumbi, Jason Sluzevich, Leah Swanson, Dennis Murphree, Aaron Mangold, Ravishankar Iyer
arXiv:2608.08374v3 Announce Type: replace
Abstract: Transformer models are increasingly used for whole-slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural ima...
By Zhiyuan Yang, Jiahao Cheng, Vincent Quoc-Huy Trinh, Mahdi S. Hosseini
MDSkin-Net is a multi‑task skin lesion analysis framework that integrates Pattern Analysis priors into a hybrid CNN‑Transformer architecture. It introduces a Pattern Analysis‑Guided Attention Module (PAGAM) with improved Efficient Channel Attention, Multi‑Scale Spatial Attention, and Biased Asymmetry Attention, along with a multi‑scale spatial alignment regularization that uses segmentation masks as soft supervision. Trained only on the ISIC 2017 training split, the model achieves high segmentation and classification performance on multiple datasets, demonstrating strong zero‑shot generalization across different cohorts.
By Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas, Nikolaos Papanikolopoulos
arXiv:2607. 00385v2 Announce Type: replace-cross Abstract: Automated malaria diagnosis from blood smear microscopy is a critical global health AI challenge; expert scarcity remains the primary diagnostic bottleneck.
By Kaysarul Anas Apurba, Md Hasibul Hasan, Mohammed Ali, Tanzilur Rahman
SheafStain introduces a sheaf-theoretic Schr"odinger Bridge framework to improve virtual staining of gigapixel whole-slide images. By treating Vision Foundation Model features as sheaf-like sections, it integrates class and patch tokens to enforce spatial and biological coherence, mitigating patch-boundary artifacts. The method is evaluated on HER2, ER, PR, and Ki‑67 stains, outperforming six prior approaches on stitched 1024 × 1024 outputs.
By Hyeongyeol Lim, Hongjun Yoon, Eunjin Jang, Daeky Jeong, Won June Cho, Hwamin Lee
arXiv:2608.30420v1 Announce Type: cross
Abstract: Automating the analysis of whole-slide images has high clinical value, since characterizing cancers requires examining them in detail. Such analysis...
By Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Benoit Macq, Christophe De Vleeschouwer
arXiv:2609.15150v1 Announce Type: new
Abstract: Unpaired cross-modal distillation transfers grade structure from histopathology into a micro-ultrasound (micro-US) encoder by aligning a pooled needle-...
By Obed Korshie Dzikunu, Emma Willis, Mohammad Mahdi Abootorabi, Mohamed Harmanani, Zhuoxin Guo, Ferdinand Luger, Adam Kinnaird, Brian Wodlinger, Parvin Mousavi, Purang Abolmaesumi
The paper introduces an attention‑guided fusion framework that combines global and lesion‑focused local information for image classification. Using a three‑branch architecture built on DenseNet‑121, the model generates attention maps with Grad‑CAM, refines local features with CBAM, and adaptively fuses the two representations. Experiments on synthetic and real datasets, including skin, guava leaf, and grape leaf images, show that the fusion branch outperforms individual branches, achieving up to 97.75% accuracy on skin lesions and 99.64% on guava leaves.
By Mst Shafia Tasnima, Md Samaun Elaheea, Tanjim Taharat Aurpab, Md Musfique Anwar
The paper introduces PRISM, a Compositional Reward Model framework that decomposes image quality into multiple verifier‑grounded stages for conditional medical image generation. By assigning distinct rewards for fine‑to‑coarse properties—such as intensity, texture, structural alignment, and semantic fidelity—and combining them via a Hierarchical Constrained Propagation mechanism, PRISM addresses shortcomings of single‑scalar reward approaches. Experiments on PanNuke, CeDeM, and ISIC datasets show that data generated with PRISM improves downstream model performance, achieving higher mDice, lower MRE, and increased F1 scores compared to baseline methods.
By Aayush Kumar Tyagi, Prathosh A. P., Mausam
The study evaluates whether disease can be identified from reactive, non‑lesional brain tissue in intracranial biopsies. Using four foundation‑model encoders within an attention‑based multiple‑instance learning framework on 245 whole‑slide images, the authors find that disease labels remain predictive even after controlling for slide size and sampling bias, and that performance is similar across all encoders. Signed instance‑contribution maps and expert review confirm that predictive signals localize to reactive parenchyma rather than artifacts such as blood.
"whyItMatters":"The findings demonstrate that weakly supervised models can recover disease signals from tissue traditionally considered non‑diagnostic, highlighting the need for provenance‑only baselines in computational pathology benchmarks."
By Jan Schnorrenberg, Jan Ernsting, Enrico K\"ullenberg, Tim Hahn, Benjamin Risse, Christian Thomas
arXiv:2607.10851v2 Announce Type: replace
Abstract: Medical image classification models are ideally expected to identify diagnostically relevant regions while making predictions, yet standard classif...
By Tonmoy Hossain, Atiqur Rahman, Farhana Hossain Swarnali, Miaomiao Zhang
arXiv:2608.22066v1 Announce Type: cross
Abstract: Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is effective for slide-level tasks, but exhaustive inf...
By Duncan Stothers, Ren-Chin Wu, William Lotter