arXiv:2606. 00092v1 Announce Type: cross Abstract: Weakly-supervised classification of whole-slide images with attention-based multiple instance learning (ABMIL) on top of foundation features now reaches near-saturation on Camelyon16 slide-level performance, but the corresponding attention maps are an imperfect localization signal: in clinical interpretation, a model that classifies correctly without firing on the actual lesion is hard to trust.
By Devansh Lalwani, Swapnil Bhat, Maulik Shah
arXiv:2609.36429v1 Announce Type: new
Abstract: Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods...
By Zijun Gao, Chunbin Gu, Jinxi Xiang, Xiangde Luo, Pheng-Ann Heng
arXiv:2606. 03644v1 Announce Type: new Abstract: Comprehensive molecular profiling is essential for modern precision oncology but remains hindered by prohibitive costs, specimen exhaustion, and protracted turnaround times.
By Fengtao Zhou, Yingxue Xu, Zhengyu Zhang, Yihui Wang, Zhengrui Guo, Ling Liang, Jiabo Ma, Cheng Jin, Ziyi Liu, Huajun Zhou, Hongyi Wang, Du Cai, Chenglong Zhao, Xi Wang, Can Yang, Yu Wang, Wenbin Li, Feng Gao, Zhe Wang, Zhenhui Li, Xiuming Zhang, Li Liang, Hao Chen
HERO (Histology Encoder for Robust Representation in Oncology) is a ViT‑G/14 pathology foundation model trained with DINO and iBOT objectives and refined using high‑resolution Gram anchoring on a 500‑million‑tile corpus from about 575,000 clinical whole‑slide images. It demonstrates superior robustness to center, scanner, and stain variation compared to other state‑of‑the‑art foundation models, while maintaining competitive performance on tile‑level classification, segmentation, and gene‑expression prediction. Across 39 slide‑level clinical tasks, HERO ranks first on average and achieves the best average rank across six benchmark frameworks under an equal‑weighted analysis.
By Zhi Li (Caris Life Sciences, Irving, TX, United States), Eghbal Amidi (Caris Life Sciences, Irving, TX, United States), Yating Cheng (Caris Life Sciences, Irving, TX, United States), Tyson Dawson (Caris Life Sciences, Irving, TX, United States), Gorkem Can Ates (Caris Life Sciences, Irving, TX, United States), Shuzhen Kuang (Caris Life Sciences, Irving, TX, United States), Norsang Lama (Caris Life Sciences, Irving, TX, United States), Md Ashequr Rahman (Caris Life Sciences, Irving, TX, United States), Zhiying Lu (Caris Life Sciences, Irving, TX, United States), Elisabeth K. Kong (Caris Life Sciences, Irving, TX, United States), Milan Radovich (Caris Life Sciences, Irving, TX, United States), David Spetzler (Caris Life Sciences, Irving, TX, United States), Matthew Oberley (Caris Life Sciences, Irving, TX, United States), George W. Sledge (Caris Life Sciences, Irving, TX, United States), Ming Chen (Caris Life Sciences, Irving, TX, United States)
arXiv:2608.30420v1 Announce Type: cross
Abstract: Automating the analysis of whole-slide images has high clinical value, since characterizing cancers requires examining them in detail. Such analysis...
By Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Benoit Macq, Christophe De Vleeschouwer
arXiv:2606. 17702v1 Announce Type: cross Abstract: Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extraction, and interpretable clinical reporting.
By Wan Siti Halimatul Munirah Wan Ahmad, Faris Syahmi Samidi, Mohammad Badal Ahmmed, Vimal Angela Thiviyanathan, Selvam James Thavaraj, Anwar P. P. Abdul Majeed
AtlasPatch is a scalable, high‑throughput whole‑slide image preprocessing method that uses a foundation‑model‑based tissue detector operating at thumbnail resolution. By updating only 0.076% of the SAM2 model weights and leveraging a curated dataset of 30,000 thumbnail‑mask pairs, it generates accurate tissue masks and directly produces patch coordinates at the desired magnification, eliminating repeated patch‑level inference. The approach achieves 0.986 precision, is up to 16× faster than existing deep‑learning methods, and maintains downstream multiple‑instance learning performance across six slide‑level classification tasks.
By Ahmed Alagha, Christopher Leclerc, Yousef Kotp, Omar Metwally, Calvin Moras, Peter Rentopoulos, Ghodsiyeh Rostami, Bich Ngoc Nguyen, Jumanah Baig, Abdelhakim Khellaf, Vincent Quoc-Huy Trinh, Rabeb Mizouni, Hadi Otrok, Jamal Bentahar, Mahdi S. Hosseini
The paper introduces CytoCRF, a conditional random field framework tailored for cytology images. It adapts pairwise terms to focus on chromatin and cytology-specific staining and enriches neighborhood information by combining multiple backbone models. Across ten cytology datasets, CytoCRF surpasses existing CRF methods at all annotation budgets, achieving up to +13.6 percentage points over the best baseline and +33.7 over zero‑shot performance with only 50 annotations.
By Manon Dausort, Tiffanie Godelaine, Karim El Khoury, Maxime Zanella, Christophe De Vleeschouwer, Beno\^it Macq
The paper introduces the Consistency Memory Bank (COMB), a label‑free virtual staining framework designed to process gigapixel Whole Slide Images without the memory bottlenecks of patch‑based deep learning. COMB decouples context storage from computation, using a dynamic retrieval mechanism to fetch feature representations from adjacent tiles, local padding to resolve spatial discontinuities, and neighbor‑aware channel attention to stabilize statistical drift. The method achieves superior perceptual fidelity and tiling consistency compared to state‑of‑the‑art baselines, and its improved continuity suggests downstream benefits for tumor segmentation.
By Dou Hoon Kwark, Kianoush Falahkheirkhah, Ji-hun Oh, Shirui Luo, Volodymyr Kindratenko, Rohit Bhargava
arXiv:2603. 12433v3 Announce Type: replace-cross Abstract: Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility.
By Zheda Mai, Ke Zhang, Fu-En Wang, Zixiao Ken Wang, Albert Y. C. Chen, Lu Xia, Min Sun, Wei-Lun Chao, Cheng-Hao Kuo
LanGuSTE is a patch‑selection framework for whole slide image analysis that uses vision‑language models and large language model knowledge. It introduces Cross‑Scale Visual Prompt Tuning to align low‑resolution and high‑resolution patches, and a coarse‑to‑fine selection module that encodes only informative high‑resolution patches. Experiments show LanGuSTE cuts overall processing time to about one‑third of the baseline while matching or surpassing diagnostic performance of exhaustive and state‑of‑the‑art methods.
By Yonghan Shin, Gangsu Kim, Won-Ki Jeong
arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.
By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu