arXiv Computer Vision

PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images

PathoHR is a new pipeline for predicting breast cancer survival from high‑resolution pathological images. It uses a plug‑and‑play Vision Transformer to enhance patch‑wise whole slide image representations, evaluates multiple similarity metrics to optimize feature learning, and shows that smaller, enhanced patches can match or surpass the accuracy of larger raw patches while cutting computational cost. The authors provide experimental evidence that this approach improves both accuracy and efficiency in computational pathology.

arXiv AI
2d ago

A Lightweight CNN Integrated Compact Convolutional Transformer for Multi-Scale Feature Learning and reducing computational complexity for breast cancer mammography image detection and classification

The paper presents a lightweight CNN‑integrated Compact Convolutional Transformer (CCT) designed for multi‑scale feature learning in breast cancer mammography. With only 250,435 parameters, the model achieved 99‑100% accuracy across three datasets using 5‑fold cross‑validation, demonstrating robust generalization. Explainable AI components were added to clarify the classification process, aiming to increase clinical trust in resource‑constrained settings.

By Md Taimur Ahad (Department of Management North South University, Dhaka, Bangladesh), Ainuddin Ahmed (Department of Management North South University, Dhaka, Bangladesh)
arXiv Machine Learning
Aug 19

MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology

MagViT is an interpretable multi‑magnification transformer that classifies breast histopathology images by extracting representations from four BreakHis magnifications (40X, 100X, 200X, 400X) and fusing them with a learnable, scale‑gated mechanism that can mask missing scales. The model selects the most accurate architectural branch at the patient level using five‑fold cross‑validation, achieving high performance on BreakHis (mean image accuracy 0.9191, patient accuracy 0.9643, macro‑F1 0.9042) and demonstrating preliminary cross‑dataset generalization on BUSI and IDC. Grad‑CAM visualizations confirm that the network focuses on diagnostically relevant regions across magnifications.

By Nabil Ashab, Soumit Kumar Kundu, Saif Mahmud Parvez, Shahadat Hossain Sohag, Bidhan Biswas, Nazmus Subha
arXiv Machine Learning
Sep 3

Morphology signal in whole slide image foundation models can automatically triage slides

The paper introduces a pipeline that uses publicly available whole slide image foundation models (FMs) to automatically triage slides by ranking them based on zero‑shot classification predictions. This approach accurately identifies slides containing the most tumor, achieving top‑2 ranking for patients with up to 43 slides across multiple datasets. The study also proposes a ranked evaluation framework to benchmark FM performance in slide triage.

By Ayushi Sinha, Shashank Yadav, Benjamin Holmes, Pravat Das, Aaron W. Bogan, James S. Lewis Jr., Santiago Romero-Brufau, Andrew Y. K. Foong, Scott H. Kaufmann, Kathryn M. Van Abel, David M. Routman, Michael R. Lucas
arXiv AI
Jun 8

DaX: Learning General Pathology Representations Across Scales

arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.

By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv Machine Learning
Jul 14

BiLoG-Net: A Bi-Context Location-Guided Network for Breast Mass Segmentation and Malignancy Classification in Mammography

arXiv:2607. 10188v1 Announce Type: cross Abstract: Breast cancer remains the most commonly diagnosed malignancy among women worldwide, yet accurate detection and characterization of breast masses in mammography remain challenging due to subtle intensity variations, heterogeneous tissue densities, and indistinct lesion boundaries that complicate radiological interpretation.

By Abu Fatema Mohammad Abdun Noor, Md Imam Ahasan, Md Samiul Ahasan, Kah Ong Michael Goh, S M Hasan Mahmud, Raihana Zannat
arXiv AI
Jul 21

GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

arXiv:2607. 18218v1 Announce Type: cross Abstract: Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning transferable representations from large-scale histopathology data.

By Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer, Cliff Wong, Soohee Lee, Hao Qiu, Theodore Zhengde Zhao, Racheli Ben Shimol, Angela Crabtree, Kevin Matlock, Eduardo Alejandro Lozano Garcia, Naiteek Sangani, Alberto Santamaria-Pang, Jason Entenmann, Alexandra Q. Bartlett, Bill J. Wright, Bernard A. Fox, Brian Piening, Sheng Zhang, Sheng Wang, Tristan Naumann, Carlo Bifulco, Hoifung Poon
arXiv Computer Vision
Aug 25

Extending the Horizon of Early Diagnosis: Lung Cancer Prediction with Vision Transformers

arXiv:2608.21571v1 Announce Type: new Abstract: Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage...

By Olivera Kotevska, Ian Goethert, Michael McGee, Maria Mahbub, Sean R. Wilkinson, Rowena Yip, Myvizhi Esai Selvan, Zeynep H. Gumus, Claudia Henschke, Robert J. Klein, Providencia Morales, Samuel M Aguayo, Ioana Danciu, Mayanka Chandrashekar
arXiv Computer Vision
Sep 3

AtlasPatch: Scalable Foundation Model-based Tissue Detection and Patch Extraction for Computational Pathology

AtlasPatch is a scalable, high‑throughput whole‑slide image preprocessing method that uses a foundation‑model‑based tissue detector operating at thumbnail resolution. By updating only 0.076% of the SAM2 model weights and leveraging a curated dataset of 30,000 thumbnail‑mask pairs, it generates accurate tissue masks and directly produces patch coordinates at the desired magnification, eliminating repeated patch‑level inference. The approach achieves 0.986 precision, is up to 16× faster than existing deep‑learning methods, and maintains downstream multiple‑instance learning performance across six slide‑level classification tasks.

By Ahmed Alagha, Christopher Leclerc, Yousef Kotp, Omar Metwally, Calvin Moras, Peter Rentopoulos, Ghodsiyeh Rostami, Bich Ngoc Nguyen, Jumanah Baig, Abdelhakim Khellaf, Vincent Quoc-Huy Trinh, Rabeb Mizouni, Hadi Otrok, Jamal Bentahar, Mahdi S. Hosseini