arXiv:2607. 14703v1 Announce Type: cross Abstract: Multiple instance learning (MIL) has become the main paradigm for whole-slide image (WSI) analysis in computational pathology.
By Mingxi Fu, Jiawen Li, Renao Yan, Jiali Hu, Qiehe Sun, Tian Guan, Yonghong He
arXiv:2607. 11257v1 Announce Type: cross Abstract: Pathology Foundation Models (PFMs) offer powerful Whole Slide Image (WSI) representations but suffer from massive computational costs.
By Gangsu Kim, Won-Ki Jeong
WILSON is a vision–language foundation model that represents whole‑slide images and multi‑slide patient cases as single multi‑magnification composite images. Trained on about 189,000 Mayo Clinic slides covering 42 organs and 829 diagnostic entities, it outperforms dedicated case‑level models on internal cohorts and matches slide‑level models while using far less compute. Fine‑tuning on triple‑negative breast cancer data improves histologic subtyping and lymphocyte grading, and the model retrieves diagnostic text with high recall and generates captions closer to report references than prior methods.
By Saghir Alfasly, Wataru Uegami, Sobhan Hemati, Wenchao Han, Xiaojia Tang, Kevin Thompson, Daniel Stone, Ghazal Alabtah, Saba Yasir, Michael R. Lucas, Eric W. Klee, Cheryl L. Willman, Judy C. Boughey, Matthew P. Goetz, Krishna R. Kalari, H. R. Tizhoosh
LanGuSTE is a patch‑selection framework for whole slide image analysis that uses vision‑language models and large language model knowledge. It introduces Cross‑Scale Visual Prompt Tuning to align low‑resolution and high‑resolution patches, and a coarse‑to‑fine selection module that encodes only informative high‑resolution patches. Experiments show LanGuSTE cuts overall processing time to about one‑third of the baseline while matching or surpassing diagnostic performance of exhaustive and state‑of‑the‑art methods.
By Yonghan Shin, Gangsu Kim, Won-Ki Jeong
HERO (Histology Encoder for Robust Representation in Oncology) is a ViT‑G/14 pathology foundation model trained with DINO and iBOT objectives and refined using high‑resolution Gram anchoring on a 500‑million‑tile corpus from about 575,000 clinical whole‑slide images. It demonstrates superior robustness to center, scanner, and stain variation compared to other state‑of‑the‑art foundation models, while maintaining competitive performance on tile‑level classification, segmentation, and gene‑expression prediction. Across 39 slide‑level clinical tasks, HERO ranks first on average and achieves the best average rank across six benchmark frameworks under an equal‑weighted analysis.
By Zhi Li (Caris Life Sciences, Irving, TX, United States), Eghbal Amidi (Caris Life Sciences, Irving, TX, United States), Yating Cheng (Caris Life Sciences, Irving, TX, United States), Tyson Dawson (Caris Life Sciences, Irving, TX, United States), Gorkem Can Ates (Caris Life Sciences, Irving, TX, United States), Shuzhen Kuang (Caris Life Sciences, Irving, TX, United States), Norsang Lama (Caris Life Sciences, Irving, TX, United States), Md Ashequr Rahman (Caris Life Sciences, Irving, TX, United States), Zhiying Lu (Caris Life Sciences, Irving, TX, United States), Elisabeth K. Kong (Caris Life Sciences, Irving, TX, United States), Milan Radovich (Caris Life Sciences, Irving, TX, United States), David Spetzler (Caris Life Sciences, Irving, TX, United States), Matthew Oberley (Caris Life Sciences, Irving, TX, United States), George W. Sledge (Caris Life Sciences, Irving, TX, United States), Ming Chen (Caris Life Sciences, Irving, TX, United States)
arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.
By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv:2609.00866v1 Announce Type: cross
Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pa...
By Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balc{\i}, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan S\"okmen, Ece Tu\u{g}ba Cebeci, Ahmet Hal{\i}c{\i}, Musa Balc{\i}, Kardelen Pe\c{c}enek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn
SCOUT is a concept‑grounded multimodal transformer that generates whole‑slide pathology reports by integrating local histological patterns, whole‑slide context, and expert‑curated diagnostic concepts. It uses evolving visual representations and recursively updated slide‑ and concept‑conditioned representations, with separate attention pathways during decoding that are fused adaptively for each token. Evaluated on TCGA‑BRCA, HistAI, and REG‑2025, SCOUT outperformed existing methods, improving BLEU, METEOR, and ROUGE‑L scores and raising the Clinical Report Quality Score on REG‑2025.
By Suryakant Singh, Saarthak Kapse, Joel Saltz, Prateek Prasanna
arXiv:2609.37682v1 Announce Type: new
Abstract: The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the...
By Chu Zhang, Haoyu Jiang, Hongyuan Zhang, Hongbin Liu, Dong Yi
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the sc...
arXiv:2608.22066v1 Announce Type: cross
Abstract: Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is effective for slide-level tasks, but exhaustive inf...
By Duncan Stothers, Ren-Chin Wu, William Lotter
arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.
By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang