arXiv AI

Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images

arXiv AI
Sep 10

Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation

SCOUT is a concept‑grounded multimodal transformer that generates whole‑slide pathology reports by integrating local histological patterns, whole‑slide context, and expert‑curated diagnostic concepts. It uses evolving visual representations and recursively updated slide‑ and concept‑conditioned representations, with separate attention pathways during decoding that are fused adaptively for each token. Evaluated on TCGA‑BRCA, HistAI, and REG‑2025, SCOUT outperformed existing methods, improving BLEU, METEOR, and ROUGE‑L scores and raising the Clinical Report Quality Score on REG‑2025.

By Suryakant Singh, Saarthak Kapse, Joel Saltz, Prateek Prasanna
arXiv Computer Vision
Aug 25

LanGuSTE: Language-Guided Coarse-to-Fine Patch Selection for Efficient Whole Slide Image Analysis

LanGuSTE is a patch‑selection framework for whole slide image analysis that uses vision‑language models and large language model knowledge. It introduces Cross‑Scale Visual Prompt Tuning to align low‑resolution and high‑resolution patches, and a coarse‑to‑fine selection module that encodes only informative high‑resolution patches. Experiments show LanGuSTE cuts overall processing time to about one‑third of the baseline while matching or surpassing diagnostic performance of exhaustive and state‑of‑the‑art methods.

By Yonghan Shin, Gangsu Kim, Won-Ki Jeong
arXiv AI
Sep 2

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

arXiv:2609.00866v1 Announce Type: cross Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pa...

By Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balc{\i}, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan S\"okmen, Ece Tu\u{g}ba Cebeci, Ahmet Hal{\i}c{\i}, Musa Balc{\i}, Kardelen Pe\c{c}enek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn
arXiv Computer Vision
Sep 15

Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment

arXiv:2510.23224v2 Announce Type: replace Abstract: The rapid digitization of histopathology slides has opened new opportunities for computational tools in clinical and research workflows. Content-ba...

By Hongyi Wang, Zhengjie Zhu, Junlin Hou, Jiabo Ma, Fang Wang, Yue Shi, Qiuyu Cai, Jili Wang, Bo Luo, Zhizhong Chai, Zhengyu Zhang, Li Liang, Xiuming Zhang, Yen-Wei Chen, Lanfen Lin, Hao Chen
arXiv Computer Vision
Aug 25

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

arXiv:2409.16183v2 Announce Type: replace Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...

By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang