SlideBank is a training‑free framework that turns each whole‑slide image into a persistent, concept‑indexed evidence bank. It performs coarse‑to‑fine exploration to locate informative regions and multi‑scale views, converts them into explicit morphological observations, and anchors pathology signals to the supporting patches and slide coordinates. During inference, questions are routed to relevant signals and evidence scales, and a confidence‑based cross‑level consensus integrates global, regional, and patch evidence, achieving high accuracy on WSI‑VQA and SlideBench‑BCNB while enabling consistent re‑phrasing and reduced inference cost.
By Beidi Zhao, Gexin Huang, Ciro Zhang, Anqi Li, Yusheng Tan, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li
arXiv:2607.19261v4 Announce Type: replace-cross
Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating...
By Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin
The paper introduces SlideTIM, a transductive few‑shot classification method tailored for whole‑slide images (WSIs). SlideTIM extends the LC‑TIM approach by adding a spatial‑latent regularizer and a class‑distribution prior, ensuring that spatially and semantically similar patches receive consistent predictions and that predicted class proportions are calibrated. Experiments on four histology datasets show that SlideTIM outperforms existing TIM variants, boosting macro‑F1 scores by up to 8.1 percentage points over the best baseline and 19.4 percentage points over zero‑shot predictions at one shot.
By Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq, Christophe De Vleeschouwer
arXiv:2608. 14719v1 Announce Type: cross Abstract: Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis.
By Xiaoxiao Li, Xitong Ling, Jiawen Li, Weiming Chen, Zhenyang Cai, Xidong Wang, Tian Guan, Benyou Wang, Yonghong He
arXiv:2609.00396v1 Announce Type: new
Abstract: Histopathological whole slide images (WSIs) are central to cancer diagnosis, but their gigapixel scale, tissue heterogeneity, weak slide-level supervis...
By Chad Wong, Sicheng Chen, Tianyi Zhang, Enhui Chai, Yueming Jin, Zeyu Liu, Fei Xia
arXiv:2609.16284v1 Announce Type: cross
Abstract: Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However,...
By Yan Zhu, Yongbo Chen, Zhengming Ding, Rebecca Faust
The paper introduces FFM-CP, a framework that fuses multiple pathology vision‑language foundation models for few‑shot learning. It aligns heterogeneous representations with an Orthogonal Procrustes transformation, then uses a unified graph to refine support‑image features and class prototypes across backbones. Experiments on six histopathology datasets show that FFM‑CP outperforms the best single adapted model in 50 of 54 few‑shot comparisons.
By Anh-Tien Nguyen, Trung DQ. Dang, Nghiem Tuong Diep, Bui Ngoc Han Nguyen, Tan-Ha Mai, Miriam Cindy Maurer, Phuong Hoa Nguyen, Thi Thuy Uyen Nguyen, Youngjun Park, Daniel Sonntag, Duy Minh Ho Nguyen, Anne-Christin Hauschild
arXiv:2608.30420v1 Announce Type: cross
Abstract: Automating the analysis of whole-slide images has high clinical value, since characterizing cancers requires examining them in detail. Such analysis...
By Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Benoit Macq, Christophe De Vleeschouwer
arXiv:2608. 04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis.
By Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology benchmarks evaluate models on pre-cropped patches or pre-extracted slide features, leaving their ability to acquire evidence directly from gigapixel WSIs largely untested.
Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does...
arXiv:2607. 19261v1 Announce Type: cross Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence.
By Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin